Unlock Full Resume Report

New offer - be the first one to apply!

September 5, 2026

AI Engineer (Python + AWS + AI/LLM)

Senior • Remote

353,600 - 374,400 PLN/yr

Warsaw, MZ, Poland

Rate: up to 180 PLN/h on B2B

Project Overview

We are building CaaS (Content as a Service), a platform that transforms publisher content—PDF textbooks and Excel manifests—into structured, enriched, AI-ready data.

The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.

Key use cases:

  • RAG-based Teacher Assistant
  • Editorial tooling
  • Future AI-powered student-facing products

The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.

Responsibilities

Data Engineering & Pipelines

  • Build and maintain multi-stage data ingestion pipelines.
  • Design and implement idempotent, restartable batch processing workflows.
  • Use S3 as the core storage layer for raw and processed data.
  • Implement pipeline stages including:
    • Content ingestion and book identity assignment
    • PDF-to-markdown conversion (AI OCR)
    • Table of contents and structure extraction
    • Hierarchical chunking
    • Embedding generation

AI / LLM Processing

  • Use LLMs and OCR models to extract structured data from PDFs.
  • Design prompts and context strategies for consistent outputs.
  • Generate structured metadata and enrich content for downstream use cases.

Data Storage & Consistency

  • Maintain PostgreSQL (Aurora) as the system of record.
  • Design and maintain SQL schemas and versioned migrations.
  • Ensure data consistency across S3, PostgreSQL (Aurora), and the Weaviate vector database.
  • Implement reconciliation logic across distributed systems.

Retrieval & Vector Search

  • Work with Weaviate for vector search and semantic retrieval.
  • Support RAG-based applications.
  • Design data organization strategies by subject, country, and client.

APIs & Integration

  • Build REST APIs using FastAPI.
  • Expose content as a service for multiple downstream applications.
  • Integrate with internal and external systems.

Engineering Practices

  • Write strongly typed Python code using mypy.
  • Follow CI/CD processes with automated checks using ruff and pytest.
  • Work across development, staging, and production environments.
  • Debug distributed data inconsistencies.

Key Requirements

Must-have

  • Strong Python development experience in production systems.
  • AWS experience with S3, Glue, and Aurora.
  • Experience with data pipelines, ETL, or batch processing.
  • Strong SQL and PostgreSQL experience.
  • Experience with schema design and migrations.
  • Production experience with LLMs for OCR, content processing, or enrichment.

Nice-to-have

  • Prompt engineering or context engineering.
  • Experience with vector databases: Weaviate, Pinecone, Qdrant, pgvector.
  • Knowledge of embeddings, semantic search, and RAG.
  • Experience with FastAPI.
  • Experience with Airflow or MWAA.
  • Experience building data platforms serving multiple consumers.

Similar jobs you might like