Unlock Full Resume Report

New offer - be the first one to apply!

September 24, 2026

Data Engineer (Databricks, dbt)

Mid • On-site

Warsaw, Poland

Quick Facts

  • Role: Regular Data Engineer (Databricks, dbt)

  • Focus: LLM-enabled document processing and automated decision support via training datasets

  • Work model: Hybrid (at least 1 day/week on-site)

Description

Build end-to-end data pipelines on a modern cloud data platform (Databricks) to create training datasets for machine learning and enable automated decision-making powered by Large Language Models (LLMs). You will design ingestion, transformation, storage, and consumption layers, maintain high-quality training datasets, and apply best practices in data engineering, testing, and deployment. Work closely with data scientists, engineers, and business stakeholders, including direct interaction with business users.

Responsibilities

  • Build a training dataset for document processing and decision support

  • Design and implement end-to-end data pipelines (ingestion, transformation, storage, and consumption)

  • Prepare and maintain high-quality training datasets for machine learning models

  • Work with large-scale data on Databricks

  • Apply best practices in data engineering, testing, and deployment

  • Collaborate with data scientists, engineers, and business stakeholders

  • Continuously improve performance, reliability, and automation of data workflows

Requirements

  • Regular-level experience in Data Engineering

  • Bachelor’s or Master’s degree in a technical field (e.g., Computer Science, Engineering) or equivalent experience

  • Fluent English (written and spoken)

  • DevOps mindset (“you build it, you run it”)

  • Ability to understand complex requirements and translate them into actionable solutions

  • Strong communication skills and ability to work with cross-functional teams

  • Attention to detail, especially regarding data quality and business logic

  • Proactive attitude, ownership, and willingness to learn new technologies

  • Strong experience with Databricks (DBX)

  • Advanced knowledge of Python

  • Very good knowledge of SQL and relational databases

  • Experience building and optimizing ETL/ELT pipelines

  • Experience with PySpark

Benefits

  • Contract under Polish law: B2B or Umowa o Pracę

  • Private medical care

  • Group insurance

  • Multisport card

  • English classes available

  • Hybrid work (at least 1 day/week on-site)

  • Opportunity to work with excellent professionals and focus on high code quality

  • Continuous learning and growth in an international team

Similar jobs you might like