Unlock Full Resume Report

New offer - be the first one to apply!

September 18, 2026

Senior Data Engineer (Spark / Data Lakehouse Developer)

Senior • On-site

1,100 - 1,300 PLN/yr

Wrocław, DS, Poland

Quick Facts

  • Focus: Data Lakehouse development for big data processing, reconciliation, and data quality
  • Stack: Apache Spark, MongoDB, Apache Iceberg

Description

You will design and develop solutions for processing, synchronizing, and reconciling data in a Big Data environment. The role includes co-creating a modern Data Lakehouse architecture and ensuring quality, performance, and reliability of data processes.

Responsibilities

  • Design and implement data reconciliation processes in Apache Spark (PySpark, Spark SQL)
  • Compare ODS data (MongoDB) with source systems and identify discrepancies
  • Create and develop data quality rules, detect missing records, duplicates, and anomalies
  • Build and maintain Raw/Bronze/Silver/Gold Data Lakehouse layers
  • Work with Apache Iceberg and manage versioned datasets
  • Optimize Spark performance and data processing costs
  • Build and maintain CI/CD pipelines for Spark applications
  • Implement monitoring, metrics, and alerting for data pipelines
  • Automate deployments using Docker, Kubernetes, and Spark Operator
  • Troubleshoot data quality, consistency, and availability issues
  • Create technical documentation, architecture diagrams, and runbooks
  • Collaborate with Data Engineers, DevOps Engineers, Business Analysts, and product teams
  • Participate in code reviews and mentor less experienced team members

Requirements

  • Minimum 4 years of commercial experience with Apache Spark
  • Very good knowledge of PySpark, Spark SQL, DataFrame API, and Structured Streaming
  • Experience integrating Apache Spark with MongoDB using the MongoDB Spark Connector
  • Practical knowledge of Apache Iceberg and managing versioned data tables
  • Experience building Data Lake or Data Lakehouse solutions
  • Good knowledge of Python
  • Experience with Jenkins and building CI/CD pipelines
  • Knowledge of monitoring and observability tools, especially Prometheus and Grafana
  • Experience with Agile methodologies (Scrum or Kanban)
  • Practical knowledge of Jira and Confluence
  • Willingness to work from the office 1–2 days per week (Wrocław)

Benefits

  • Opportunity to co-create a modern Data Lakehouse architecture and work on end-to-end data processing and data quality automation

Similar jobs you might like