April 8, 2026
Data Engineer ( Spark / Streaming / Java)
Senior • Remote
160 - 200 PLN
Warsaw, Poland
#Data Engineer #Apache Spark #Databricks #Java #Apache Kafka #Batch Processing #Structured Streaming #Azure #SQL #Microservices #CI/CD #Docker #DDD
Are you ready to join our international team as a Data Engineer? We shall tell you why you should...
What product do we develop?
We are building an innovative solution, KMD Elements, on Microsoft Azure cloud dedicated to the energy distribution market (electrical energy, gas, water, utility, and similar types of business). Our customers include institutions and companies operating in the energy market as transmission service operators, market regulators, distribution service operators, energy trading, and retail companies.
KMD Elements delivers components allowing implementation of the full lifecycle of a customer on the energy market: meter data processing, connection to the network, physical network management, change of operator, full billing process support, payment, and debt management, customer communication, and finishing on customer account termination and network disconnection.
The key market advantage of KMD Elements is its ability to support highly flexible, complex billing models as well as scalability to support large volumes of data. Our solution enables energy companies to promote efficient energy generation and usage patterns, supporting sustainable and green energy generation and consumption.
We work with always up-to-date versions of:
• Apache Spark on Azure Databricks
• Apache Kafka
• Delta Lake
• Java
• MS SQL Server and NoSQL storages like Elastic Search, Redis, Azure Data Explorer
• Docker containers
• Azure DevOps and fully automated CI/CD pipelines with Databricks Asset Bundles, ArgoCD, GitOps, Helm charts
• Automated tests
How do we work?
#Agile #Scrum #Teamwork #CleanCode #CodeReview #Feedback #BestPracticies
• We follow Scrum principles in our work – we work in biweekly iterations and produce production-ready functionalities at the end of each iteration – every 3 iterations we plan the next product release
• We have end-to-end responsibility for the features we develop – from business requirements, through design and implementation up to running features on production
• More than 75% of our work is spent on new product features
• Our teams are cross-functional (7-8 persons) – they develop, test and maintain features they have built
• Teams’ own domains in the solution and the corresponding system components
• We value feedback and continuously seek improvements
• We value software best practices and craftsmanship
Product principles:
• Domain model created using domain-driven design principles
• Distributed event-driven architecture / microservices
• Large-scale system for large volumes of data (>100TB data), processed by Apache Spark streaming and batch jobs powered by Databricks platform
Your responsibilities:
• Develop and maintain the leading IT solution for the energy market using Apache Spark, Databricks, Delta Lake, and Apache Kafka
• Have end-to-end responsibility for the full lifecycle of features you develop
• Design technical solutions for business requirements from the product roadmap
• Maintain alignment with architectural principles defined on the project and organizational level
• Ensure optimal performance through continuous monitoring and code optimization.
• Refactor existing code and enhance system architecture to improve maintainability and scalability.
• Design and evolve the test automation strategy, including technology stack and solution architecture.
• Prepare reviews, participate in retrospectives, estimate user stories, and refine features ensuring their readiness for development.
Personal requirements:
• Have 4+ years of Apache Spark experience and have faced various data engineering challenges in batch or streaming
• Have an interest in stream processing with Apache Spark Structured Streaming on top of Apache Kafka
• Have experience leading technical solution designs
• Have experience with distributed systems on a cloud platform
• Have experience with large-scale systems in a microservice architecture
• Are familiar with Git and CI/CD practices and can design or implement the deployment process for your data pipelines
• Possess a proactive approach and can-do attitude
• Are excellent in English and Polish, both written and spoken
• Have a higher education in computer science or a related field
• Are a team player with strong communication skills
Nice to have requirements:
• Apache Spark Structured Streaming
• Azure
• Domain Driven Development
• Docker containers and Kubernetes
• Message brokers (i.e. Kafka) and event-driven architecture
• Agile/Scrum
Our offer:
• Contract type: B2B
• Work Mode: Flexible — this role supports on-site, hybrid, and remote arrangements, depending on your individual preferences.
• Occasional on-site presence may be required — for example, onboard new team members, explore new business domains, or refine requirements in close collaboration with stakeholders or team building activities.
What does the recruitment process look like?
• Phone conversation with Recruitment Partner
• Technical interview with the Hiring Team
• Cognitive test
• Offer
Similar jobs you might like
Technology
Datumo
Data Engineer (GCP)
Mid
Remote
Warsaw, Poland
16,000 - 28,000 PLN
🏢 Summary: The offer is for a Data Engineer responsible for building and optimizing scalable data platforms in cloud environments, mainly on Google Cloud Platform and Azure. The role focuses on designing big data solutions, developing ETL processes, and working with distributed data processing frameworks across international projects. It includes full remote work and opportunities for continuous technical development. 🗂️ Requirements: 3-4 years of commercial programming experience, Experience with Google Cloud Platform, Strong knowledge of Python, Strong knowledge of SQL, Strong knowledge of Scala or Java or Kotlin, Experience with BigQuery, Understanding of big data storage, modeling, processing and scheduling, Experience with Apache Spark or similar framework, Experience in data modeling and data storage, Experience with automated testing, CI/CD and code review, Experience collaborating with business stakeholders, English proficiency at minimum B2 level, Proficiency in Polish 📃 Skills: Python, SQL, Scala, Java, Kotlin, GCP, BigQuery, Spark, CI/CD, Git, Azure, Databricks, Kafka, Airflow, Docker, Kubernetes, dbt, PubSub, Dataflow, Flink 🏢 Description: We’re looking for a Data Engineer ready to push boundaries and grow with us. Datumo specializes in providing Data Engineering and Cloud Computing consulting services to clients from all over the world, primarily in Western Europe, Poland and the USA. Core industries we support include e-commerce 🛒, telecommunications 📡 and life sciences 🧬. Our team consists of exceptional people whose commitment allows us to conduct highly demanding projects. Our team members tend to stick around for more than 3 years, and when a project wraps up, we don't let them go - we embark on a journey to discover exciting new challenges for them. It's not just a workplace; it's a community that grows together! Must-have: ✅ at least 3-4 years of commercial experience in programming, ✅ proven record with a cloud provider - Google Cloud Platform, ✅ strong knowledge of Python, SQL and JVM languages (Scala or Java or Kotlin), ✅ experience in BigQuery data warehousing solution, ✅ in-depth understanding of big data aspects like data storage, modeling , processing , scheduling etc., ✅ understanding of Apache Spark (or similar distributed data processing framework), ✅ data modeling and data storage experience, ✅ ensuring solution quality through automatic tests, CI/CD and code review, ✅ proven collaboration with businesses, ✅ English proficiency at min. B2 level, proficient in Polish. Nice to have: 🌟 knowledge of dbt, Docker and Kubernetes, Apache Kafka, 🌟 familiarity with Apache Airflow or similar pipeline orchestrator, 🌟 another JVM (Java/Scala/Kotlin) programming language, 🌟 experience in Machine Learning projects, 🌟 familiarity with one of BI tools: Power BI/Looker/Tableau, 🌟 willingness to share knowledge (conferences, articles, open-source projects). What’s on offer: 🔥 100% remote work, with workation opportunity, 🔥 20 free days, 🔥 onboarding with a dedicated mentor, 🔥 project switching possible after a certain period, 🔥 individual budget for training and conferences, 🔥 benefits: Medicover Private Medical Care , co-financing of the Medicover Sport card, 🔥 opportunity to learn English with a native speaker, 🔥 regular company trips and informal get-togethers. Development opportunities in Datumo: 🚀 participation in industry conferences, 🚀 establishing Datumo's online brand presence, 🚀 support in obtaining certifications (e.g. GCP, Azure, Snowflake), 🚀 involvement in internal initiatives, like building technological roadmaps, 🚀 training budget, 🚀 access to internal technological training repositories. Discover our exemplary project: 🔌 IoT data ingestion to cloud The project integrates data from edge devices into the cloud using Azure services. The platform supports data streaming via either the IoT Edge environment with Java or Python modules, or direct connection using Kafka protocol to Event Hubs. It also facilitates batch data transmission to ADLS. Data transformation from raw telemetry to structured tables is done through Spark jobs in Databricks or data connections and update policies in Azure Data Explorer. ☁️ Petabyte-scale data platform migration to Google Cloud The goal of the project is to improve scalability and performance of the data platform by transitioning over a thousand active pipelines to GCP. The main focus is on rearchitecting existing Spark applications to either Cloud Dataproc or Cloud BigQuery SQL, depending on the Client’s requirements and automate it using Cloud Composer. 📈 Data analytics platform for investing company The project centers on developing and overseeing a data platform for an asset management company focused on ESG investing. Databricks is the central component. The platform, built on Azure cloud, integrates various Azure services for diverse functionalities. The primary task involves implementing and extending complex ETL processes that enrich investment data, using Spark jobs in Scala. Integrations with external data providers, as well as solutions for improving data quality and optimizing cloud resources, have been implemented. 🛒 Realtime Consumer Data Platform The initiative involves constructing a consumer data platform (CDP) for a major Polish retail company. Datumo actively participates from the project’s start, contributing to planning the platform’s architecture. The CDP is built on Google Cloud Platform (GCP), utilizing services like Pub/Sub, Dataflow and BigQuery. Open-source tools, including a Kubernetes cluster with Apache Kafka, Apache Airflow and Apache Flink, are used to meet specific requirements. This combination offers significant possibilities for the platform. Recruitment process: 1️⃣Tech quiz - 15 minutes 2️⃣ Soft skills interview - 30 minutes 3️⃣ Technical interview - 60 minutes Find out more by visiting our website - https://www.datumo.io If you like what we do and you dream about creating this world with us - don’t wait, apply now!
Technology
Datumo
Data Engineer (GCP)
Mid
Remote
Warsaw, Poland
16,000 - 28,000 PLN
🏢 Summary: The offer is for a Data Engineer responsible for designing and developing scalable data platforms and pipelines in cloud environments, primarily on Google Cloud Platform. The role focuses on big data processing, data modeling, and building high-quality data solutions using modern distributed frameworks. The position involves working on advanced cloud-based analytics and data warehousing projects. 🗂️ Requirements: 3-4 years commercial programming experience, Experience with Google Cloud Platform, Strong knowledge of Python, Strong knowledge of SQL, Proficiency in JVM language (Scala or Java or Kotlin), Experience with BigQuery, Understanding of big data processing and storage, Experience with Apache Spark or similar framework, Experience in data modeling, Experience with CI/CD pipelines, Experience with automated testing, Experience with code review practices 📃 Skills: Python, SQL, Scala, Java, Kotlin, GCP, BigQuery, Spark, JVM, BigData, DataModeling, CICD, Testing, CodeReview 🏢 Description: We’re looking for a Data Engineer ready to push boundaries and grow with us. Datumo specializes in providing Data Engineering and Cloud Computing consulting services to clients from all over the world, primarily in Western Europe, Poland and the USA. Core industries we support include e-commerce 🛒, telecommunications 📡 and life sciences 🧬. Our team consists of exceptional people whose commitment allows us to conduct highly demanding projects. Our team members tend to stick around for more than 3 years, and when a project wraps up, we don't let them go - we embark on a journey to discover exciting new challenges for them. It's not just a workplace; it's a community that grows together! Must-have: ✅ at least 3-4 years of commercial experience in programming, ✅ proven record with a cloud provider - Google Cloud Platform, ✅ strong knowledge of Python, SQL and JVM languages (Scala or Java or Kotlin), ✅ experience in BigQuery data warehousing solution, ✅ in-depth understanding of big data aspects like data storage, modeling , processing , scheduling etc., ✅ understanding of Apache Spark (or similar distributed data processing framework), ✅ data modeling and data storage experience, ✅ ensuring solution quality through automatic tests, CI/CD and code review, ✅ proven collaboration with businesses, ✅ English proficiency at min. B2 level, proficient in Polish. Nice to have: 🌟 knowledge of dbt, Docker and Kubernetes, Apache Kafka, 🌟 familiarity with Apache Airflow or similar pipeline orchestrator, 🌟 another JVM (Java/Scala/Kotlin) programming language, 🌟 experience in Machine Learning projects, 🌟 familiarity with one of BI tools: Power BI/Looker/Tableau, 🌟 willingness to share knowledge (conferences, articles, open-source projects). What’s on offer: 🔥 100% remote work with workation opportunity (you need to be based in Poland), 🔥 20 free days, 🔥 onboarding with a dedicated mentor, 🔥 project switching possible after a certain period, 🔥 individual budget for training and conferences, 🔥 benefits: Medicover Private Medical Care , co-financing of the Medicover Sport card, 🔥 opportunity to learn English with a native speaker, 🔥 regular company trips and informal get-togethers. Development opportunities in Datumo: 🚀 participation in industry conferences, 🚀 establishing Datumo's online brand presence, 🚀 support in obtaining certifications (e.g. GCP, Azure, Snowflake), 🚀 involvement in internal initiatives, like building technological roadmaps, 🚀 training budget, 🚀 access to internal technological training repositories. Discover our exemplary project: 🔌 IoT data ingestion to cloud The project integrates data from edge devices into the cloud using Azure services. The platform supports data streaming via either the IoT Edge environment with Java or Python modules, or direct connection using Kafka protocol to Event Hubs. It also facilitates batch data transmission to ADLS. Data transformation from raw telemetry to structured tables is done through Spark jobs in Databricks or data connections and update policies in Azure Data Explorer. ☁️ Petabyte-scale data platform migration to Google Cloud The goal of the project is to improve scalability and performance of the data platform by transitioning over a thousand active pipelines to GCP. The main focus is on rearchitecting existing Spark applications to either Cloud Dataproc or Cloud BigQuery SQL, depending on the Client’s requirements and automate it using Cloud Composer. 📈 Data analytics platform for investing company The project centers on developing and overseeing a data platform for an asset management company focused on ESG investing. Databricks is the central component. The platform, built on Azure cloud, integrates various Azure services for diverse functionalities. The primary task involves implementing and extending complex ETL processes that enrich investment data, using Spark jobs in Scala. Integrations with external data providers, as well as solutions for improving data quality and optimizing cloud resources, have been implemented. 🛒 Realtime Consumer Data Platform The initiative involves constructing a consumer data platform (CDP) for a major Polish retail company. Datumo actively participates from the project’s start, contributing to planning the platform’s architecture. The CDP is built on Google Cloud Platform (GCP), utilizing services like Pub/Sub, Dataflow and BigQuery. Open-source tools, including a Kubernetes cluster with Apache Kafka, Apache Airflow and Apache Flink, are used to meet specific requirements. This combination offers significant possibilities for the platform. Recruitment process: 1️⃣Tech quiz - 15 minutes 2️⃣ Soft skills interview - 30 minutes 3️⃣ Technical interview - 60 minutes Find out more by visiting our website - https://www.datumo.io If you like what we do and you dream about creating this world with us - don’t wait, apply now!
Technology
Datumo
Data Engineer
Mid
Remote
Warsaw, Poland
16,000 - 28,000 PLN
🏢 Summary: The offer is for a Data Engineer role focused on designing, building, and optimizing scalable data platforms and pipelines in cloud environments. The position involves working with big data technologies, data warehousing solutions, and distributed processing frameworks across international projects. The role emphasizes high-quality delivery through automated testing, CI/CD, and close collaboration with business stakeholders. 🗂️ Requirements: Minimum 3 years of commercial programming experience, Experience with at least one cloud provider: GCP, Azure or AWS, Knowledge of JVM language (Scala or Java or Kotlin), Knowledge of Python, Knowledge of SQL, Experience with data warehousing solution: BigQuery or Snowflake or Databricks or similar, Understanding of big data concepts: storage, modeling, processing, scheduling, Experience in data modeling, Experience in data storage solutions, Experience with automated testing and CI/CD, Experience with code review practices, Experience in collaboration with business stakeholders, English proficiency at B2 level, Communicative Polish 📃 Skills: GCP, Azure, AWS, Scala, Java, Kotlin, Python, SQL, BigQuery, Snowflake, Databricks, Spark, Dataproc, CloudComposer, PubSub, Dataflow, Kafka, Airflow, Flink, Kubernetes, Docker, BigData, CI/CD, ETL 🏢 Description: We’re looking for a Data Engineer ready to push boundaries and grow with us. Datumo specializes in providing Data Engineering and Cloud Computing consulting services to clients from all over the world, primarily in Western Europe, Poland and the USA. Core industries we support include e-commerce 🛒, telecommunications 📡 and life sciences 🧬. Our team consists of exceptional people whose commitment allows us to conduct highly demanding projects. Our team members tend to stick around for more than 3 years, and when a project wraps up, we don't let them go - we embark on a journey to discover exciting new challenges for them. It's not just a workplace; it's a community that grows together! Must-have: ✅ at least 3 years of commercial experience in programming ✅ proven record with a selected cloud provider GCP (preferred), Azure or AWS ✅ good knowledge of JVM languages (Scala or Java or Kotlin), Python, SQL ✅ experience in one of data warehousing solutions: BigQuery/Snowflake/Databricks or similar ✅ in-depth understanding of big data aspects like data storage, modeling , processing , scheduling etc. ✅ data modeling and data storage experience ✅ ensuring solution quality through automatic tests, CI/CD and code review ✅ proven collaboration with businesses ✅ English proficiency at B2 level, communicative in Polish Nice to have: 🌟 knowledge of dbt, Docker and Kubernetes, Apache Kafka 🌟 familiarity with Apache Airflow or similar pipeline orchestrator 🌟 another JVM (Java/Scala/Kotlin) programming language 🌟 experience in Machine Learning projects 🌟 understanding of Apache Spark or similar distributed data processing framework 🌟 familiarity with one of BI tools: Power BI/Looker/Tableau 🌟 willingness to share knowledge (conferences, articles, open-source projects) What’s on offer: 🔥 100% remote work, with workation opportunity 🔥 20 free days 🔥 onboarding with a dedicated mentor 🔥 project switching possible after a certain period 🔥 individual budget for training and conferences 🔥 benefits: Medicover Private Medical Care , co-financing of the Medicover Sport card 🔥 opportunity to learn English with a native speaker 🔥 regular company trips and informal get-togethers Development opportunities in Datumo: 🚀 participation in industry conferences 🚀 establishing Datumo's online brand presence 🚀 support in obtaining certifications (e.g. GCP, Azure, Snowflake) 🚀 involvement in internal initiatives, like building technological roadmaps 🚀 training budget 🚀 access to internal technological training repositories Discover our exemplary project: 🔌 IoT data ingestion to cloud The project integrates data from edge devices into the cloud using Azure services. The platform supports data streaming via either the IoT Edge environment with Java or Python modules, or direct connection using Kafka protocol to Event Hubs. It also facilitates batch data transmission to ADLS. Data transformation from raw telemetry to structured tables is done through Spark jobs in Databricks or data connections and update policies in Azure Data Explorer. ☁️ Petabyte-scale data platform migration to Google Cloud The goal of the project is to improve scalability and performance of the data platform by transitioning over a thousand active pipelines to GCP. The main focus is on rearchitecting existing Spark applications to either Cloud Dataproc or Cloud BigQuery SQL, depending on the Client’s requirements and automate it using Cloud Composer. 📈 Data analytics platform for investing company The project centers on developing and overseeing a data platform for an asset management company focused on ESG investing. Databricks is the central component. The platform, built on Azure cloud, integrates various Azure services for diverse functionalities. The primary task involves implementing and extending complex ETL processes that enrich investment data, using Spark jobs in Scala. Integrations with external data providers, as well as solutions for improving data quality and optimizing cloud resources, have been implemented. 🛒 Realtime Consumer Data Platform The initiative involves constructing a consumer data platform (CDP) for a major Polish retail company. Datumo actively participates from the project’s start, contributing to planning the platform’s architecture. The CDP is built on Google Cloud Platform (GCP), utilizing services like Pub/Sub, Dataflow and BigQuery. Open-source tools, including a Kubernetes cluster with Apache Kafka, Apache Airflow and Apache Flink, are used to meet specific requirements. This combination offers significant possibilities for the platform. Recruitment process: 1️⃣Quiz - 15 minutes 2️⃣ Soft skills interview - 30 minutes 3️⃣ Technical interview - 60 minutes Find out more by visiting our website - https://www.datumo.io If you like what we do and you dream about creating this world with us - don’t wait, apply now!
Technology
TechTree
Senior Data Platform Engineer
Senior
Remote
Krakow, Poland
208,000 - 312,000 PLN/yr
🏢 Summary: Senior Data Platform Engineer role focused on building and optimising a cloud-native lakehouse platform for large-scale analytics and reporting. The position involves designing distributed data pipelines, enabling self-service analytics, and implementing governance and observability frameworks using modern data technologies. You will work with Spark-based systems and integrated data warehousing solutions to deliver scalable, reliable data platforms. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production environments, Experience with Delta Lake and/or Apache Iceberg in production, Practical experience with dbt for data transformations, Experience with Databricks and Snowflake, Understanding of data governance and lineage in large-scale environments, Familiarity with Kubernetes and Docker, Experience with CI/CD and automated testing practices, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platforms you build will have real consequence across the organisation. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As a Senior Data Platform Engineer, you'll combine strong software engineering principles with deep data expertise to build robust, cloud-native platforms that process large-scale datasets efficiently and enable internal teams to build reporting and analytics on top of them. The role emphasises cloud-native architecture, lakehouse integration, data warehousing, and governance best practices. You'll work on systems using Apache Spark, Delta Lake, and Iceberg, and help deliver curated data models and self-service analytics capabilities to internal stakeholders. You'll also participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems Design and implement scalable data pipelines and distributed systems using Spark and Python to process and transform large-scale datasets for analytics and reporting. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Iceberg, ensuring data reliability, versioning, and performance optimisation at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Collaborate with internal teams to deliver curated datasets and self-service analytics capabilities for reporting and advanced use cases. Data warehousing optimisation Integrate and optimise Databricks and Snowflake for scalable storage and query performance. Drive performance tuning and cost optimisation across Spark jobs and cloud-native environments. Governance and observability Implement observability and governance frameworks including data lineage, quality checks, and compliance controls. Build platforms that allow secure and compliant access to diverse data sources. Engineering best practices Apply and champion clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in both Python and SQL, applied to production data platform work at scale. Apache Spark Solid hands-on experience with Spark for distributed data processing, including performance tuning in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You've applied these in production and understand the trade-offs in real-world scenarios. dbt and analytics tooling Practical experience with dbt for transformation workflows. Familiarity with Databricks and Snowflake for large-scale analytics workloads. Data governance and compliance Understanding of data governance, lineage tracking, and compliance requirements in large-scale, multi-tenant data environments. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 208,000 – 312,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology
SoftBlue
Data Engineer (Python & AWS)
Senior
Remote
Bydgoszcz, Poland
150 - 180 PLN
🏢 Summary: Senior Data Engineer role focused on designing and building scalable serverless data ingestion pipelines in AWS within the healthcare domain. The position emphasizes strong Python engineering, cloud architecture leadership, and implementation of modern data platforms and DevOps practices. The role involves driving technical excellence and delivering reliable, high-impact data solutions in an international environment. 🗂️ Requirements: 10+ years of experience in Python programming, Strong software engineering skills in data processing, Extensive experience with AWS Cloud and Serverless Architecture, Hands-on experience with AWS Lambda, S3, and Cognito, Experience building E2E automated tests for data pipelines, Practical knowledge of Data Mesh and Medallion Architecture, Experience with Infrastructure as Code using AWS CDK or Terraform, Experience with CI/CD pipelines using GitLab or GitHub Actions, Experience with ETL/ELT processes and dbt, Experience working with GraphQL, Minimum B2 level English proficiency 📃 Skills: Python, AWS, Lambda, S3, Cognito, Boto3, DataMesh, Medallion, CDK, Terraform, GitLab, GitHubActions, ETL, ELT, dbt, GraphQL, CI/CD 🏢 Description: We are looking for a highly skilled Data Engineer to join our client in the healthcare sector. Our requirements: Technical Expertise: Python Programming: 10+ years of experience with strong Software Engineering skills focused on data processing. AWS & Serverless: Extensive experience with AWS Cloud, specifically focusing on Serverless Architecture and services (including AWS Lambda , AWS S3 Tables , and AWS Cognito ). Automated Testing: Proven experience in developing End-to-End (E2E) automated tests to ensure pipeline reliability, utilizing tools such as Boto3 for AWS resource validation. Data Concepts: Practical knowledge of Data Mesh and Medallion Architecture , along with general data processing and analysis. DevOps & IaC: Hands-on experience with Infrastructure as Code ( AWS CDK or Terraform ) and CI/CD pipelines ( GitLab pipelines or GitHub Actions ). Modern Tooling: Experience with ETL/ELT solutions, dbt , and GraphQL . Communication & Soft Skills: English Language: Minimum B2 level , enabling smooth daily technical and business communication in a global environment. Collaboration: Excellent communication skills and the ability to thrive in a collaborative, international team. Standards: A strong commitment to high standards of ethics, quality (Clean Code), and reliable delivery. Nice to have: Experience with Snowflake and SQL . Knowledge of Data Vault 2.0 modeling. Experience with Databricks . Familiarity with the Microsoft ecosystem: C# / .Net, T-SQL, SQL Server , and Azure DevOps . Experience with Star Schema database modeling. Knowledge of Descriptive Statistics. Your responsibilites: Design and build scalable Data Ingestion pipelines within the AWS cloud ecosystem. Lead technical delivery and implementation of core platform components, ensuring architectural integrity across the entire data lifecycle. Collaborate with Engineering Managers and cross-functional teams across the globe and Poland. Drive technical excellence by improving team processes, architecture standards, and engineering best practices. Support and consult with stakeholders to ensure successful delivery of high-impact, data-driven solutions. Contribute to the growth and maturity of the team’s cloud and data engineering capabilities. We offer: Challenging role within the company that creates innovative solutions. Work in international environment on demanding projects. Remote work model. Subsidized private medical care, life insurance, multisport card. Integration meetings. Employee referral program. If you have a deep expertise in Python and AWS , and building scalable Serverless data architectures is where you truly excel, this is the perfect role for you!
Technology
Harvey Nash Technology
Senior Data Engineer (cloud&ai)
Senior
On-site
Warsaw, Poland
30,000 - 40,000 PLN
🏢 Summary: Design and scale high-throughput data pipelines on cloud platforms to support advanced analytics and AI-driven products. The role focuses on building distributed data architectures in AWS and Databricks, ensuring performance, governance, and data quality. You will collaborate with AI/ML teams to deliver scalable, production-grade data solutions. 🗂️ Requirements: 3+ years of data engineering experience, Strong Python programming skills, Experience with Spark or Scala, Experience building distributed data pipelines in cloud environments, Knowledge of data modeling and data warehousing principles, Bachelor’s or Master’s degree in Computer Science or Engineering 📃 Skills: Python, Spark, Scala, AWS, Glue, EMR, Fargate, StepFunctions, Databricks, SQL, APIs, GenAI, GraphDB 🏢 Description: Data Engineer – Cloud & AI Platforms We’re looking for a Data Engineer to design and scale high-throughput data pipelines supporting advanced analytics and AI-driven products. What You’ll Do Architect and maintain distributed data pipelines in Databricks and AWS (Glue, EMR, Fargate, Step Functions) Ingest and process large volumes of structured and unstructured data (internal, market, third-party, alternative sources) Collaborate with AI/ML and engineering teams to design scalable data architectures and APIs Optimize performance and cost using Spark and cloud-native best practices Implement data governance, privacy, lineage, and access controls Build automated validation, monitoring, and data quality frameworks Evaluate emerging GenAI and data tooling to enhance platform capabilities What You Bring 3+ years of experience in data engineering Strong Python and experience with Spark or Scala Proven experience building distributed pipelines in cloud environments Solid understanding of data modeling, architecture, and warehousing principles Innovative problem-solving mindset Bachelor’s or Master’s degree in Computer Science or Engineering Nice to have: Experience with graph databases.
Technology
TechTree
Advanced Data Platform Engineer
Senior
Remote
Krakow, MA, Poland
160,000 - 240,000 PLN/yr
🏢 Summary: The role focuses on designing and building scalable, cloud-native data platforms to enable advanced analytics and reporting across a large internal data ecosystem. It involves developing distributed data pipelines, lakehouse architectures, and optimised data warehousing solutions with strong emphasis on performance, governance, and reliability. The position requires deep technical expertise in big data technologies and cloud-native infrastructure, including on-call responsibility for platform stability. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Commercial experience with Apache Spark for distributed data processing, Hands-on experience with Delta Lake and/or Apache Iceberg in production, Experience with dbt for SQL transformations, Experience with Databricks and Snowflake, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Understanding of performance tuning and cost optimisation in large-scale data systems, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. Over 75% of our business has transitioned to our cloud platform, and we are making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you build will have real consequence. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As an Advanced Data Platform Engineer, you'll design and implement scalable, cloud-native data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services to support diverse analytical use cases at enterprise scale. The role emphasises technical depth — performance optimisation, governance best practices, and the kind of engineering rigour that keeps vast datasets accessible, secure, and compliant. You'll work closely with internal teams to deliver curated datasets and self-service analytics capabilities, and you'll participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems design Design and implement complex data pipelines and distributed systems using Spark and Python, applying clean code principles, modular design, CI/CD, automated testing, and thorough code reviews. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Apache Iceberg, ensuring reliability, performance, and long-term maintainability at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Deliver curated datasets and self-service analytics capabilities that empower internal stakeholders to explore data independently. Data warehousing optimisation Optimise Databricks and Snowflake environments for performance and scalability. Drive cost optimisation and performance tuning across Spark jobs and cloud-native infrastructure. Observability and governance Implement observability and governance frameworks including data lineage tracking and compliance controls, ensuring data remains secure and auditable. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in Python and SQL — the foundation for everything you'll build here. Apache Spark Solid experience with Spark for distributed data processing at scale, including performance tuning and optimisation. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the tradeoffs and have used these in production environments. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and SQL transformation pipelines. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Scalability and cost optimisation Understanding of performance tuning, scalability strategies, and cost optimisation for large-scale data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 160,000 – 240,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology
MOTIFE
Senior Data Platform Engineer
Senior
Hybrid
Warsaw, Poland
23,000 - 30,000 PLN/mo
🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, platform reliability, and cloud-native data infrastructure to improve performance, scalability, and observability of global data ecosystems. Hybrid work model with competitive salary and comprehensive benefits. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend engineering skills (Python, Java, or Kotlin), Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience in cloud environments (AWS preferred), Experience with relational or NoSQL databases, Hands-on experience with large-scale data pipelines, Experience with event-driven or streaming architectures, Understanding of end-to-end data flow (ingestion, transformation, storage, access), Experience designing scalable and reliable data systems 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .
Technology
MOTIFE
Senior Data Platform Engineer
Senior
Hybrid
Warsaw, Poland
23,000 - 30,000 PLN/mo
🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, cloud-native infrastructure, and data platform reliability to improve performance and scalability of global data systems. It is a hands-on role working with large-scale data pipelines, streaming, and cloud environments. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend programming skills in Python, Java, or Kotlin, Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience with cloud environments, preferably AWS, Experience with relational or NoSQL databases, Hands-on experience with large-scale data pipelines, Experience with event-driven architectures or streaming systems, Understanding of end-to-end data flow and data modeling, Experience designing scalable and reliable data systems, Experience with infrastructure automation tools 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .
Technology
MOTIFE
Senior Data Platform Engineer
Senior
Hybrid
Warsaw, Poland
23,000 - 30,000 PLN/mo
🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, cloud-native infrastructure, and data platform reliability to support global, data-intensive applications. The engineer will enhance performance, scalability, and observability of production datastores across a multi-region environment. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend programming skills (Python, Java, or Kotlin), Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience with cloud environments, preferably AWS, Hands-on experience with relational or NoSQL databases, Experience with large-scale data pipelines and data processing systems, Experience with event-driven architectures or streaming/batch processing, Understanding of end-to-end data flow (ingestion, transformation, storage, access), Experience designing scalable, high-performance, reliable data systems 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .