April 8, 2026

Senior Data Platform Engineer

Senior • Remote

208,000 - 312,000 PLN/yr

Krakow, Poland

ABOUT THE COMPANY

We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response.

We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platforms you build will have real consequence across the organisation.

ABOUT THE ROLE

We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As a Senior Data Platform Engineer, you'll combine strong software engineering principles with deep data expertise to build robust, cloud-native platforms that process large-scale datasets efficiently and enable internal teams to build reporting and analytics on top of them.

The role emphasises cloud-native architecture, lakehouse integration, data warehousing, and governance best practices. You'll work on systems using Apache Spark, Delta Lake, and Iceberg, and help deliver curated data models and self-service analytics capabilities to internal stakeholders. You'll also participate in on-call rotations as part of shared team responsibility.

WHAT YOU'LL WORK ON

Data pipeline and distributed systems

Design and implement scalable data pipelines and distributed systems using Spark and Python to process and transform large-scale datasets for analytics and reporting.

Lakehouse platform development

Develop and maintain lakehouse capabilities with Delta Lake and Iceberg, ensuring data reliability, versioning, and performance optimisation at scale.

Analytics workflow enablement

Integrate dbt for SQL transformations running on Spark. Collaborate with internal teams to deliver curated datasets and self-service analytics capabilities for reporting and advanced use cases.

Data warehousing optimisation

Integrate and optimise Databricks and Snowflake for scalable storage and query performance. Drive performance tuning and cost optimisation across Spark jobs and cloud-native environments.

Governance and observability

Implement observability and governance frameworks including data lineage, quality checks, and compliance controls. Build platforms that allow secure and compliant access to diverse data sources.

Engineering best practices

Apply and champion clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work.

On-call participation

Participate in on-call rotations as part of shared team responsibility for platform reliability.

WHAT WE LOOK FOR

Python and SQL

Strong programming skills in both Python and SQL, applied to production data platform work at scale.

Apache Spark

Solid hands-on experience with Spark for distributed data processing, including performance tuning in production environments.

Lakehouse architecture

Expertise in Delta Lake and/or Apache Iceberg. You've applied these in production and understand the trade-offs in real-world scenarios.

dbt and analytics tooling

Practical experience with dbt for transformation workflows. Familiarity with Databricks and Snowflake for large-scale analytics workloads.

Data governance and compliance

Understanding of data governance, lineage tracking, and compliance requirements in large-scale, multi-tenant data environments.

Infrastructure and containerisation

Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments.

Software engineering fundamentals

Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems.

Bonus

Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust.

THE TEAM

You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing.

COMPENSATION & BENEFITS

Salary

208,000 – 312,000 PLN per year, plus an annual performance bonus and long-term incentives.

Health coverage

Comprehensive health, dental, and vision plans.

Parental leave

Parental leave available for both primary and secondary caregivers.

Flexible working

Flexible work arrangements with a remote-first model.

Company breaks

Two week-long company-wide breaks per year, plus additional time off.

Training investment

Dedicated training investment programme to support ongoing professional development.

Similar jobs you might like

Technology

TechTree

Advanced Data Platform Engineer

Senior

Remote

Krakow, MA, Poland

160,000 - 240,000 PLN/yr

🏢 Summary: The role focuses on designing and building scalable, cloud-native data platforms to enable advanced analytics and reporting across a large internal data ecosystem. It involves developing distributed data pipelines, lakehouse architectures, and optimised data warehousing solutions with strong emphasis on performance, governance, and reliability. The position requires deep technical expertise in big data technologies and cloud-native infrastructure, including on-call responsibility for platform stability. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Commercial experience with Apache Spark for distributed data processing, Hands-on experience with Delta Lake and/or Apache Iceberg in production, Experience with dbt for SQL transformations, Experience with Databricks and Snowflake, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Understanding of performance tuning and cost optimisation in large-scale data systems, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. Over 75% of our business has transitioned to our cloud platform, and we are making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you build will have real consequence. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As an Advanced Data Platform Engineer, you'll design and implement scalable, cloud-native data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services to support diverse analytical use cases at enterprise scale. The role emphasises technical depth — performance optimisation, governance best practices, and the kind of engineering rigour that keeps vast datasets accessible, secure, and compliant. You'll work closely with internal teams to deliver curated datasets and self-service analytics capabilities, and you'll participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems design Design and implement complex data pipelines and distributed systems using Spark and Python, applying clean code principles, modular design, CI/CD, automated testing, and thorough code reviews. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Apache Iceberg, ensuring reliability, performance, and long-term maintainability at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Deliver curated datasets and self-service analytics capabilities that empower internal stakeholders to explore data independently. Data warehousing optimisation Optimise Databricks and Snowflake environments for performance and scalability. Drive cost optimisation and performance tuning across Spark jobs and cloud-native infrastructure. Observability and governance Implement observability and governance frameworks including data lineage tracking and compliance controls, ensuring data remains secure and auditable. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in Python and SQL — the foundation for everything you'll build here. Apache Spark Solid experience with Spark for distributed data processing at scale, including performance tuning and optimisation. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the tradeoffs and have used these in production environments. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and SQL transformation pipelines. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Scalability and cost optimisation Understanding of performance tuning, scalability strategies, and cost optimisation for large-scale data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 160,000 – 240,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

TechTree

Lead Distributed Data Platform Engineer

Senior

Remote

Warsaw, Poland

270,000 - 406,000 PLN/yr

🏢 Summary: Lead Distributed Data Platform Engineer responsible for architecting and delivering enterprise-scale lakehouse and distributed data platforms to enable advanced analytics and reporting. The role combines hands-on technical leadership with team mentorship, driving scalable, secure, and high-performance data solutions in cloud-native environments. You will guide architectural decisions, enforce engineering best practices, and ensure platform reliability at scale. 🗂️ Requirements: Proven experience leading data engineering or platform teams, Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production, Experience with Delta Lake and/or Apache Iceberg in production, Experience designing distributed systems and lakehouse architectures, Experience building scalable data pipelines, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Experience working in cloud-native environments 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD, Java, Scala, Rust 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platform you lead will underpin how the entire organisation accesses and acts on its data. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As Lead Distributed Data Platform Engineer, you'll combine deep technical expertise with hands-on team leadership — guiding a team in designing and maintaining data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services at enterprise scale. You'll lead architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable solutions. The role emphasises technical leadership, governance best practices, and a culture of innovation and continuous improvement. You'll also participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data platform engineers, promoting collaboration, knowledge sharing, and professional growth. Set and maintain high engineering standards across the team. Distributed systems architecture Drive architectural decisions for distributed systems and lakehouse platforms using Spark, Delta Lake, and Iceberg. Facilitate architecture reviews and contribute to design decisions for fault-tolerant, future-ready systems. Data pipeline and platform delivery Oversee design and implementation of scalable data pipelines and analytics workflows, ensuring they are reliable, performant, and maintainable at scale. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all platform engineering work. Performance and cost optimisation Manage performance tuning, scalability strategies, and cost optimisation across cloud-native environments and large-scale distributed workloads. Governance and observability Champion governance, observability, and compliance frameworks across all data platforms — ensuring data remains accessible, secure, and auditable. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure delivery aligns with business objectives and analytics needs. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or platform development teams — mentoring engineers, owning architectural decisions, and driving delivery outcomes. Python and SQL Strong programming skills in both Python and SQL applied to production data platform work at scale. Apache Spark Hands-on experience with Spark for distributed data processing, including performance tuning and optimisation in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the trade-offs and have applied these technologies in production at scale. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and large-scale data processing. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data platform systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Communication and stakeholder management Strong communication skills with the confidence to operate across engineering teams, cross-functional partners, and senior leadership. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. Exposure to service mesh and advanced orchestration patterns. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

TechTree

Lead Data Engineer

Senior

Remote

Krakow, Poland

270,000 - 406,000 PLN/yr

🏢 Summary: Lead Data Engineer role focused on driving architecture and leading a team to build scalable, secure ETL/ELT pipelines and analytics-ready data models on modern cloud platforms. The position combines hands-on engineering with technical leadership, ensuring high standards in governance, observability, and performance optimisation. You will shape data infrastructure that supports large-scale analytics across the organisation. 🗂️ Requirements: Proven experience leading data engineering or analytics engineering teams, Strong programming skills in SQL, Strong programming skills in Python, Hands-on experience with Airflow or Prefect in production, Deep practical experience with dbt, Experience with Snowflake or Databricks at scale, Strong knowledge of dimensional modelling and SCD strategies, Experience implementing data quality and governance frameworks, Experience with CI/CD and automated testing in data systems 📃 Skills: SQL, Python, Airflow, Prefect, dbt, Snowflake, Databricks, ETL, ELT, CI/CD, SCD, Dimensional, Git, IaC 🏢 Description: ABOUT THE COMPANY Our client is a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you lead will have real consequence across the organisation. ABOUT THE ROLE We're looking for a Lead Data Engineer to combine deep technical expertise with hands-on team leadership, guiding a team of data engineers building and maintaining ETL/ELT pipelines, data models, and governance frameworks that power analytics and reporting across the organisation. This is a technical leadership role — you'll drive architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable data solutions. You'll collaborate closely with stakeholders to align technical work with business objectives, champion governance and observability standards, and foster a culture of continuous improvement. The expectation is that you're equally effective in an architecture review as you are pairing with an engineer on a tricky pipeline problem. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data engineers, promoting collaboration, knowledge sharing, and professional growth. Set the standard for engineering quality and hold the bar consistently. Architecture and pipeline design Drive architectural decisions for ETL/ELT pipelines, orchestration frameworks (Airflow/Prefect), and transformation layers (dbt). Facilitate architecture reviews and contribute to design decisions for scalable, fault-tolerant systems. Analytics data modelling Oversee design and implementation of analytics-ready data models — dimensional schemas, SCD strategies, and semantic layers — that internal teams can build on reliably. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. Platform optimisation Manage performance tuning and cost optimisation for Snowflake, Databricks, and related cloud data platforms at scale. Governance and observability Champion governance, observability, and compliance frameworks across all data workflows — including data quality, lineage tracking, and multi-tenant environment controls. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure timely delivery aligned with business objectives. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or analytics-focused development teams — mentoring engineers, driving architectural decisions, and owning delivery outcomes. SQL and Python Strong programming skills in both SQL and Python, applied to production data systems at scale. ETL/ELT orchestration Hands-on experience with orchestration tools — Airflow and/or Prefect — in production pipeline environments. dbt expertise Deep practical experience with dbt for transformation workflows and analytics modelling, including testing, documentation, and modular project design. Snowflake and Databricks Familiarity with Snowflake and/or Databricks for large-scale data processing, including performance tuning and cost management. Data modelling principles Solid understanding of data modelling principles, incremental strategies, and schema design for analytics — dimensional modelling, SCDs, and semantic layer design. Governance and data quality Knowledge of data quality frameworks, lineage tracking, and governance in multi-tenant environments. Software engineering practices Familiarity with CI/CD, automated testing, and infrastructure-as-code practices applied to data systems. Communication and stakeholder management Strong communication skills with the ability to operate confidently across technical teams and business stakeholders. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements, hybrid model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

MOTIFE

Software Engineer (Data)

Senior

Hybrid

Krakow, Poland

20,000 - 23,000 PLN/mo

🏢 Summary: The offer is for a Software Engineer (Data) role focused on building scalable data infrastructure and production-ready ML systems within an AI-driven environment. The position combines backend engineering, data engineering, and MLOps to design high-throughput pipelines and support AI solutions from experimentation to production. The role involves close collaboration with Data Science teams to shape architecture and engineering standards for next-generation data platforms. 🗂️ Requirements: 4+ years of software engineering experience, Strong proficiency in Python, Experience building scalable production systems, Experience with data-intensive applications and SQL databases, Knowledge of data modeling and query optimization, Experience with Terraform, Kubernetes, Docker or similar tools, Understanding of CI/CD pipelines and Infrastructure as Code, Experience implementing automated testing and clean architecture principles, Ability to productionize ML solutions with Data Science teams 📃 Skills: Python, MySQL, PostgreSQL, Spark, Terraform, Kubernetes, Docker, Airflow, SQL, CI/CD, AWS, Snowflake, DBT, MLOps 🏢 Description: Our client helps small teams power big businesses with the must-have platform for intelligent marketing automation. Customers from over 170 countries depend on the Client’s mix of pre-built automation and integration to power personalized marketing, transactional emails, and one-to-one CRM interactions throughout the customer lifecycle. We’re looking for a Software Engineer (Data) to join our client’s growing AI and Data organization and help build the scalable foundations behind next-generation data and ML systems. In this role, you won’t just be working with data infrastructure; you’ll be shaping the software architecture, engineering standards, and production-ready platforms that power the company’s AI ecosystem. This is more than a traditional backend or data engineering position. You’ll work at the intersection of software engineering, data, and AI, partnering closely with Data Science and AI teams to bridge the gap between experimentation and production. From designing resilient systems to building scalable MLOps pipelines, your work will directly influence how AI solutions are developed, deployed, and scaled across the organization. Key takeaways: Stack: Python, MySQL, PostgreSQL, Spark, Terraform, Kubernetes, Docker Salary : 20 000 - 23 000 PLN gross/month, Contract of employment (+10% annual bonus, 75% Creative Tax) Working model: Hybrid, once a week in the office Location: Krakow, ul. Konopnickiej Recruitment process: Call with MOTIFE recruiter (30 min) Interview with Hiring Manager (45 min) Technical interview, live coding (1h) Cross-functional interview (1h) Responsibilities: Design, develop, and maintain scalable high-throughput data pipelines across complex data ecosystems. Build and optimize data models, schemas, and database structures to ensure long-term scalability and performance. Implement engineering best practices, including automated testing, CI/CD pipelines, and Infrastructure as Code (Terraform). Partner closely with AI and Data Science teams to productionize machine learning solutions and develop scalable MLOps pipelines. Engineer and maintain feature stores and data infrastructure supporting AI-driven initiatives. Monitor, maintain, and improve the reliability and efficiency of containerized environments using Kubernetes, Docker, and Airflow. Ensure platform stability, observability, and operational excellence through proactive system monitoring and health checks. Collaborate cross-functionally with engineering and business stakeholders to translate complex technical concepts into actionable insights. Contribute to the architectural direction and scalability of the organization’s AI and data platforms. Drive the adoption of robust software engineering standards across data and infrastructure projects. Requirements: 4+ years of experience in software engineering, with strong hands-on experience in backend development and building scalable production systems. Strong proficiency in Python and solid software engineering fundamentals, including clean architecture, testing, and maintainable code practices. Experience working with data-intensive applications, databases, and SQL, including data modeling and query optimization. Exposure to modern data engineering, cloud, or infrastructure environments, with familiarity in tools such as Terraform, Kubernetes, Docker, or similar technologies. Understanding of CI/CD pipelines, Infrastructure as Code, and general engineering best practices. Interest in AI/ML ecosystems and willingness to work closely with Data Science and AI teams on productionizing ML solutions. Familiarity with cloud platforms and modern data stack technologies such as AWS, Snowflake, Spark, or DBT is considered a strong plus. Ownership mindset and comfort working in evolving, fast-moving environments where systems and processes are still being built. What we offer: Health Benefits 1. Medical Full coverage for employees and their dependents through LUX MED. Employees have access to the “Premium” package, providing enhanced coverage and greater access to care. A client pays 100% of the premium for employees and 50% for dependents. 2. Dental No additional cost for dental coverage- integrated into LUX MED medical plan. 3. Vision Reimbursement for vision expenses up to 400 PLN every 2 years. Mental Health Tools Access to TELUS Health EAP to provide support and resources in a time of need. Additional Benefits 10% annual bonus 75% Creative Tax Vacation: 26 days. Home Office Stipend: One-time $150 equivalent home office stipend to outfit their home office. Calm Subscription: Premium subscription access to Calm, the #1 app for sleep, meditation, and relaxation. Hub Perks: Receive meal and transportation benefits when traveling to the Poland Hub. Baby Swag: If you have a baby or adopt, you’ll receive a company-branded first bath bundle. Sabbatical Program: After 5 years of employment, receive a month-long paid sabbatical leave, with a sabbatical leave bonus.

Technology

MOTIFE

Senior Data Platform Engineer

Senior

Hybrid

Warsaw, Poland

23,000 - 30,000 PLN/mo

🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, platform reliability, and cloud-native data infrastructure to improve performance, scalability, and observability of global data ecosystems. Hybrid work model with competitive salary and comprehensive benefits. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend engineering skills (Python, Java, or Kotlin), Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience in cloud environments (AWS preferred), Experience with relational or NoSQL databases, Hands-on experience with large-scale data pipelines, Experience with event-driven or streaming architectures, Understanding of end-to-end data flow (ingestion, transformation, storage, access), Experience designing scalable and reliable data systems 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .

Technology

MOTIFE

Senior Data Platform Engineer

Senior

Hybrid

Warsaw, Poland

23,000 - 30,000 PLN/mo

🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, cloud-native infrastructure, and data platform reliability to improve performance and scalability of global data systems. It is a hands-on role working with large-scale data pipelines, streaming, and cloud environments. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend programming skills in Python, Java, or Kotlin, Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience with cloud environments, preferably AWS, Experience with relational or NoSQL databases, Hands-on experience with large-scale data pipelines, Experience with event-driven architectures or streaming systems, Understanding of end-to-end data flow and data modeling, Experience designing scalable and reliable data systems, Experience with infrastructure automation tools 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .

Technology

MOTIFE

Senior Data Platform Engineer

Senior

Hybrid

Warsaw, Poland

23,000 - 30,000 PLN/mo

🏢 Summary: Senior Data Platform Engineer role focused on designing, building, and operating scalable, highly available data persistence systems for distributed services. The position combines backend engineering, cloud-native infrastructure, and data platform reliability to support global, data-intensive applications. The engineer will enhance performance, scalability, and observability of production datastores across a multi-region environment. 🗂️ Requirements: 5+ years of software engineering experience in production systems, Strong backend programming skills (Python, Java, or Kotlin), Experience with large-scale, data-intensive systems, Solid understanding of distributed systems fundamentals, Experience with cloud environments, preferably AWS, Hands-on experience with relational or NoSQL databases, Experience with large-scale data pipelines and data processing systems, Experience with event-driven architectures or streaming/batch processing, Understanding of end-to-end data flow (ingestion, transformation, storage, access), Experience designing scalable, high-performance, reliable data systems 📃 Skills: Python, Java, Kotlin, Kafka, Spark, PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch, AWS, Terraform, ETL 🏢 Description: We are hiring on behalf of our client, a global innovator in fitness and wellness technology. Their mission is to empower people to live fit, strong, long, and happy lives by delivering integrated experiences to millions of members anytime, anywhere. We are looking for a Senior Data Platform Engineer to join the Datastores team. This team is responsible for building and operating the core data persistence layer used by application services across the organization. In this role, you will design and improve the systems that store, access, and scale critical data across distributed services. It is a hands-on engineering position where you will work at the intersection of backend engineering, platform reliability, and cloud-native data infrastructure. Your work will directly influence the scalability, performance, and reliability of the company’s global data ecosystem. Key takeaways: Stack: Python, Kafka, Spark, PostgreSQL, AWS Salary: 23.000 - 30 000 PLN gross per month on Employment Contract Working model: hybrid - 3x weekly from the office Location: ul. Grzybowska 60, Warsaw Recruitment process: A call with Motife recruiter (30 min) Coding Interview (1h) Interview panel: architecture & system design discussion; Hiring Manager meeting (up to 2h in total) Responsibilities: Data Infrastructure Engineering Design, build, and operate backend systems that rely on scalable and highly available data persistence layers. Contribute to architectural decisions around distributed data systems, multi-region persistence, and global scalability. Improve the reliability and performance of production datastores used by critical services. Data Performance & Optimization Partner with service teams to improve database schema design, query performance, and data modelling. Optimize data access patterns and indexing strategies for relational and NoSQL databases. Support teams in designing systems that scale efficiently under high load. Developer Experience & Platform Tooling Build and maintain self-service tooling that enables engineers to provision and manage databases and caching layers. Contribute to infrastructure automation using tools such as Terraform and internal developer platforms. Improve observability and operational insight into datastore performance and reliability. Platform Reliability & Observability Implement monitoring, metrics, and tracing strategies to improve visibility into production data systems. Develop autoscaling and performance optimization strategies for critical data infrastructure. Support operational excellence by reducing manual processes and improving system resilience. Requirements: Technical Expertise 5+ years of experience in software engineering, building and operating production systems Strong backend engineering fundamentals (e.g. Python, Java, or Kotlin) Experience working with large-scale, data-intensive systems Solid understanding of distributed systems fundamentals (e.g. scalability, latency, reliability, data consistency) Experience working in cloud environments (preferably AWS) Familiarity with relational or NoSQL databases (e.g. PostgreSQL, MySQL, DynamoDB, Redis, Elasticsearch) Data Systems & Architecture Hands-on experience with large-scale data pipelines and data processing systems Exposure to event-driven architectures, streaming or batch processing (e.g. Kafka, Spark, ETL workflows) Understanding of end-to-end data flow: ingestion (how data enters the system) transformation (how it is processed) storage & access (how other services consume it) Experience designing systems where data performance, scalability, and reliability are critical Collaboration & Engineering Mindset Ability to work cross-functionally with service teams to improve system design and data access patterns. Strong problem-solving skills with a focus on performance, scalability, and reliability. Clear communication skills and a collaborative engineering approach. What we offer: 100% paid medical care Multisport Creative tax (KUP) Home office allowance MacBook Pro Apply now If you’re excited about building developer platforms that scale, empower teams, and set new standards for engineering excellence, we’d love to hear from you. Apply via our careers page and please submit your CV in English .

Technology

Nimble Robotics

Data Engineer II/III

Mid

On-site

San Francisco, CA

140,004 - 200,004 USD/yr

🏢 Summary: Data Engineer role focused on building and scaling data infrastructure for advanced robotics systems. The position involves designing reliable batch and real-time pipelines, optimizing ETL/ELT processes, and implementing data governance across cloud platforms. You will collaborate cross-functionally to deliver high-quality, scalable data solutions supporting company-wide analytics and operations. 🗂️ Requirements: BS/MS/PhD in Computer Science, Mathematics, Computer Engineering or related field, or equivalent practical experience, 1–5 years of professional experience, Proficiency in Python and SQL, Hands-on experience with Rust, Go, Java, C#, or C++, Experience with Kafka, Spark, and lakehouse formats (Icebert, DeltaLake), Experience with AWS, GCP, or Azure, Strong debugging and problem-solving skills, Willingness to work extended hours and weekends if needed, Ability to work full time onsite in San Francisco 📃 Skills: Python, SQL, Rust, Go, Java, C#, C++, Kafka, Spark, Icebert, DeltaLake, AWS, GCP, Azure, Clickhouse, Flink, Databricks 🏢 Description: About the Role Help us advance our robotics moonshot by scaling our data engineering efforts. Drive design and development of data infrastructure across our products and internal tools. You will play a critical role in working with a cross-functional team to architect and build advanced robotic systems. Responsibilities - Design, build, and maintain scalable, reliable data pipelines that support both batch and real-time analytics. - Develop and optimize ETL/ELT processes to ingest, transform, and integrate data from multiple sources into data warehouses and lakehouse platforms. - Drive and apply best practices in data modeling, pipeline architecture, query optimization, data quality, and data engineering standards. - Collaborate closely with finance, bizops, and engineering teams to understand data needs and deliver solutions. - Provide data engineering support to ensure data accessibility and usability company-wide. - Implement and maintain data governance frameworks to support compliance with internal policies and external regulatory requirements. - Evaluate and adopt modern data engineering technologies, tools, and best practices to improve scalability, reliability, and operational efficiency. - Document data engineering processes, designs, and architectures. Qualifications - BS/MS/PhD in Computer Science, Mathematics, Computer Engineering, or a related field, or equivalent practical experience. - 1-5 years of professional experience. - Proficiency in writing production-grade code in Python and SQL. - Hands-on experience with one of the following languages: Rust, Go, Java, C#, C++. - Strong debugging skills and the ability to diagnose and resolve issues efficiently. - Experience with Kafka, Spark and common lakehouse formats like Icebert, DeltaLake. - Experience working with cloud platforms like AWS, GCP, Azure. Preferred Experience - Experience with Clickhouse, Apache Flink, Databricks. - Experience on financial reporting and understanding of accountability. Additional Requirements - Willing to work extended hours and weekends if needed. - This position is based full time in our San Francisco headquarters. Compensation The pay range for this position at the start of employment is expected to be between $140,000 - $200,000/year. The exact offer may vary depending on job-related knowledge, skills, and experience. In addition to cash compensation, this position will also receive generous equity. Benefits - Unlimited Flexible Time Off. - Health insurance (medical, dental, vision). - Paid parental leave. - Commuter benefits including fully paid parking spots. - Referral bonus. - 401k retirement plan. - Equity program.

Technology

KMD Poland

Data Engineer ( Spark / Streaming / Java)

Senior

Remote

Warsaw, Poland

160 - 200 PLN

🏢 Summary: The offer is for a Data Engineer role focused on building and maintaining a large-scale, cloud-based energy market solution on Microsoft Azure. The position involves developing batch and streaming data processing pipelines using Apache Spark and Databricks within a distributed, event-driven microservices architecture. The role includes end-to-end responsibility for designing, implementing, optimizing, and maintaining scalable data solutions. 🗂️ Requirements: 4+ years of experience with Apache Spark, Experience with batch and streaming data processing, Experience with Apache Spark Structured Streaming, Experience with Apache Kafka, Experience designing technical solutions, Experience with distributed systems on cloud platforms, Experience with large-scale systems in microservices architecture, Experience with Git, Experience with CI/CD pipelines, Ability to design or implement deployment processes for data pipelines, Higher education in computer science or related field, Fluent English and Polish 📃 Skills: Spark, Databricks, Kafka, Delta, Java, SQL, Azure, Docker, Git, CI/CD, MS SQL, ElasticSearch, Redis, Azure Data Explorer, Helm, ArgoCD, GitOps, Microservices, DDD 🏢 Description: #Data Engineer #Apache Spark #Databricks #Java #Apache Kafka #Batch Processing #Structured Streaming #Azure #SQL #Microservices #CI/CD #Docker #DDD Are you ready to join our international team as a Data Engineer ? We shall tell you why you should... What product do we develop? We are building an innovative solution, KMD Elements , on Microsoft Azure cloud dedicated to the energy distribution market (electrical energy, gas, water, utility, and similar types of business). Our customers include institutions and companies operating in the energy market as transmission service operators, market regulators, distribution service operators , energy trading, and retail companies. KMD Elements delivers components allowing implementation of the full lifecycle of a customer on the energy market: meter data processing , connection to the network, physical network management, change of operator, full billing process support, payment, and debt management, customer communication, and finishing on customer account termination and network disconnection. The key market advantage of KMD Elements is its ability to support highly flexible, complex billing models as well as scalability to support large volumes of data. Our solution enables energy companies to promote efficient energy generation and usage patterns, supporting sustainable and green energy generation and consumption. We work with always up-to-date versions of: • Apache Spark on Azure Databricks • Apache Kafka • Delta Lake • Java • MS SQL Server and NoSQL storages like Elastic Search, Redis, Azure Data Explorer • Docker containers • Azure DevOps and fully automated CI/CD pipelines with Databricks Asset Bundles, ArgoCD, GitOps, Helm charts • Automated tests How do we work? #Agile #Scrum #Teamwork #CleanCode #CodeReview #Feedback #BestPracticies • We follow Scrum principles in our work – we work in biweekly iterations and produce production-ready functionalities at the end of each iteration – every 3 iterations we plan the next product release • We have end-to-end responsibility for the features we develop – from business requirements, through design and implementation up to running features on production • More than 75% of our work is spent on new product features • Our teams are cross-functional (7-8 persons) – they develop, test and maintain features they have built • Teams’ own domains in the solution and the corresponding system components • We value feedback and continuously seek improvements • We value software best practices and craftsmanship Product principles: • Domain model created using domain-driven design principles • Distributed event-driven architecture / microservices • Large-scale system for large volumes of data (>100TB data), processed by Apache Spark streaming and batch jobs powered by Databricks platform Your responsibilities: • Develop and maintain the leading IT solution for the energy market using Apache Spark, Databricks, Delta Lake, and Apache Kafka • Have end-to-end responsibility for the full lifecycle of features you develop • Design technical solutions for business requirements from the product roadmap • Maintain alignment with architectural principles defined on the project and organizational level • Ensure optimal performance through continuous monitoring and code optimization. • Refactor existing code and enhance system architecture to improve maintainability and scalability. • Design and evolve the test automation strategy, including technology stack and solution architecture. • Prepare reviews, participate in retrospectives, estimate user stories, and refine features ensuring their readiness for development. Personal requirements: • Have 4+ years of Apache Spark experience and have faced various data engineering challenges in batch or streaming • Have an interest in stream processing with Apache Spark Structured Streaming on top of Apache Kafka • Have experience leading technical solution designs • Have experience with distributed systems on a cloud platform • Have experience with large-scale systems in a microservice architecture • Are familiar with Git and CI/CD practice s and can design or implement the deployment process for your data pipelines • Possess a proactive approach and can-do attitude • Are excellent in English and Polish, both written and spoken • Have a higher education in computer science or a related field • Are a team player with strong communication skills Nice to have requirements: • Apache Spark Structured Streaming • Azure • Domain Driven Development • Docker containers and Kubernetes • Message brokers (i.e. Kafka) and event-driven architecture • Agile/Scrum Our offer: • Contract type: B2B • Work Mode : Flexible — this role supports on-site , hybrid , and remote arrangements, depending on your individual preferences. • Occasional on-site presence may be required — for example, onboard new team members, explore new business domains, or refine requirements in close collaboration with stakeholders or team building activities. What does the recruitment process look like? • Phone conversation with Recruitment Partner • Technical interview with the Hiring Team • Cognitive test • Offer

Technology

Datumo

Data Engineer (GCP)

Mid

Remote

Warsaw, Poland

16,000 - 28,000 PLN

🏢 Summary: The offer is for a Data Engineer responsible for building and optimizing scalable data platforms in cloud environments, mainly on Google Cloud Platform and Azure. The role focuses on designing big data solutions, developing ETL processes, and working with distributed data processing frameworks across international projects. It includes full remote work and opportunities for continuous technical development. 🗂️ Requirements: 3-4 years of commercial programming experience, Experience with Google Cloud Platform, Strong knowledge of Python, Strong knowledge of SQL, Strong knowledge of Scala or Java or Kotlin, Experience with BigQuery, Understanding of big data storage, modeling, processing and scheduling, Experience with Apache Spark or similar framework, Experience in data modeling and data storage, Experience with automated testing, CI/CD and code review, Experience collaborating with business stakeholders, English proficiency at minimum B2 level, Proficiency in Polish 📃 Skills: Python, SQL, Scala, Java, Kotlin, GCP, BigQuery, Spark, CI/CD, Git, Azure, Databricks, Kafka, Airflow, Docker, Kubernetes, dbt, PubSub, Dataflow, Flink 🏢 Description: We’re looking for a Data Engineer ready to push boundaries and grow with us. Datumo specializes in providing Data Engineering and Cloud Computing consulting services to clients from all over the world, primarily in Western Europe, Poland and the USA. Core industries we support include e-commerce 🛒, telecommunications 📡 and life sciences 🧬. Our team consists of exceptional people whose commitment allows us to conduct highly demanding projects. Our team members tend to stick around for more than 3 years, and when a project wraps up, we don't let them go - we embark on a journey to discover exciting new challenges for them. It's not just a workplace; it's a community that grows together! Must-have: ✅ at least 3-4 years of commercial experience in programming, ✅ proven record with a cloud provider - Google Cloud Platform, ✅ strong knowledge of Python, SQL and JVM languages (Scala or Java or Kotlin), ✅ experience in BigQuery data warehousing solution, ✅ in-depth understanding of big data aspects like data storage, modeling , processing , scheduling etc., ✅ understanding of Apache Spark (or similar distributed data processing framework), ✅ data modeling and data storage experience, ✅ ensuring solution quality through automatic tests, CI/CD and code review, ✅ proven collaboration with businesses, ✅ English proficiency at min. B2 level, proficient in Polish. Nice to have: 🌟 knowledge of dbt, Docker and Kubernetes, Apache Kafka, 🌟 familiarity with Apache Airflow or similar pipeline orchestrator, 🌟 another JVM (Java/Scala/Kotlin) programming language, 🌟 experience in Machine Learning projects, 🌟 familiarity with one of BI tools: Power BI/Looker/Tableau, 🌟 willingness to share knowledge (conferences, articles, open-source projects). What’s on offer: 🔥 100% remote work, with workation opportunity, 🔥 20 free days, 🔥 onboarding with a dedicated mentor, 🔥 project switching possible after a certain period, 🔥 individual budget for training and conferences, 🔥 benefits: Medicover Private Medical Care , co-financing of the Medicover Sport card, 🔥 opportunity to learn English with a native speaker, 🔥 regular company trips and informal get-togethers. Development opportunities in Datumo: 🚀 participation in industry conferences, 🚀 establishing Datumo's online brand presence, 🚀 support in obtaining certifications (e.g. GCP, Azure, Snowflake), 🚀 involvement in internal initiatives, like building technological roadmaps, 🚀 training budget, 🚀 access to internal technological training repositories. Discover our exemplary project: 🔌 IoT data ingestion to cloud The project integrates data from edge devices into the cloud using Azure services. The platform supports data streaming via either the IoT Edge environment with Java or Python modules, or direct connection using Kafka protocol to Event Hubs. It also facilitates batch data transmission to ADLS. Data transformation from raw telemetry to structured tables is done through Spark jobs in Databricks or data connections and update policies in Azure Data Explorer. ☁️ Petabyte-scale data platform migration to Google Cloud The goal of the project is to improve scalability and performance of the data platform by transitioning over a thousand active pipelines to GCP. The main focus is on rearchitecting existing Spark applications to either Cloud Dataproc or Cloud BigQuery SQL, depending on the Client’s requirements and automate it using Cloud Composer. 📈 Data analytics platform for investing company The project centers on developing and overseeing a data platform for an asset management company focused on ESG investing. Databricks is the central component. The platform, built on Azure cloud, integrates various Azure services for diverse functionalities. The primary task involves implementing and extending complex ETL processes that enrich investment data, using Spark jobs in Scala. Integrations with external data providers, as well as solutions for improving data quality and optimizing cloud resources, have been implemented. 🛒 Realtime Consumer Data Platform The initiative involves constructing a consumer data platform (CDP) for a major Polish retail company. Datumo actively participates from the project’s start, contributing to planning the platform’s architecture. The CDP is built on Google Cloud Platform (GCP), utilizing services like Pub/Sub, Dataflow and BigQuery. Open-source tools, including a Kubernetes cluster with Apache Kafka, Apache Airflow and Apache Flink, are used to meet specific requirements. This combination offers significant possibilities for the platform. Recruitment process: 1️⃣Tech quiz - 15 minutes 2️⃣ Soft skills interview - 30 minutes 3️⃣ Technical interview - 60 minutes Find out more by visiting our website - https://www.datumo.io If you like what we do and you dream about creating this world with us - don’t wait, apply now!