April 30, 2026
Principal Data Platform Engineer
Senior • On-site
San Diego, CA
Exactera has offices in New York City, Tarrytown NY, San Diego, CA, London, and Argentina.
The Role
As Principal Data Platform Engineer, you'll architect and implement our centralized data platform on Databricks. You'll establish governance patterns using Unity Catalog, optimize for cost and performance at scale, and enable our existing Data Engineers to build confidently on the platform. This is a data infrastructure role—focused on pipelines, storage, governance, and platform operations.
The Business Challenge
We operate multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning), each with distinct databases containing enterprise financial data—journal entries, general ledgers, and financial statements. Our immediate challenge is migrating multi-terabyte datasets from legacy systems to a unified Databricks lakehouse while establishing governance patterns that enable multi-product operations at scale.
What You'll Build
- Data Structuring: Design data models and implement unified schemas across multiple disparate product lines.
- Unity Catalog Architecture: Design and implement multi-catalog governance strategy supporting data isolation, cross-product data sharing, and comprehensive lineage tracking across our product portfolio
- Delta Lake Optimization: Establish patterns for Z-ordering, compaction, and liquid clustering at multi-TB scale. Define table structures, partitioning strategies, and retention policies that balance query performance with storage costs
- ETL Pipeline Framework: Build declarative pipeline patterns using Delta Live Tables. Create orchestration workflows for ingesting data from internal sources such as SQL databases and S3
- Third Party Integrations: Integrate with third party data sources such as ERP systems (Netsuite etc.) and external data providers (S&P etc.) with automated ingest, robust error handling and monitoring.
- Platform Operations: Implement cost monitoring and optimization strategies, establish data quality frameworks, create self-service patterns enabling Data Engineers to work independently while maintaining governance standards
Business Problems You'll Solve
- Key Legacy Product Migrations: Lead the architecture for migrating multi-terabyte datasets from legacy systems to Databricks—establishing patterns that will be reused across multiple product lines
- Multi-Product Data Architecture: Design Unity Catalog structures enabling secure data separation between product lines while allowing controlled cross-product analytics where appropriate
- Cost-Efficient Scale: Build infrastructure that scales efficiently—through intelligent caching, query optimization, and compute management strategies that avoid linear cost growth
- Platform Reliability: Establish monitoring, alerting, and data quality validation ensuring the platform operates reliably as foundation for both analytics and AI workloads
Required Experience
Databricks Expertise (Required)
- Unity Catalog: Production experience with multi-catalog governance, metastore design, and lineage tracking.
- Data Structuring: Experience designing and building unified schemas across multiple disparate product lines.
- Delta Lake: Expert-level experience with Z-ordering, compaction, liquid clustering, and performance tuning at multi-TB scale
- Delta Live Tables: Strong hands-on experience building declarative ETL pipelines, including change data capture and expectations/constraints
- Databricks Workflows: Experience with job orchestration, scheduling, and operational monitoring
- Business Intelligence: Experience enabling company-wide analytics and reporting with modern business intelligence tools and maintaining source of truth data and metrics.
- PySpark & Databricks SQL: Strong proficiency for code review, performance tuning, and query optimization
Core Platform Engineering
- 5-8 years in data engineering or data platform roles, with 3+ years hands-on Databricks experience
- Track record leading at least one significant platform build or migration project
- AWS experience (S3, IAM, VPC) with ability to collaborate on infrastructure decisions
- Infrastructure-as-code experience (Terraform preferred)
Technical Leadership
- Demonstrated ability architecting data platforms from first principles and defending technical decisions
- Strong written and verbal communication— document architecture decisions and present to both technical and business stakeholders
Preferred But Not Required
- Experience with financial data, accounting systems (NetSuite), or enterprise ERP platforms
- Background building platforms that serve AI/ML workloads (experience preparing data for downstream ML consumption, RAG and retrieval, and LLMs.
- Understand advanced intelligence concepts such as relationship surfacing with knowledge graphs
- Familiarity with data governance frameworks and compliance requirements for regulated industries
What We Offer:
(The following only applies to US-based positions)
-
A collaborative team culture with opportunities for career development.
-
Ample opportunities to be recognized, build valuable skills, and grow your career.
-
Generous vacation policy, including paid parental leave.
-
Comprehensive health plans with FSA and HSA options.
-
401(k) retirement plan.
-
Life and disability insurance coverage.
-
Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program.
About Us:
Similar jobs you might like
Technology

Exactera
Staff Data Engineer
Senior
On-site
San Diego, CA
🏢 Summary: Senior hands-on Staff Data Engineer role focused on owning and scaling a centralized Databricks lakehouse platform, building production-grade data pipelines, and leading multi-terabyte legacy migrations. The position combines architectural decision-making, data modeling, and technical leadership to support multiple financial product lines. The engineer acts as the primary onshore technical authority ensuring platform reliability, governance, and scalability. 🗂️ Requirements: 5+ years of data engineering experience, Strong proficiency in SQL, Python, and PySpark, Hands-on experience with Databricks (Delta Lake, Workflows, Databricks SQL), Experience building and maintaining large-scale ETL/ELT pipelines, Proven ability to design scalable data models, Experience migrating multi-terabyte datasets, AWS experience (S3, IAM, VPC), Infrastructure-as-code experience (Terraform preferred), Knowledge of data governance, lineage, and access controls, Ability to make independent architectural decisions, Experience mentoring data engineers, US-based with timezone overlap for stakeholder collaboration 📃 Skills: SQL, Python, PySpark, Databricks, Delta, Lake, Workflows, DatabricksSQL, UnityCatalog, AWS, S3, IAM, VPC, Terraform, ETL, ELT, DataModeling, Governance 🏢 Description: Exactera has offices in New York City, Tarrytown NY, San Diego, CA, London, and Argentina. The Role As Staff Data Engineer, you will provide senior onshore technical leadership for the data engineering team. You will own a defined slice of our centralized Databricks data platform with full accountability for decisions and delivery, serve as a technical counterpart to the Principal Data Platform Engineer, and drive architectural judgment and independent problem-solving as platform complexity scales post-migration. This is a hands-on data engineering role focused on building and maintaining production pipelines, exercising architectural judgment on data modeling and pipeline design, and serving as the onshore escalation point and institutional knowledge backup for platform decisions. The Business Challenge We operate multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning), each with distinct databases containing enterprise financial data—journal entries, general ledgers, and financial statements. Our immediate challenge is migrating multi-terabyte datasets from legacy systems to a unified Databricks lakehouse while establishing governance patterns that enable multi-product operations at scale. As the platform matures, the data engineering team needs senior onshore technical presence to drive architecture ownership and maintain platform quality. What You'll Build Production Data Pipelines: Build and maintain production data pipelines within the patterns and governance established by the Lead Data Platform Engineer, ensuring reliability and performance at multi-terabyte scale. Data Modeling & Architecture: Exercise architectural judgment on data modeling, pipeline design, and platform usage—translating complex business requirements into scalable data solutions across our product portfolio. Stakeholder Engagement: Engage proactively with product and engineering stakeholders to translate requirements into data solutions, serving as the primary onshore technical point of contact for data engineering needs. Platform Quality: Drive platform quality through code reviews, testing practices, and engineering standards that ensure the team delivers reliable, maintainable data infrastructure. Knowledge & Continuity: Serve as onshore escalation point and institutional knowledge backup for platform decisions, reducing single-point-of-failure risk and building onshore technical depth as the platform scales. Business Problems You'll Solve Multi-Product Data Delivery: Implement data pipelines that serve multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning) with distinct data requirements, ensuring each product gets the data it needs reliably and on schedule. Legacy Migration Execution: Lead pipeline implementation for migrating multi-terabyte datasets from legacy systems to Databricks, working within the architecture defined by the Lead Data Platform Engineer. Onshore Technical Leadership: Provide the senior judgment layer the current nearshore team cannot—owning problems end-to-end, making independent architectural decisions, and mentoring engineers to raise the quality bar across the team. Cross-Team Coordination: Bridge the gap between product teams and data infrastructure, translating business requirements into data solutions and ensuring the data platform delivers on product commitments. Required Experience Core Data Engineering SQL, Python, and PySpark—production pipeline implementation and performance optimization Databricks experience—Delta Lake, Workflows, and Databricks SQL; Unity Catalog familiarity preferred 5+ years in data engineering with demonstrated ability to own problems end-to-end without close direction Experience building and maintaining ETL/ELT pipelines at scale, including error handling, monitoring, and data quality validation Strong data modeling skills across structured and semi-structured data sources Platform & Infrastructure AWS experience (S3, IAM, VPC) with ability to collaborate on infrastructure decisions Infrastructure-as-code experience (Terraform preferred) Familiarity with data governance patterns (Unity Catalog, data lineage, access controls) Technical Leadership Demonstrated ability to exercise independent architectural judgment—not just ticket execution Experience mentoring or guiding junior and mid-level data engineers Strong written and verbal communication—able to document architecture decisions and engage directly with both technical and business stakeholders Onshore (US-based)—role requires timezone overlap, async-light communication, and direct stakeholder engagement Preferred But Not Required Experience with financial data, accounting systems (NetSuite), or enterprise ERP platforms Background building pipelines that serve AI/ML workloads (preparing data for downstream ML consumption, RAG, and LLMs) Familiarity with data governance frameworks and compliance requirements for regulated industries Experience working alongside or transitioning from nearshore engineering teams What We Offer (The following only applies to US-based positions) A collaborative team culture with opportunities for career development. Ample opportunities to be recognized, build valuable skills, and grow your career. Generous vacation policy, including paid parental leave. Comprehensive health plans with FSA and HSA options. 401(k) retirement plan. Life and disability insurance coverage. Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program. About Us At Exactera, a FinTech SaaS start-up founded in 2016, we stand at the intersection of human and machine intelligence. Our corporate tax solutions are powered by AI and cloud-based technologies, serving customers worldwide. We are committed to diversity, inclusion, and equal opportunities for all. What We Offer:(The following only applies to US-based positions) A collaborative team culture with opportunities for career development. Ample opportunities to be recognized, build valuable skills, and grow your career. Generous vacation policy, including paid parental leave. Comprehensive health plans with FSA and HSA options. 401(k) retirement plan. Life and disability insurance coverage. Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program. About Us: At Exactera, a FinTech SaaS start-up founded in 2016, we stand at the intersection of human and machine intelligence. Our corporate tax solutions are powered by AI and cloud-based technologies, serving customers worldwide. With over $100 million in funding from Savant Venture Fund and Insight Partners, we are poised for growth. We are committed to diversity, inclusion, and equal opportunities for all.
Technology

Exactera
Principal Product Manager (Agentic Platform)
Senior
On-site
San Diego, CA
🏢 Summary: Principal-level Product Manager role owning the strategy and roadmap for an AI-native, agentic platform that powers compliance-grade tax analysis on trusted data. The position defines how AI agents and practitioners access, use, and govern data, ensuring defensible, auditable outputs in regulated tax workflows. You will work closely with AI and Data Engineering leadership to shape platform capabilities, guardrails, and measurable practitioner outcomes. 🗂️ Requirements: 10+ years in product management, Experience owning technical platform, data, or ML products, Proven ownership of AI or ML platform from ambiguity to production, Strong understanding of lakehouse architectures and data contracts, Knowledge of agentic AI systems, retrieval, and tool calling, Experience designing guardrails, evaluation, and human-in-the-loop systems, Ability to define compliance-grade, auditable product requirements, Experience working cross-functionally without positional authority, Hands-on use of AI tools in daily work 📃 Skills: AI, ML, Lakehouse, Databricks, Delta, Retrieval, MCP, Data, Lineage, Evaluation, Guardrails, SaaS 🏢 Description: About the Role Exactera helps large multinationals handle complex tax work: transfer pricing, R&D tax credits, indirect tax, and audit defense. The company is moving from a software-and-services model to an AI-native platform that performs the work itself, with tax practitioners applying judgment where defensibility requires it. This role owns the product direction for the agentic platform behind that shift: the system that enables AI agents and practitioner tools to run tax analysis on trusted, current data. The data foundation (a lakehouse, external sources, and the contracts that keep them reliable) is the substrate. The product defines what agents and practitioners can do with it. You work directly with the Principal AI Engineer, who owns the AI/ML architecture, and the Principal Data Engineer, who owns the data tier. You own what the platform does, in what order, and why. Because you define how practitioners and agents use AI, you work deeply in these tools yourself. What You’ll Do The six outcomes that define success in the first 18 months: O1. Own the product strategy and roadmap for the platform. You own what the platform is for and what it builds next. The roadmap ties to practitioner outcomes rather than feature counts. When priorities compete across engineering, product, and practice teams, you decide and keep everyone aligned. O2. Define what compliance-grade means for this platform. Tax work must be defensible. You set product requirements that make platform output trustworthy: where results must be deterministic, what guardrails constrain agent behavior, how tool access is scoped, and what must be auditable. You determine where human judgment remains in the loop. O3. Own the agentic access surface. You own the interface AI agents and practitioner tools use to reach data and act on it, the contracts behind it, and the limits on agent capabilities. How this surface works determines how the entire platform is used. O4. Treat the data foundation as a product. With the Principal Data Engineer, you define contracts between data producers and consumers, including schema, quality, and freshness commitments practitioners can rely on. O5. Make external data a platform capability. You decide which regulatory, financial, and third-party data sources matter, how they appear in the product, and how to balance coverage, cost, and reliability. O6. Measure success by practitioner outcomes. Success is defined by faster, more accurate work with less manual effort. A production pipeline that does not improve a practitioner’s day is not successful. What We’re Looking For Required 10+ years in product management, with significant experience in technical platform, data, or ML products. Experience owning a platform or major product area at Principal-level scope across teams without positional authority. Ability to set direction across a modern data platform, including understanding of lakehouse and medallion patterns, data contracts, quality, and lineage. Track record of owning AI or ML platform initiatives from ambiguous problem definition to production. Strong understanding of agentic AI systems and data access patterns, including retrieval and tool calling. Experience designing AI systems where correctness and defensibility matter, including evaluation, guardrails, and human-in-the-loop workflows. Ability to define and measure platform success through practitioner results. Clear communication across engineering, product, and practice teams. Active use of AI tools in day-to-day work. Preferred Experience in tax, financial reporting, regulatory compliance, or other regulated, data-intensive domains. Experience building or governing data products, including external data sourcing and licensing trade-offs. Hands-on familiarity with Databricks, Delta Lake, or comparable lakehouse platforms in production. Background in tech-enabled services or SaaS-to-services transitions. Experience building a platform or major product area from an early stage. What We Offer A collaborative team culture with opportunities for career development. Ample opportunities to be recognized, build valuable skills, and grow your career. Generous vacation policy, including paid parental leave. Comprehensive health plans with FSA and HSA options. 401(k) retirement plan. Life and disability insurance coverage. Supplemental benefits such as a dependent care savings plan, pet insurance, will preparation, and an employee assistance program.
Technology
TechTree
Lead Data Engineer
Senior
Remote
Krakow, Poland
270,000 - 406,000 PLN/yr
🏢 Summary: Lead Data Engineer role focused on driving architecture and leading a team to build scalable, secure ETL/ELT pipelines and analytics-ready data models on modern cloud platforms. The position combines hands-on engineering with technical leadership, ensuring high standards in governance, observability, and performance optimisation. You will shape data infrastructure that supports large-scale analytics across the organisation. 🗂️ Requirements: Proven experience leading data engineering or analytics engineering teams, Strong programming skills in SQL, Strong programming skills in Python, Hands-on experience with Airflow or Prefect in production, Deep practical experience with dbt, Experience with Snowflake or Databricks at scale, Strong knowledge of dimensional modelling and SCD strategies, Experience implementing data quality and governance frameworks, Experience with CI/CD and automated testing in data systems 📃 Skills: SQL, Python, Airflow, Prefect, dbt, Snowflake, Databricks, ETL, ELT, CI/CD, SCD, Dimensional, Git, IaC 🏢 Description: ABOUT THE COMPANY Our client is a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you lead will have real consequence across the organisation. ABOUT THE ROLE We're looking for a Lead Data Engineer to combine deep technical expertise with hands-on team leadership, guiding a team of data engineers building and maintaining ETL/ELT pipelines, data models, and governance frameworks that power analytics and reporting across the organisation. This is a technical leadership role — you'll drive architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable data solutions. You'll collaborate closely with stakeholders to align technical work with business objectives, champion governance and observability standards, and foster a culture of continuous improvement. The expectation is that you're equally effective in an architecture review as you are pairing with an engineer on a tricky pipeline problem. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data engineers, promoting collaboration, knowledge sharing, and professional growth. Set the standard for engineering quality and hold the bar consistently. Architecture and pipeline design Drive architectural decisions for ETL/ELT pipelines, orchestration frameworks (Airflow/Prefect), and transformation layers (dbt). Facilitate architecture reviews and contribute to design decisions for scalable, fault-tolerant systems. Analytics data modelling Oversee design and implementation of analytics-ready data models — dimensional schemas, SCD strategies, and semantic layers — that internal teams can build on reliably. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. Platform optimisation Manage performance tuning and cost optimisation for Snowflake, Databricks, and related cloud data platforms at scale. Governance and observability Champion governance, observability, and compliance frameworks across all data workflows — including data quality, lineage tracking, and multi-tenant environment controls. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure timely delivery aligned with business objectives. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or analytics-focused development teams — mentoring engineers, driving architectural decisions, and owning delivery outcomes. SQL and Python Strong programming skills in both SQL and Python, applied to production data systems at scale. ETL/ELT orchestration Hands-on experience with orchestration tools — Airflow and/or Prefect — in production pipeline environments. dbt expertise Deep practical experience with dbt for transformation workflows and analytics modelling, including testing, documentation, and modular project design. Snowflake and Databricks Familiarity with Snowflake and/or Databricks for large-scale data processing, including performance tuning and cost management. Data modelling principles Solid understanding of data modelling principles, incremental strategies, and schema design for analytics — dimensional modelling, SCDs, and semantic layer design. Governance and data quality Knowledge of data quality frameworks, lineage tracking, and governance in multi-tenant environments. Software engineering practices Familiarity with CI/CD, automated testing, and infrastructure-as-code practices applied to data systems. Communication and stakeholder management Strong communication skills with the ability to operate confidently across technical teams and business stakeholders. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements, hybrid model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology
TechTree
Advanced Data Platform Engineer
Senior
Remote
Krakow, MA, Poland
160,000 - 240,000 PLN/yr
🏢 Summary: The role focuses on designing and building scalable, cloud-native data platforms to enable advanced analytics and reporting across a large internal data ecosystem. It involves developing distributed data pipelines, lakehouse architectures, and optimised data warehousing solutions with strong emphasis on performance, governance, and reliability. The position requires deep technical expertise in big data technologies and cloud-native infrastructure, including on-call responsibility for platform stability. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Commercial experience with Apache Spark for distributed data processing, Hands-on experience with Delta Lake and/or Apache Iceberg in production, Experience with dbt for SQL transformations, Experience with Databricks and Snowflake, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Understanding of performance tuning and cost optimisation in large-scale data systems, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. Over 75% of our business has transitioned to our cloud platform, and we are making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you build will have real consequence. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As an Advanced Data Platform Engineer, you'll design and implement scalable, cloud-native data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services to support diverse analytical use cases at enterprise scale. The role emphasises technical depth — performance optimisation, governance best practices, and the kind of engineering rigour that keeps vast datasets accessible, secure, and compliant. You'll work closely with internal teams to deliver curated datasets and self-service analytics capabilities, and you'll participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems design Design and implement complex data pipelines and distributed systems using Spark and Python, applying clean code principles, modular design, CI/CD, automated testing, and thorough code reviews. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Apache Iceberg, ensuring reliability, performance, and long-term maintainability at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Deliver curated datasets and self-service analytics capabilities that empower internal stakeholders to explore data independently. Data warehousing optimisation Optimise Databricks and Snowflake environments for performance and scalability. Drive cost optimisation and performance tuning across Spark jobs and cloud-native infrastructure. Observability and governance Implement observability and governance frameworks including data lineage tracking and compliance controls, ensuring data remains secure and auditable. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in Python and SQL — the foundation for everything you'll build here. Apache Spark Solid experience with Spark for distributed data processing at scale, including performance tuning and optimisation. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the tradeoffs and have used these in production environments. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and SQL transformation pipelines. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Scalability and cost optimisation Understanding of performance tuning, scalability strategies, and cost optimisation for large-scale data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 160,000 – 240,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology
TechTree
Senior Data Platform Engineer
Senior
Remote
Krakow, Poland
208,000 - 312,000 PLN/yr
🏢 Summary: Senior Data Platform Engineer role focused on building and optimising a cloud-native lakehouse platform for large-scale analytics and reporting. The position involves designing distributed data pipelines, enabling self-service analytics, and implementing governance and observability frameworks using modern data technologies. You will work with Spark-based systems and integrated data warehousing solutions to deliver scalable, reliable data platforms. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production environments, Experience with Delta Lake and/or Apache Iceberg in production, Practical experience with dbt for data transformations, Experience with Databricks and Snowflake, Understanding of data governance and lineage in large-scale environments, Familiarity with Kubernetes and Docker, Experience with CI/CD and automated testing practices, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platforms you build will have real consequence across the organisation. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As a Senior Data Platform Engineer, you'll combine strong software engineering principles with deep data expertise to build robust, cloud-native platforms that process large-scale datasets efficiently and enable internal teams to build reporting and analytics on top of them. The role emphasises cloud-native architecture, lakehouse integration, data warehousing, and governance best practices. You'll work on systems using Apache Spark, Delta Lake, and Iceberg, and help deliver curated data models and self-service analytics capabilities to internal stakeholders. You'll also participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems Design and implement scalable data pipelines and distributed systems using Spark and Python to process and transform large-scale datasets for analytics and reporting. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Iceberg, ensuring data reliability, versioning, and performance optimisation at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Collaborate with internal teams to deliver curated datasets and self-service analytics capabilities for reporting and advanced use cases. Data warehousing optimisation Integrate and optimise Databricks and Snowflake for scalable storage and query performance. Drive performance tuning and cost optimisation across Spark jobs and cloud-native environments. Governance and observability Implement observability and governance frameworks including data lineage, quality checks, and compliance controls. Build platforms that allow secure and compliant access to diverse data sources. Engineering best practices Apply and champion clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in both Python and SQL, applied to production data platform work at scale. Apache Spark Solid hands-on experience with Spark for distributed data processing, including performance tuning in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You've applied these in production and understand the trade-offs in real-world scenarios. dbt and analytics tooling Practical experience with dbt for transformation workflows. Familiarity with Databricks and Snowflake for large-scale analytics workloads. Data governance and compliance Understanding of data governance, lineage tracking, and compliance requirements in large-scale, multi-tenant data environments. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 208,000 – 312,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology

SugarCRM
Senior Data Engineer - Databricks
Senior
Hybrid
Denver, CO
155,004 - 185,004 USD
🏢 Summary: Senior Data Engineer role focused on building and optimizing Databricks-based data pipelines for the Sugar Predict platform, integrating ERP and CRM data across Azure and AWS environments. The position involves production support, legacy ETL modernization, multi-tenant data architecture, and collaboration with ML and product teams to deliver scalable revenue intelligence solutions. Hybrid work arrangement with in-office collaboration in Denver, CO. 🗂️ Requirements: 4+ years of data engineering experience, 2+ years of Databricks or Apache Spark experience, Proficiency in PySpark, Proficiency in SQL, Proficiency in Python, Experience building production-grade data pipelines, Hands-on experience with Delta Lake, Experience with pipeline performance tuning, Knowledge of PostgreSQL, Experience maintaining legacy ETL tooling, Experience with multi-tenant architectures, Understanding of data governance and security principles, Ability to support on-call incident response, Experience with CI/CD practices 📃 Skills: Databricks, Spark, PySpark, SQL, Python, DeltaLake, PostgreSQL, SSIS, Informatica, Azure, AWS, ETL, ELT, CI/CD, Serverless, SQLServer, Feast 🏢 Description: About SugarAI SugarAI is redefining CRM for the age of AI. We help teams turn fragmented customer and revenue signals into clear, prioritized action through intelligent CRM solutions. Where You Fit In: The Sugar Predict platform powers revenue intelligence for mid-market enterprises by fusing ERP and CRM data into actionable insights. As a Senior Data Engineer, you will own the Databricks pipelines that make this possible, driving production reliability, cost efficiency, and platform growth through customer onboarding and legacy modernization. You will work closely with ML engineers, product teams, and the Enterprise Architecture team to ensure the data backbone behind Sugar Predict is always fast, clean, and ready to deliver at a global scale. This role operates on a hybrid model, with a mix of remote work and in-office collaboration at our Denver, CO location, specifically, working in-office a minimum of 2-3 days per week. Impact You Will Make in the Role: • Own Databricks production support for the Sugar Predict data platform, including monitoring, alerting, and incident response across all production data flows • Maintain and report on SLA performance metrics for data pipeline delivery • Identify and implement pipeline optimizations that reduce Databricks compute costs and improve throughput • Migrate legacy ETL/ELT pipelines to Databricks and build automation tooling • Support new customer onboarding by provisioning and validating tenant data pipelines • Design and build high-performance Databricks pipelines across Azure and AWS environments • Own Delta Lake architecture including schema design and incremental processing patterns • Enforce data security best practices including role-based access control and secrets management • Implement data quality monitoring and observability across pipeline health and ML model inputs • Apply and enforce multi-tenant data isolation patterns • Partner with the Enterprise Architecture team on platform integration • Support globally distributed operations through on-call rotation and incident response • Maintain technical documentation, runbooks, and architectural decision records • Apply CI/CD best practices including version control, automated testing, and deployment tooling What You Will Bring: • 4+ years of data engineering experience • At least 2 years on Databricks or Apache Spark across Azure and/or AWS • Proficiency in PySpark, SQL, and Python • Hands-on experience with Delta Lake including schema evolution and ACID transactions • Experience with pipeline performance tuning and compute optimization • Solid working knowledge of PostgreSQL • Experience supporting legacy ETL tooling such as SSIS, Informatica, or custom Python/SQL pipelines • Experience supporting large-scale multi-tenant architectures • Proven ability to collaborate across data science, product, and infrastructure teams • Strong understanding of data governance, security, and compliance principles Preferred Qualifications/Experience: • Experience operating Databricks workspaces across Azure and AWS • Experience optimizing Databricks workloads in a Serverless environment • Experience with Microsoft SQL Server • Exposure to ML feature engineering or feature stores • Experience with customer onboarding automation or IaC patterns • Databricks Certified Data Engineer Associate or Professional certification Benefits and Perks: • Excellent healthcare package for you and your family • 401(k) match • Unlimited Paid Time Off • Paid Parental Leave • Online Legal Services • Financial Planning Services • Discounted Pet Insurance • Corporate Benefit Program with travel and entertainment discounts • Health and Wellness Reimbursement Program • Travel Discounts • Educational Resources and Career Development Program • Employee Referral Bonus Program • Merit-based career growth opportunities Our company uses E-Verify to confirm the employment eligibility of all newly hired employees.
Technology

Nimble Robotics
Data Engineer II/III
Mid
On-site
San Francisco, CA
140,004 - 200,004 USD/yr
🏢 Summary: Data Engineer role focused on building and scaling data infrastructure for advanced robotics systems. The position involves designing reliable batch and real-time pipelines, optimizing ETL/ELT processes, and implementing data governance across cloud platforms. You will collaborate cross-functionally to deliver high-quality, scalable data solutions supporting company-wide analytics and operations. 🗂️ Requirements: BS/MS/PhD in Computer Science, Mathematics, Computer Engineering or related field, or equivalent practical experience, 1–5 years of professional experience, Proficiency in Python and SQL, Hands-on experience with Rust, Go, Java, C#, or C++, Experience with Kafka, Spark, and lakehouse formats (Icebert, DeltaLake), Experience with AWS, GCP, or Azure, Strong debugging and problem-solving skills, Willingness to work extended hours and weekends if needed, Ability to work full time onsite in San Francisco 📃 Skills: Python, SQL, Rust, Go, Java, C#, C++, Kafka, Spark, Icebert, DeltaLake, AWS, GCP, Azure, Clickhouse, Flink, Databricks 🏢 Description: About the Role Help us advance our robotics moonshot by scaling our data engineering efforts. Drive design and development of data infrastructure across our products and internal tools. You will play a critical role in working with a cross-functional team to architect and build advanced robotic systems. Responsibilities - Design, build, and maintain scalable, reliable data pipelines that support both batch and real-time analytics. - Develop and optimize ETL/ELT processes to ingest, transform, and integrate data from multiple sources into data warehouses and lakehouse platforms. - Drive and apply best practices in data modeling, pipeline architecture, query optimization, data quality, and data engineering standards. - Collaborate closely with finance, bizops, and engineering teams to understand data needs and deliver solutions. - Provide data engineering support to ensure data accessibility and usability company-wide. - Implement and maintain data governance frameworks to support compliance with internal policies and external regulatory requirements. - Evaluate and adopt modern data engineering technologies, tools, and best practices to improve scalability, reliability, and operational efficiency. - Document data engineering processes, designs, and architectures. Qualifications - BS/MS/PhD in Computer Science, Mathematics, Computer Engineering, or a related field, or equivalent practical experience. - 1-5 years of professional experience. - Proficiency in writing production-grade code in Python and SQL. - Hands-on experience with one of the following languages: Rust, Go, Java, C#, C++. - Strong debugging skills and the ability to diagnose and resolve issues efficiently. - Experience with Kafka, Spark and common lakehouse formats like Icebert, DeltaLake. - Experience working with cloud platforms like AWS, GCP, Azure. Preferred Experience - Experience with Clickhouse, Apache Flink, Databricks. - Experience on financial reporting and understanding of accountability. Additional Requirements - Willing to work extended hours and weekends if needed. - This position is based full time in our San Francisco headquarters. Compensation The pay range for this position at the start of employment is expected to be between $140,000 - $200,000/year. The exact offer may vary depending on job-related knowledge, skills, and experience. In addition to cash compensation, this position will also receive generous equity. Benefits - Unlimited Flexible Time Off. - Health insurance (medical, dental, vision). - Paid parental leave. - Commuter benefits including fully paid parking spots. - Referral bonus. - 401k retirement plan. - Equity program.
Technology

Vulcan Elements
Data Engineer
Senior
On-site
Research Triangle Park, NC
🏢 Summary: Data Engineer role focused on designing and scaling data infrastructure, ETL pipelines, and Lakehouse architecture for a manufacturing environment supporting analytics and AI workloads. The position involves building operational data systems, ensuring data quality, and integrating industrial and manufacturing data sources. Candidates will collaborate cross-functionally and help establish scalable data architecture standards for future facility growth. 🗂️ Requirements: 8+ years of data engineering or data infrastructure experience, Experience designing data lakes or Lakehouse platforms, Experience building ETL/ELT pipelines, Strong data modeling expertise, Experience with relational databases, Strong SQL skills, Ability to document architecture and technical decisions, Experience collaborating with technical and non-technical stakeholders, U.S. Person status for export-controlled access 📃 Skills: SQL, PostgreSQL, SQLServer, ETL, ELT, Lakehouse, Python, Airflow, Prefect, dbt, InfluxDB, TimescaleDB, MQTT, DeltaLake, ApacheIceberg, AWS, Azure, GCP 🏢 Description: Vulcan Elements is manufacturing American rare-earth permanent magnets for a secure, resilient future. With a focus on national security and economic resiliency, we serve critical industries such as defense, aerospace, and automotive, powering a high-technology future. Vulcan Elements is building a team of ambitious professionals committed to Mission Focus, Technical Excellence, and Transparency. As the Data Engineer, you will design and build the data infrastructure that makes Vulcan's operational and business data useful — first at pilot scale, and then as the foundation for a 10,000 ton/year facility. You will work from architecture to implementation: evaluating and selecting platforms, designing data models and pipelines, and building the systems that collect, contextualize, and deliver data to the teams and tools that depend on it. You will collaborate closely with cross-functional stakeholders to translate operational requirements into a durable, scalable data architecture. As Vulcan grows, this role has the opportunity to expand into a team leadership position. Responsibilities Architecture & Platform Design - Design and own Vulcan's data architecture from operational data stores through ETL pipelines to the analytics and AI layer - Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, weighing scalability, compliance requirements, operational burden, and cost - Review, refine, and implement data architecture design documents, ensuring designs are technically sound and account for CUI and ITAR data handling requirements - Make and document key platform and design decisions with enough clarity that future team members can understand the reasoning and build on it - Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign - Apply sound engineering practices to everything you build: version control, testing, observability, and documentation, and hold those standards as the data team grows Data Pipeline & Integration - Design and build ETL pipelines that move data from operational data stores into the data Lakehouse with full contextual enrichment, making it ready for analytics and AI workloads - Build reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems - Collaborate across engineering, operations, and IT to understand data flows, dependencies, and integration requirements, and translate them into pipeline and architecture decisions - Identify and eliminate manual data workflows, replacing them with monitored, reliable pipelines - Diagnose and resolve data quality issues across the stack, and build monitoring into pipelines so problems surface early Data Modeling & Quality - Define data models that support operational queries, analytical workloads, and future AI and ML applications - Own data contextualization standards ensuring every data point carries the metadata needed to make it meaningful - Contribute to schema design and payload definitions for operational data stores, working toward consistency and legibility across the organization - Support the development of reporting and visibility tools that give operations and leadership clear insight into process and quality data - Write clear technical documentation for architecture decisions, data models, pipeline designs, and operational runbooks Responsibilities and tasks outlined are not exhaustive and may change as determined by the needs of the business. Qualifications - 8+ years of experience in data engineering, data infrastructure, or a closely related technical role with a track record of owning and delivering production systems - Demonstrated experience designing and building data lakes, Lakehouses, or analytical data stores; understands the tradeoffs between platforms and can make and defend platform selection decisions - Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data - Deep fluency with data modeling for both operational and analytical workloads; can design schemas that serve present needs without foreclosing future ones - Experience with relational databases (PostgreSQL, SQL Server, or similar); writes and debugs SQL confidently - Comfortable working in a fast-moving environment with a small team, making decisions with incomplete information and documenting them clearly for future colleagues - Strong communicator who can work across technical and non-technical stakeholders and translate between operational requirements and data architecture decisions - Must be a U.S. Person due to required access to U.S. export-controlled information or facilities Desired Skills - Experience with time-series databases (InfluxDB, TimescaleDB, or similar) common in industrial and IoT environments - Familiarity with industrial data concepts — historian data, process tags, OT/IT integration — and the data challenges specific to manufacturing environments - Experience working on or alongside a Unified Namespace or MQTT-based data architecture; understands how industrial messaging infrastructure relates to the data layer - Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar) - Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar) - Comfort with scripting and lightweight development (Python, SQL, or similar) for pipeline development and data quality tooling - Familiarity with cloud platforms (AWS, Azure, or GCP) and experience evaluating on-premises vs. cloud tradeoffs for data infrastructure - Experience working in a controlled information environment; familiarity with the handling requirements for Controlled Unclassified Information (CUI) or export-controlled technical data under ITAR or EAR - Experience in a manufacturing, industrial, or operations-heavy environment
Technology

Vulcan Elements
Data Engineer
Senior
On-site
Durham, NC
🏢 Summary: Data Engineer role focused on designing and scaling data infrastructure, ETL pipelines, and Lakehouse architecture for manufacturing operations supporting analytics and AI workloads. The position involves building reliable industrial data systems, defining data models, and collaborating across engineering and operations teams in a secure, compliance-driven environment. There is potential for future leadership responsibilities as the organization grows. 🗂️ Requirements: 8+ years of experience in data engineering or data infrastructure, Experience designing and building data lakes or Lakehouse platforms, Experience building ETL/ELT pipelines, Strong data modeling experience for operational and analytical workloads, Experience with relational databases, Strong SQL skills, Ability to work in fast-moving environments with small teams, Ability to communicate across technical and non-technical stakeholders, U.S. Person status for access to export-controlled information 📃 Skills: PostgreSQL, SQLServer, SQL, ETL, ELT, Lakehouse, Python, InfluxDB, TimescaleDB, MQTT, DeltaLake, Iceberg, Airflow, Prefect, dbt, AWS, Azure, GCP 🏢 Description: Vulcan Elements is manufacturing American rare-earth permanent magnets for a secure, resilient future. With a focus on national security and economic resiliency, the company serves critical industries such as defense, aerospace, and automotive. As the Data Engineer, you will design and build the data infrastructure that makes operational and business data useful — first at pilot scale, and then as the foundation for a 10,000 ton/year facility. You will work from architecture to implementation: evaluating and selecting platforms, designing data models and pipelines, and building the systems that collect, contextualize, and deliver data to the teams and tools that depend on it. You will collaborate closely with cross-functional stakeholders to translate operational requirements into a durable, scalable data architecture. As the organization grows, this role has the opportunity to expand into a team leadership position. Responsibilities Architecture & Platform Design - Design and own data architecture from operational data stores through ETL pipelines to the analytics and AI layer - Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, weighing scalability, compliance requirements, operational burden, and cost - Review, refine, and implement data architecture design documents, ensuring designs are technically sound and account for CUI and ITAR data handling requirements - Make and document key platform and design decisions with enough clarity that future team members can understand the reasoning and build on it - Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign - Apply sound engineering practices to everything you build: version control, testing, observability, and documentation, and hold those standards as the data team grows Data Pipeline & Integration - Design and build ETL pipelines that move data from operational data stores into the data Lakehouse with full contextual enrichment, making it ready for analytics and AI workloads - Build reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems - Collaborate across engineering, operations, and IT to understand data flows, dependencies, and integration requirements, and translate them into pipeline and architecture decisions - Identify and eliminate manual data workflows, replacing them with monitored, reliable pipelines - Diagnose and resolve data quality issues across the stack, and build monitoring into pipelines so problems surface early Data Modeling & Quality - Define data models that support operational queries, analytical workloads, and future AI and ML applications - Own data contextualization standards ensuring every data point carries the metadata needed to make it meaningful - Contribute to schema design and payload definitions for operational data stores, working toward consistency and legibility across the organization - Support the development of reporting and visibility tools that give operations and leadership clear insight into process and quality data - Write clear technical documentation for architecture decisions, data models, pipeline designs, and operational runbooks Responsibilities and tasks outlined are not exhaustive and may change as determined by business needs. Qualifications - 8+ years of experience in data engineering, data infrastructure, or a closely related technical role with a track record of owning and delivering production systems - Demonstrated experience designing and building data lakes, Lakehouses, or analytical data stores; understands the tradeoffs between platforms and can make and defend platform selection decisions - Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data - Deep fluency with data modeling for both operational and analytical workloads - Experience with relational databases (PostgreSQL, SQL Server, or similar) - Writes and debugs SQL confidently - Comfortable working in a fast-moving environment with a small team - Strong communicator able to work across technical and non-technical stakeholders - Must be a U.S. Person due to required access to U.S. export-controlled information or facilities Desired Skills - Experience with time-series databases (InfluxDB, TimescaleDB, or similar) - Familiarity with industrial data concepts including historian data, process tags, and OT/IT integration - Experience with Unified Namespace or MQTT-based data architecture - Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar) - Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar) - Comfort with scripting and lightweight development (Python, SQL, or similar) - Familiarity with cloud platforms (AWS, Azure, or GCP) - Experience working in controlled information environments with CUI, ITAR, or EAR requirements - Experience in manufacturing, industrial, or operations-heavy environments
Technology
TechTree
Lead Distributed Data Platform Engineer
Senior
Remote
Warsaw, Poland
270,000 - 406,000 PLN/yr
🏢 Summary: Lead Distributed Data Platform Engineer responsible for architecting and delivering enterprise-scale lakehouse and distributed data platforms to enable advanced analytics and reporting. The role combines hands-on technical leadership with team mentorship, driving scalable, secure, and high-performance data solutions in cloud-native environments. You will guide architectural decisions, enforce engineering best practices, and ensure platform reliability at scale. 🗂️ Requirements: Proven experience leading data engineering or platform teams, Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production, Experience with Delta Lake and/or Apache Iceberg in production, Experience designing distributed systems and lakehouse architectures, Experience building scalable data pipelines, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Experience working in cloud-native environments 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD, Java, Scala, Rust 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platform you lead will underpin how the entire organisation accesses and acts on its data. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As Lead Distributed Data Platform Engineer, you'll combine deep technical expertise with hands-on team leadership — guiding a team in designing and maintaining data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services at enterprise scale. You'll lead architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable solutions. The role emphasises technical leadership, governance best practices, and a culture of innovation and continuous improvement. You'll also participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data platform engineers, promoting collaboration, knowledge sharing, and professional growth. Set and maintain high engineering standards across the team. Distributed systems architecture Drive architectural decisions for distributed systems and lakehouse platforms using Spark, Delta Lake, and Iceberg. Facilitate architecture reviews and contribute to design decisions for fault-tolerant, future-ready systems. Data pipeline and platform delivery Oversee design and implementation of scalable data pipelines and analytics workflows, ensuring they are reliable, performant, and maintainable at scale. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all platform engineering work. Performance and cost optimisation Manage performance tuning, scalability strategies, and cost optimisation across cloud-native environments and large-scale distributed workloads. Governance and observability Champion governance, observability, and compliance frameworks across all data platforms — ensuring data remains accessible, secure, and auditable. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure delivery aligns with business objectives and analytics needs. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or platform development teams — mentoring engineers, owning architectural decisions, and driving delivery outcomes. Python and SQL Strong programming skills in both Python and SQL applied to production data platform work at scale. Apache Spark Hands-on experience with Spark for distributed data processing, including performance tuning and optimisation in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the trade-offs and have applied these technologies in production at scale. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and large-scale data processing. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data platform systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Communication and stakeholder management Strong communication skills with the confidence to operate across engineering teams, cross-functional partners, and senior leadership. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. Exposure to service mesh and advanced orchestration patterns. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.