April 30, 2026

Staff Data Engineer

Senior • On-site

San Diego, CA

Exactera has offices in New York City, Tarrytown NY, San Diego, CA, London, and Argentina.

The Role

As Staff Data Engineer, you will provide senior onshore technical leadership for the data engineering team. You will own a defined slice of our centralized Databricks data platform with full accountability for decisions and delivery, serve as a technical counterpart to the Principal Data Platform Engineer, and drive architectural judgment and independent problem-solving as platform complexity scales post-migration.

This is a hands-on data engineering role focused on building and maintaining production pipelines, exercising architectural judgment on data modeling and pipeline design, and serving as the onshore escalation point and institutional knowledge backup for platform decisions.

The Business Challenge

We operate multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning), each with distinct databases containing enterprise financial data—journal entries, general ledgers, and financial statements. Our immediate challenge is migrating multi-terabyte datasets from legacy systems to a unified Databricks lakehouse while establishing governance patterns that enable multi-product operations at scale. As the platform matures, the data engineering team needs senior onshore technical presence to drive architecture ownership and maintain platform quality.

What You'll Build

  • Production Data Pipelines: Build and maintain production data pipelines within the patterns and governance established by the Lead Data Platform Engineer, ensuring reliability and performance at multi-terabyte scale.
  • Data Modeling & Architecture: Exercise architectural judgment on data modeling, pipeline design, and platform usage—translating complex business requirements into scalable data solutions across our product portfolio.
  • Stakeholder Engagement: Engage proactively with product and engineering stakeholders to translate requirements into data solutions, serving as the primary onshore technical point of contact for data engineering needs.
  • Platform Quality: Drive platform quality through code reviews, testing practices, and engineering standards that ensure the team delivers reliable, maintainable data infrastructure.
  • Knowledge & Continuity: Serve as onshore escalation point and institutional knowledge backup for platform decisions, reducing single-point-of-failure risk and building onshore technical depth as the platform scales.

Business Problems You'll Solve

  • Multi-Product Data Delivery: Implement data pipelines that serve multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning) with distinct data requirements, ensuring each product gets the data it needs reliably and on schedule.
  • Legacy Migration Execution: Lead pipeline implementation for migrating multi-terabyte datasets from legacy systems to Databricks, working within the architecture defined by the Lead Data Platform Engineer.
  • Onshore Technical Leadership: Provide the senior judgment layer the current nearshore team cannot—owning problems end-to-end, making independent architectural decisions, and mentoring engineers to raise the quality bar across the team.
  • Cross-Team Coordination: Bridge the gap between product teams and data infrastructure, translating business requirements into data solutions and ensuring the data platform delivers on product commitments.

Required Experience

Core Data Engineering

  • SQL, Python, and PySpark—production pipeline implementation and performance optimization
  • Databricks experience—Delta Lake, Workflows, and Databricks SQL; Unity Catalog familiarity preferred
  • 5+ years in data engineering with demonstrated ability to own problems end-to-end without close direction
  • Experience building and maintaining ETL/ELT pipelines at scale, including error handling, monitoring, and data quality validation
  • Strong data modeling skills across structured and semi-structured data sources

Platform & Infrastructure

  • AWS experience (S3, IAM, VPC) with ability to collaborate on infrastructure decisions
  • Infrastructure-as-code experience (Terraform preferred)
  • Familiarity with data governance patterns (Unity Catalog, data lineage, access controls)

Technical Leadership

  • Demonstrated ability to exercise independent architectural judgment—not just ticket execution
  • Experience mentoring or guiding junior and mid-level data engineers
  • Strong written and verbal communication—able to document architecture decisions and engage directly with both technical and business stakeholders
  • Onshore (US-based)—role requires timezone overlap, async-light communication, and direct stakeholder engagement

Preferred But Not Required

  • Experience with financial data, accounting systems (NetSuite), or enterprise ERP platforms
  • Background building pipelines that serve AI/ML workloads (preparing data for downstream ML consumption, RAG, and LLMs)
  • Familiarity with data governance frameworks and compliance requirements for regulated industries
  • Experience working alongside or transitioning from nearshore engineering teams

What We Offer

(The following only applies to US-based positions)

  • A collaborative team culture with opportunities for career development.
  • Ample opportunities to be recognized, build valuable skills, and grow your career.
  • Generous vacation policy, including paid parental leave.
  • Comprehensive health plans with FSA and HSA options.
  • 401(k) retirement plan.
  • Life and disability insurance coverage.
  • Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program.

About Us

At Exactera, a FinTech SaaS start-up founded in 2016, we stand at the intersection of human and machine intelligence. Our corporate tax solutions are powered by AI and cloud-based technologies, serving customers worldwide. We are committed to diversity, inclusion, and equal opportunities for all.

What We Offer:
(The following only applies to US-based positions)

  • A collaborative team culture with opportunities for career development.
  • Ample opportunities to be recognized, build valuable skills, and grow your career.
  • Generous vacation policy, including paid parental leave.
  • Comprehensive health plans with FSA and HSA options.
  • 401(k) retirement plan.
  • Life and disability insurance coverage.
  • Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program.

About Us:

At Exactera, a FinTech SaaS start-up founded in 2016, we stand at the intersection of human and machine intelligence. Our corporate tax solutions are powered by AI and cloud-based technologies, serving customers worldwide. With over $100 million in funding from Savant Venture Fund and Insight Partners, we are poised for growth. We are committed to diversity, inclusion, and equal opportunities for all.

Similar jobs you might like

Technology

Exactera

Principal Data Platform Engineer

Senior

On-site

San Diego, CA

🏢 Summary: Principal Data Platform Engineer role focused on architecting and implementing a centralized Databricks lakehouse platform, including governance, ETL frameworks, and large-scale data migration. The position centers on building scalable, cost-efficient, and well-governed data infrastructure across multiple product lines. It involves leading multi-terabyte legacy migrations and establishing robust platform operations for analytics and AI workloads. 🗂️ Requirements: 5-8 years in data engineering or data platform roles, 3+ years hands-on experience with Databricks, Production experience with Unity Catalog and multi-catalog governance, Expert experience with Delta Lake optimization at multi-TB scale, Strong hands-on experience with Delta Live Tables, Experience with Databricks Workflows for orchestration and monitoring, Strong proficiency in PySpark and Databricks SQL, Experience designing unified schemas across disparate data sources, AWS experience including S3, IAM, VPC, Infrastructure-as-code experience (Terraform preferred), Experience leading at least one significant platform build or migration project, Ability to architect data platforms and document technical decisions 📃 Skills: Databricks, UnityCatalog, DeltaLake, DeltaLiveTables, PySpark, DatabricksSQL, AWS, S3, IAM, VPC, Terraform, SQL, ETL, Z-ordering, Compaction, Clustering 🏢 Description: Exactera has offices in New York City, Tarrytown NY, San Diego, CA, London, and Argentina. The Role As Principal Data Platform Engineer, you'll architect and implement our centralized data platform on Databricks. You'll establish governance patterns using Unity Catalog, optimize for cost and performance at scale, and enable our existing Data Engineers to build confidently on the platform. This is a data infrastructure role—focused on pipelines, storage, governance, and platform operations. The Business Challenge We operate multiple product lines (Transfer Pricing, R&D Services, RoyaltyStat, Provisioning), each with distinct databases containing enterprise financial data—journal entries, general ledgers, and financial statements. Our immediate challenge is migrating multi-terabyte datasets from legacy systems to a unified Databricks lakehouse while establishing governance patterns that enable multi-product operations at scale. What You'll Build Data Structuring: Design data models and implement unified schemas across multiple disparate product lines. Unity Catalog Architecture: Design and implement multi-catalog governance strategy supporting data isolation, cross-product data sharing, and comprehensive lineage tracking across our product portfolio Delta Lake Optimization: Establish patterns for Z-ordering, compaction, and liquid clustering at multi-TB scale. Define table structures, partitioning strategies, and retention policies that balance query performance with storage costs ETL Pipeline Framework: Build declarative pipeline patterns using Delta Live Tables. Create orchestration workflows for ingesting data from internal sources such as SQL databases and S3 Third Party Integrations: Integrate with third party data sources such as ERP systems (Netsuite etc.) and external data providers (S&P etc.) with automated ingest, robust error handling and monitoring. Platform Operations: Implement cost monitoring and optimization strategies, establish data quality frameworks, create self-service patterns enabling Data Engineers to work independently while maintaining governance standards Business Problems You'll Solve Key Legacy Product Migrations: Lead the architecture for migrating multi-terabyte datasets from legacy systems to Databricks—establishing patterns that will be reused across multiple product lines Multi-Product Data Architecture: Design Unity Catalog structures enabling secure data separation between product lines while allowing controlled cross-product analytics where appropriate Cost-Efficient Scale: Build infrastructure that scales efficiently—through intelligent caching, query optimization, and compute management strategies that avoid linear cost growth Platform Reliability: Establish monitoring, alerting, and data quality validation ensuring the platform operates reliably as foundation for both analytics and AI workloads Required Experience Databricks Expertise (Required) Unity Catalog: Production experience with multi-catalog governance, metastore design, and lineage tracking. Data Structuring: Experience designing and building unified schemas across multiple disparate product lines. Delta Lake: Expert-level experience with Z-ordering, compaction, liquid clustering, and performance tuning at multi-TB scale Delta Live Tables: Strong hands-on experience building declarative ETL pipelines, including change data capture and expectations/constraints Databricks Workflows: Experience with job orchestration, scheduling, and operational monitoring Business Intelligence: Experience enabling company-wide analytics and reporting with modern business intelligence tools and maintaining source of truth data and metrics. PySpark & Databricks SQL: Strong proficiency for code review, performance tuning, and query optimization Core Platform Engineering 5-8 years in data engineering or data platform roles, with 3+ years hands-on Databricks experience Track record leading at least one significant platform build or migration project AWS experience (S3, IAM, VPC) with ability to collaborate on infrastructure decisions Infrastructure-as-code experience (Terraform preferred) Technical Leadership Demonstrated ability architecting data platforms from first principles and defending technical decisions Strong written and verbal communication— document architecture decisions and present to both technical and business stakeholders Preferred But Not Required Experience with financial data, accounting systems (NetSuite), or enterprise ERP platforms Background building platforms that serve AI/ML workloads (experience preparing data for downstream ML consumption, RAG and retrieval, and LLMs. Understand advanced intelligence concepts such as relationship surfacing with knowledge graphs Familiarity with data governance frameworks and compliance requirements for regulated industries What We Offer:(The following only applies to US-based positions) A collaborative team culture with opportunities for career development. Ample opportunities to be recognized, build valuable skills, and grow your career. Generous vacation policy, including paid parental leave. Comprehensive health plans with FSA and HSA options. 401(k) retirement plan. Life and disability insurance coverage. Supplemental benefits like a dependent care savings plan, pet insurance, will preparation, and an employee assistance program. About Us: At Exactera, a FinTech SaaS start-up founded in 2016, we stand at the intersection of human and machine intelligence. Our corporate tax solutions are powered by AI and cloud-based technologies, serving customers worldwide. With over $100 million in funding from Savant Venture Fund and Insight Partners, we are poised for growth. We are committed to diversity, inclusion, and equal opportunities for all.

Technology

Exactera

Principal Product Manager (Agentic Platform)

Senior

On-site

San Diego, CA

🏢 Summary: Principal-level Product Manager role owning the strategy and roadmap for an AI-native, agentic platform that powers compliance-grade tax analysis on trusted data. The position defines how AI agents and practitioners access, use, and govern data, ensuring defensible, auditable outputs in regulated tax workflows. You will work closely with AI and Data Engineering leadership to shape platform capabilities, guardrails, and measurable practitioner outcomes. 🗂️ Requirements: 10+ years in product management, Experience owning technical platform, data, or ML products, Proven ownership of AI or ML platform from ambiguity to production, Strong understanding of lakehouse architectures and data contracts, Knowledge of agentic AI systems, retrieval, and tool calling, Experience designing guardrails, evaluation, and human-in-the-loop systems, Ability to define compliance-grade, auditable product requirements, Experience working cross-functionally without positional authority, Hands-on use of AI tools in daily work 📃 Skills: AI, ML, Lakehouse, Databricks, Delta, Retrieval, MCP, Data, Lineage, Evaluation, Guardrails, SaaS 🏢 Description: About the Role Exactera helps large multinationals handle complex tax work: transfer pricing, R&D tax credits, indirect tax, and audit defense. The company is moving from a software-and-services model to an AI-native platform that performs the work itself, with tax practitioners applying judgment where defensibility requires it. This role owns the product direction for the agentic platform behind that shift: the system that enables AI agents and practitioner tools to run tax analysis on trusted, current data. The data foundation (a lakehouse, external sources, and the contracts that keep them reliable) is the substrate. The product defines what agents and practitioners can do with it. You work directly with the Principal AI Engineer, who owns the AI/ML architecture, and the Principal Data Engineer, who owns the data tier. You own what the platform does, in what order, and why. Because you define how practitioners and agents use AI, you work deeply in these tools yourself. What You’ll Do The six outcomes that define success in the first 18 months: O1. Own the product strategy and roadmap for the platform. You own what the platform is for and what it builds next. The roadmap ties to practitioner outcomes rather than feature counts. When priorities compete across engineering, product, and practice teams, you decide and keep everyone aligned. O2. Define what compliance-grade means for this platform. Tax work must be defensible. You set product requirements that make platform output trustworthy: where results must be deterministic, what guardrails constrain agent behavior, how tool access is scoped, and what must be auditable. You determine where human judgment remains in the loop. O3. Own the agentic access surface. You own the interface AI agents and practitioner tools use to reach data and act on it, the contracts behind it, and the limits on agent capabilities. How this surface works determines how the entire platform is used. O4. Treat the data foundation as a product. With the Principal Data Engineer, you define contracts between data producers and consumers, including schema, quality, and freshness commitments practitioners can rely on. O5. Make external data a platform capability. You decide which regulatory, financial, and third-party data sources matter, how they appear in the product, and how to balance coverage, cost, and reliability. O6. Measure success by practitioner outcomes. Success is defined by faster, more accurate work with less manual effort. A production pipeline that does not improve a practitioner’s day is not successful. What We’re Looking For Required 10+ years in product management, with significant experience in technical platform, data, or ML products. Experience owning a platform or major product area at Principal-level scope across teams without positional authority. Ability to set direction across a modern data platform, including understanding of lakehouse and medallion patterns, data contracts, quality, and lineage. Track record of owning AI or ML platform initiatives from ambiguous problem definition to production. Strong understanding of agentic AI systems and data access patterns, including retrieval and tool calling. Experience designing AI systems where correctness and defensibility matter, including evaluation, guardrails, and human-in-the-loop workflows. Ability to define and measure platform success through practitioner results. Clear communication across engineering, product, and practice teams. Active use of AI tools in day-to-day work. Preferred Experience in tax, financial reporting, regulatory compliance, or other regulated, data-intensive domains. Experience building or governing data products, including external data sourcing and licensing trade-offs. Hands-on familiarity with Databricks, Delta Lake, or comparable lakehouse platforms in production. Background in tech-enabled services or SaaS-to-services transitions. Experience building a platform or major product area from an early stage. What We Offer A collaborative team culture with opportunities for career development. Ample opportunities to be recognized, build valuable skills, and grow your career. Generous vacation policy, including paid parental leave. Comprehensive health plans with FSA and HSA options. 401(k) retirement plan. Life and disability insurance coverage. Supplemental benefits such as a dependent care savings plan, pet insurance, will preparation, and an employee assistance program.

Technology

TechTree

Senior Data Platform Engineer

Senior

Remote

Krakow, Poland

208,000 - 312,000 PLN/yr

🏢 Summary: Senior Data Platform Engineer role focused on building and optimising a cloud-native lakehouse platform for large-scale analytics and reporting. The position involves designing distributed data pipelines, enabling self-service analytics, and implementing governance and observability frameworks using modern data technologies. You will work with Spark-based systems and integrated data warehousing solutions to deliver scalable, reliable data platforms. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production environments, Experience with Delta Lake and/or Apache Iceberg in production, Practical experience with dbt for data transformations, Experience with Databricks and Snowflake, Understanding of data governance and lineage in large-scale environments, Familiarity with Kubernetes and Docker, Experience with CI/CD and automated testing practices, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platforms you build will have real consequence across the organisation. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As a Senior Data Platform Engineer, you'll combine strong software engineering principles with deep data expertise to build robust, cloud-native platforms that process large-scale datasets efficiently and enable internal teams to build reporting and analytics on top of them. The role emphasises cloud-native architecture, lakehouse integration, data warehousing, and governance best practices. You'll work on systems using Apache Spark, Delta Lake, and Iceberg, and help deliver curated data models and self-service analytics capabilities to internal stakeholders. You'll also participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems Design and implement scalable data pipelines and distributed systems using Spark and Python to process and transform large-scale datasets for analytics and reporting. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Iceberg, ensuring data reliability, versioning, and performance optimisation at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Collaborate with internal teams to deliver curated datasets and self-service analytics capabilities for reporting and advanced use cases. Data warehousing optimisation Integrate and optimise Databricks and Snowflake for scalable storage and query performance. Drive performance tuning and cost optimisation across Spark jobs and cloud-native environments. Governance and observability Implement observability and governance frameworks including data lineage, quality checks, and compliance controls. Build platforms that allow secure and compliant access to diverse data sources. Engineering best practices Apply and champion clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in both Python and SQL, applied to production data platform work at scale. Apache Spark Solid hands-on experience with Spark for distributed data processing, including performance tuning in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You've applied these in production and understand the trade-offs in real-world scenarios. dbt and analytics tooling Practical experience with dbt for transformation workflows. Familiarity with Databricks and Snowflake for large-scale analytics workloads. Data governance and compliance Understanding of data governance, lineage tracking, and compliance requirements in large-scale, multi-tenant data environments. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 208,000 – 312,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

Nimble Robotics

Data Engineer II/III

Mid

On-site

San Francisco, CA

140,004 - 200,004 USD/yr

🏢 Summary: Data Engineer role focused on building and scaling data infrastructure for advanced robotics systems. The position involves designing reliable batch and real-time pipelines, optimizing ETL/ELT processes, and implementing data governance across cloud platforms. You will collaborate cross-functionally to deliver high-quality, scalable data solutions supporting company-wide analytics and operations. 🗂️ Requirements: BS/MS/PhD in Computer Science, Mathematics, Computer Engineering or related field, or equivalent practical experience, 1–5 years of professional experience, Proficiency in Python and SQL, Hands-on experience with Rust, Go, Java, C#, or C++, Experience with Kafka, Spark, and lakehouse formats (Icebert, DeltaLake), Experience with AWS, GCP, or Azure, Strong debugging and problem-solving skills, Willingness to work extended hours and weekends if needed, Ability to work full time onsite in San Francisco 📃 Skills: Python, SQL, Rust, Go, Java, C#, C++, Kafka, Spark, Icebert, DeltaLake, AWS, GCP, Azure, Clickhouse, Flink, Databricks 🏢 Description: About the Role Help us advance our robotics moonshot by scaling our data engineering efforts. Drive design and development of data infrastructure across our products and internal tools. You will play a critical role in working with a cross-functional team to architect and build advanced robotic systems. Responsibilities - Design, build, and maintain scalable, reliable data pipelines that support both batch and real-time analytics. - Develop and optimize ETL/ELT processes to ingest, transform, and integrate data from multiple sources into data warehouses and lakehouse platforms. - Drive and apply best practices in data modeling, pipeline architecture, query optimization, data quality, and data engineering standards. - Collaborate closely with finance, bizops, and engineering teams to understand data needs and deliver solutions. - Provide data engineering support to ensure data accessibility and usability company-wide. - Implement and maintain data governance frameworks to support compliance with internal policies and external regulatory requirements. - Evaluate and adopt modern data engineering technologies, tools, and best practices to improve scalability, reliability, and operational efficiency. - Document data engineering processes, designs, and architectures. Qualifications - BS/MS/PhD in Computer Science, Mathematics, Computer Engineering, or a related field, or equivalent practical experience. - 1-5 years of professional experience. - Proficiency in writing production-grade code in Python and SQL. - Hands-on experience with one of the following languages: Rust, Go, Java, C#, C++. - Strong debugging skills and the ability to diagnose and resolve issues efficiently. - Experience with Kafka, Spark and common lakehouse formats like Icebert, DeltaLake. - Experience working with cloud platforms like AWS, GCP, Azure. Preferred Experience - Experience with Clickhouse, Apache Flink, Databricks. - Experience on financial reporting and understanding of accountability. Additional Requirements - Willing to work extended hours and weekends if needed. - This position is based full time in our San Francisco headquarters. Compensation The pay range for this position at the start of employment is expected to be between $140,000 - $200,000/year. The exact offer may vary depending on job-related knowledge, skills, and experience. In addition to cash compensation, this position will also receive generous equity. Benefits - Unlimited Flexible Time Off. - Health insurance (medical, dental, vision). - Paid parental leave. - Commuter benefits including fully paid parking spots. - Referral bonus. - 401k retirement plan. - Equity program.

Technology

TechTree

Lead Data Engineer

Senior

Remote

Krakow, Poland

270,000 - 406,000 PLN/yr

🏢 Summary: Lead Data Engineer role focused on driving architecture and leading a team to build scalable, secure ETL/ELT pipelines and analytics-ready data models on modern cloud platforms. The position combines hands-on engineering with technical leadership, ensuring high standards in governance, observability, and performance optimisation. You will shape data infrastructure that supports large-scale analytics across the organisation. 🗂️ Requirements: Proven experience leading data engineering or analytics engineering teams, Strong programming skills in SQL, Strong programming skills in Python, Hands-on experience with Airflow or Prefect in production, Deep practical experience with dbt, Experience with Snowflake or Databricks at scale, Strong knowledge of dimensional modelling and SCD strategies, Experience implementing data quality and governance frameworks, Experience with CI/CD and automated testing in data systems 📃 Skills: SQL, Python, Airflow, Prefect, dbt, Snowflake, Databricks, ETL, ELT, CI/CD, SCD, Dimensional, Git, IaC 🏢 Description: ABOUT THE COMPANY Our client is a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you lead will have real consequence across the organisation. ABOUT THE ROLE We're looking for a Lead Data Engineer to combine deep technical expertise with hands-on team leadership, guiding a team of data engineers building and maintaining ETL/ELT pipelines, data models, and governance frameworks that power analytics and reporting across the organisation. This is a technical leadership role — you'll drive architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable data solutions. You'll collaborate closely with stakeholders to align technical work with business objectives, champion governance and observability standards, and foster a culture of continuous improvement. The expectation is that you're equally effective in an architecture review as you are pairing with an engineer on a tricky pipeline problem. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data engineers, promoting collaboration, knowledge sharing, and professional growth. Set the standard for engineering quality and hold the bar consistently. Architecture and pipeline design Drive architectural decisions for ETL/ELT pipelines, orchestration frameworks (Airflow/Prefect), and transformation layers (dbt). Facilitate architecture reviews and contribute to design decisions for scalable, fault-tolerant systems. Analytics data modelling Oversee design and implementation of analytics-ready data models — dimensional schemas, SCD strategies, and semantic layers — that internal teams can build on reliably. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. Platform optimisation Manage performance tuning and cost optimisation for Snowflake, Databricks, and related cloud data platforms at scale. Governance and observability Champion governance, observability, and compliance frameworks across all data workflows — including data quality, lineage tracking, and multi-tenant environment controls. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure timely delivery aligned with business objectives. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or analytics-focused development teams — mentoring engineers, driving architectural decisions, and owning delivery outcomes. SQL and Python Strong programming skills in both SQL and Python, applied to production data systems at scale. ETL/ELT orchestration Hands-on experience with orchestration tools — Airflow and/or Prefect — in production pipeline environments. dbt expertise Deep practical experience with dbt for transformation workflows and analytics modelling, including testing, documentation, and modular project design. Snowflake and Databricks Familiarity with Snowflake and/or Databricks for large-scale data processing, including performance tuning and cost management. Data modelling principles Solid understanding of data modelling principles, incremental strategies, and schema design for analytics — dimensional modelling, SCDs, and semantic layer design. Governance and data quality Knowledge of data quality frameworks, lineage tracking, and governance in multi-tenant environments. Software engineering practices Familiarity with CI/CD, automated testing, and infrastructure-as-code practices applied to data systems. Communication and stakeholder management Strong communication skills with the ability to operate confidently across technical teams and business stakeholders. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements, hybrid model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

TechTree

Lead Distributed Data Platform Engineer

Senior

Remote

Warsaw, Poland

270,000 - 406,000 PLN/yr

🏢 Summary: Lead Distributed Data Platform Engineer responsible for architecting and delivering enterprise-scale lakehouse and distributed data platforms to enable advanced analytics and reporting. The role combines hands-on technical leadership with team mentorship, driving scalable, secure, and high-performance data solutions in cloud-native environments. You will guide architectural decisions, enforce engineering best practices, and ensure platform reliability at scale. 🗂️ Requirements: Proven experience leading data engineering or platform teams, Strong programming skills in Python, Strong programming skills in SQL, Hands-on experience with Apache Spark in production, Experience with Delta Lake and/or Apache Iceberg in production, Experience designing distributed systems and lakehouse architectures, Experience building scalable data pipelines, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Experience working in cloud-native environments 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD, Java, Scala, Rust 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the platform you lead will underpin how the entire organisation accesses and acts on its data. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As Lead Distributed Data Platform Engineer, you'll combine deep technical expertise with hands-on team leadership — guiding a team in designing and maintaining data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services at enterprise scale. You'll lead architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable solutions. The role emphasises technical leadership, governance best practices, and a culture of innovation and continuous improvement. You'll also participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data platform engineers, promoting collaboration, knowledge sharing, and professional growth. Set and maintain high engineering standards across the team. Distributed systems architecture Drive architectural decisions for distributed systems and lakehouse platforms using Spark, Delta Lake, and Iceberg. Facilitate architecture reviews and contribute to design decisions for fault-tolerant, future-ready systems. Data pipeline and platform delivery Oversee design and implementation of scalable data pipelines and analytics workflows, ensuring they are reliable, performant, and maintainable at scale. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all platform engineering work. Performance and cost optimisation Manage performance tuning, scalability strategies, and cost optimisation across cloud-native environments and large-scale distributed workloads. Governance and observability Champion governance, observability, and compliance frameworks across all data platforms — ensuring data remains accessible, secure, and auditable. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure delivery aligns with business objectives and analytics needs. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or platform development teams — mentoring engineers, owning architectural decisions, and driving delivery outcomes. Python and SQL Strong programming skills in both Python and SQL applied to production data platform work at scale. Apache Spark Hands-on experience with Spark for distributed data processing, including performance tuning and optimisation in production environments. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the trade-offs and have applied these technologies in production at scale. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and large-scale data processing. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data platform systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Communication and stakeholder management Strong communication skills with the confidence to operate across engineering teams, cross-functional partners, and senior leadership. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. Exposure to service mesh and advanced orchestration patterns. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

TechTree

Advanced Data Platform Engineer

Senior

Remote

Krakow, MA, Poland

160,000 - 240,000 PLN/yr

🏢 Summary: The role focuses on designing and building scalable, cloud-native data platforms to enable advanced analytics and reporting across a large internal data ecosystem. It involves developing distributed data pipelines, lakehouse architectures, and optimised data warehousing solutions with strong emphasis on performance, governance, and reliability. The position requires deep technical expertise in big data technologies and cloud-native infrastructure, including on-call responsibility for platform stability. 🗂️ Requirements: Strong programming skills in Python, Strong programming skills in SQL, Commercial experience with Apache Spark for distributed data processing, Hands-on experience with Delta Lake and/or Apache Iceberg in production, Experience with dbt for SQL transformations, Experience with Databricks and Snowflake, Knowledge of CI/CD and automated testing practices, Experience with Kubernetes and Docker, Understanding of performance tuning and cost optimisation in large-scale data systems, Ability to participate in on-call rotations 📃 Skills: Python, SQL, Spark, Delta, Iceberg, dbt, Databricks, Snowflake, Kubernetes, Docker, CI/CD 🏢 Description: ABOUT THE COMPANY We are a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. Over 75% of our business has transitioned to our cloud platform, and we are making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you build will have real consequence. ABOUT THE ROLE We're building a specialised team focused on enabling advanced analytics and reporting capabilities across our internal data ecosystem. As an Advanced Data Platform Engineer, you'll design and implement scalable, cloud-native data platforms that integrate modern lakehouse technologies, distributed compute frameworks, and cloud-native services to support diverse analytical use cases at enterprise scale. The role emphasises technical depth — performance optimisation, governance best practices, and the kind of engineering rigour that keeps vast datasets accessible, secure, and compliant. You'll work closely with internal teams to deliver curated datasets and self-service analytics capabilities, and you'll participate in on-call rotations as part of shared team responsibility. WHAT YOU'LL WORK ON Data pipeline and distributed systems design Design and implement complex data pipelines and distributed systems using Spark and Python, applying clean code principles, modular design, CI/CD, automated testing, and thorough code reviews. Lakehouse platform development Develop and maintain lakehouse capabilities with Delta Lake and Apache Iceberg, ensuring reliability, performance, and long-term maintainability at scale. Analytics workflow enablement Integrate dbt for SQL transformations running on Spark. Deliver curated datasets and self-service analytics capabilities that empower internal stakeholders to explore data independently. Data warehousing optimisation Optimise Databricks and Snowflake environments for performance and scalability. Drive cost optimisation and performance tuning across Spark jobs and cloud-native infrastructure. Observability and governance Implement observability and governance frameworks including data lineage tracking and compliance controls, ensuring data remains secure and auditable. On-call participation Participate in on-call rotations as part of shared team responsibility for platform reliability. WHAT WE LOOK FOR Python and SQL Strong programming skills in Python and SQL — the foundation for everything you'll build here. Apache Spark Solid experience with Spark for distributed data processing at scale, including performance tuning and optimisation. Lakehouse architecture Expertise in Delta Lake and/or Apache Iceberg. You understand the tradeoffs and have used these in production environments. Analytics tooling Familiarity with dbt, Databricks, and Snowflake for analytics workflows and SQL transformation pipelines. Software engineering fundamentals Solid understanding of software engineering principles — CI/CD, automated testing, clean code, and modular design applied to data systems. Infrastructure and containerisation Familiarity with Kubernetes, Docker, and infrastructure-as-code tools in cloud-native environments. Scalability and cost optimisation Understanding of performance tuning, scalability strategies, and cost optimisation for large-scale data systems. Bonus Exposure to event-driven architectures and advanced analytics platforms. Experience enabling self-service analytics for internal stakeholders. Experience in Java, Scala, or Rust. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 160,000 – 240,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements with a remote-first model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.

Technology

Xometry

Staff Data Engineer

Senior

On-site

North Bethesda, MD

180,000 - 200,004 USD/yr

🏢 Summary: Senior individual contributor role responsible for designing and owning enterprise-scale data architecture and real-time data pipelines that power a strategic DFM AI + IQE partner integration. The position focuses on building scalable batch and streaming systems, defining data models, and ensuring governance, observability, and CI/CD standards across cross-system integrations. The engineer leads the digital data plane connecting internal platforms with external PLM ecosystems in a high-impact, cloud-native environment. 🗂️ Requirements: Bachelor’s degree in STEM or equivalent experience, Minimum 5 years of experience in data engineering, Deep expertise in Snowflake or similar cloud data warehouse, Expert-level SQL, Strong Python proficiency, Hands-on experience with modern data pipeline tools (dbt, Airbyte, Airflow or similar), Experience designing enterprise data architecture across multiple systems, Knowledge of batch and stream processing systems, Experience with highly scalable data stores, Experience with CI/CD, automated testing, contract testing, schema evolution, Strong knowledge of AWS data ecosystem, Experience integrating with enterprise or partner systems (e.g., PLM, ERP, SaaS) 📃 Skills: Snowflake, SQL, Python, dbt, Airbyte, Airflow, Kafka, Spark, Kinesis, Apache, Iceberg, AWS, Teamcenter, BMIDE, APIs, Looker, Streamlit, Terraform, CloudFormation, CI/CD, CDC 🏢 Description: Xometry is looking for a Staff Data Engineer to join the Data Platform team. This is a senior individual contributor role with broad technical scope and high organizational impact. You will own data architecture decisions, lead the design of scalable pipelines and platforms, and set the engineering bar for how data systems are built and operated. A defining piece of this role is owning the data architecture behind the DFM AI + IQE integration with a strategic partner. You will serve as the data engineering lead for the digital thread connecting the platform to partner ecosystems including Solid Edge, NX, Designcenter, and Teamcenter. You will build the pipelines, contracts, and observability that move quotes, parts, manufacturability signals, and pricing between systems in real time. Responsibilities Lead with technical depth – Design and drive the implementation of enterprise-scale data architecture and engineering solutions spanning multiple systems and domains. Own the partner integration data plane – Architect and build the data layer of the embedded DFM AI + IQE integration with Teamcenter and Designcenter. Own bidirectional pipelines, the joint data model for parts, BOMs, quotes, and manufacturability signals, low-latency feedback paths, and required governance, lineage, and audit controls. Build for scale – Architect and optimize reliable batch and streaming data pipelines, data models, and platforms handling complex, high-volume and event-driven data flows. Own the full lifecycle – Take end-to-end accountability from data acquisition and transformation through delivery, observability, and performance. Set the standard – Define and enforce best practices for data modeling, CI/CD, testing, code quality, contract testing, and schema evolution. Solve ambiguous problems – Navigate cross-domain technical challenges and deliver solutions meeting business and technical objectives. Develop multi-quarter roadmaps – Translate strategic priorities into technical plans and timelines. Collaborate broadly – Partner with engineering, product, data science, business stakeholders, and external partner engineering teams. Mentor and elevate – Guide engineers through design reviews, code reviews, and mentorship. Evaluate and adopt – Recommend tools, platforms, and architectural patterns within the data engineering ecosystem. Qualifications Bachelor's degree in a STEM field (or equivalent experience) and at least 5 years of experience in data engineering with ownership of large-scale data systems. Deep expertise with cloud data warehouses, preferably Snowflake, including optimization and performance tuning. Expert-level SQL and strong Python proficiency. Experience building and optimizing data pipelines and architectures using tools such as dbt, Airbyte, or Airflow. Experience planning and implementing enterprise data architecture across multiple systems and organizational boundaries. Working knowledge of queueing, batch and stream processing (Kafka, Spark, Kinesis) and scalable data stores (Apache Iceberg). Experience developing database-heavy services or APIs with focus on testability and maintainability. Strong understanding of CI/CD, automated testing, contract testing, and schema evolution in data pipelines. Strong knowledge of AWS data ecosystem and cloud-native infrastructure. Enterprise integration experience with PLM, ERP, or large SaaS systems; Teamcenter experience is a strong plus. Familiarity with data visualization tools such as Looker or Streamlit. Experience with data governance, data quality frameworks, and observability tooling. Exposure to lakehouse or data mesh architectures. Experience with infrastructure as code frameworks such as Terraform or CloudFormation. Experience with event-driven architecture, CDC pipelines, and low-latency operational data flows. Benefits Base salary range: $180,000–$200,000 annually plus commission, depending on experience and location. Competitive benefits package including 401(k) match, medical, dental, and vision insurance; life and disability insurance; generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave; employee assistance program and additional wellbeing resources.

Technology

SugarCRM

Senior Data Engineer - Databricks

Senior

Hybrid

Denver, CO

155,004 - 185,004 USD

🏢 Summary: Senior Data Engineer role focused on building and optimizing Databricks-based data pipelines for the Sugar Predict platform, integrating ERP and CRM data across Azure and AWS environments. The position involves production support, legacy ETL modernization, multi-tenant data architecture, and collaboration with ML and product teams to deliver scalable revenue intelligence solutions. Hybrid work arrangement with in-office collaboration in Denver, CO. 🗂️ Requirements: 4+ years of data engineering experience, 2+ years of Databricks or Apache Spark experience, Proficiency in PySpark, Proficiency in SQL, Proficiency in Python, Experience building production-grade data pipelines, Hands-on experience with Delta Lake, Experience with pipeline performance tuning, Knowledge of PostgreSQL, Experience maintaining legacy ETL tooling, Experience with multi-tenant architectures, Understanding of data governance and security principles, Ability to support on-call incident response, Experience with CI/CD practices 📃 Skills: Databricks, Spark, PySpark, SQL, Python, DeltaLake, PostgreSQL, SSIS, Informatica, Azure, AWS, ETL, ELT, CI/CD, Serverless, SQLServer, Feast 🏢 Description: About SugarAI SugarAI is redefining CRM for the age of AI. We help teams turn fragmented customer and revenue signals into clear, prioritized action through intelligent CRM solutions. Where You Fit In: The Sugar Predict platform powers revenue intelligence for mid-market enterprises by fusing ERP and CRM data into actionable insights. As a Senior Data Engineer, you will own the Databricks pipelines that make this possible, driving production reliability, cost efficiency, and platform growth through customer onboarding and legacy modernization. You will work closely with ML engineers, product teams, and the Enterprise Architecture team to ensure the data backbone behind Sugar Predict is always fast, clean, and ready to deliver at a global scale. This role operates on a hybrid model, with a mix of remote work and in-office collaboration at our Denver, CO location, specifically, working in-office a minimum of 2-3 days per week. Impact You Will Make in the Role: • Own Databricks production support for the Sugar Predict data platform, including monitoring, alerting, and incident response across all production data flows • Maintain and report on SLA performance metrics for data pipeline delivery • Identify and implement pipeline optimizations that reduce Databricks compute costs and improve throughput • Migrate legacy ETL/ELT pipelines to Databricks and build automation tooling • Support new customer onboarding by provisioning and validating tenant data pipelines • Design and build high-performance Databricks pipelines across Azure and AWS environments • Own Delta Lake architecture including schema design and incremental processing patterns • Enforce data security best practices including role-based access control and secrets management • Implement data quality monitoring and observability across pipeline health and ML model inputs • Apply and enforce multi-tenant data isolation patterns • Partner with the Enterprise Architecture team on platform integration • Support globally distributed operations through on-call rotation and incident response • Maintain technical documentation, runbooks, and architectural decision records • Apply CI/CD best practices including version control, automated testing, and deployment tooling What You Will Bring: • 4+ years of data engineering experience • At least 2 years on Databricks or Apache Spark across Azure and/or AWS • Proficiency in PySpark, SQL, and Python • Hands-on experience with Delta Lake including schema evolution and ACID transactions • Experience with pipeline performance tuning and compute optimization • Solid working knowledge of PostgreSQL • Experience supporting legacy ETL tooling such as SSIS, Informatica, or custom Python/SQL pipelines • Experience supporting large-scale multi-tenant architectures • Proven ability to collaborate across data science, product, and infrastructure teams • Strong understanding of data governance, security, and compliance principles Preferred Qualifications/Experience: • Experience operating Databricks workspaces across Azure and AWS • Experience optimizing Databricks workloads in a Serverless environment • Experience with Microsoft SQL Server • Exposure to ML feature engineering or feature stores • Experience with customer onboarding automation or IaC patterns • Databricks Certified Data Engineer Associate or Professional certification Benefits and Perks: • Excellent healthcare package for you and your family • 401(k) match • Unlimited Paid Time Off • Paid Parental Leave • Online Legal Services • Financial Planning Services • Discounted Pet Insurance • Corporate Benefit Program with travel and entertainment discounts • Health and Wellness Reimbursement Program • Travel Discounts • Educational Resources and Career Development Program • Employee Referral Bonus Program • Merit-based career growth opportunities Our company uses E-Verify to confirm the employment eligibility of all newly hired employees.

Technology

Air

Senior Data Platform Engineer

Senior

On-site

Pittsburgh, PA

🏢 Summary: Full-time Data Platform Engineer role focused on designing and scaling cloud-based data platforms that power AI/ML and analytics solutions for government market intelligence. The position involves building and maintaining data pipelines, enhancing platform architecture, and collaborating cross-functionally to deliver secure, reliable, and high-quality data solutions. Based in Pittsburgh with limited travel, this role requires strong expertise in Python, SQL, big data frameworks, and cloud-native architectures. 🗂️ Requirements: U.S. Citizenship, Bachelor's degree in Computer Science, Mathematics or related technical field, 5+ years experience building and maintaining scalable data platforms or services, Strong proficiency in Python, Strong proficiency in SQL, Experience architecting and implementing data pipeline frameworks including orchestration, Understanding of data security and governance best practices, Hands-on experience with big data frameworks, Knowledge of cloud-native data lake or lakehouse architectures, Ability to support AI, machine learning, data science, and web application data needs 📃 Skills: Python, SQL, Apache, Spark, AWS, AI, MachineLearning, DataScience, DataPipelines, Cloud, DataLake, Lakehouse 🏢 Description: Company Description Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. Today, a dangerous Readiness Gap exists between what the front line needs and what is delivered. Our AI-native platform, Air Enterprise Readiness, aligns development, production, delivery, and sustainment into one coordinated execution system for government agencies and industrial suppliers. By revealing true capacity, exposing real constraints, coordinating resources, and executing at the speed of operational demands, the front line gets what it needs to succeed. Job Description We are seeking an exceptional and experienced data platform engineer who shares our passion and obsession with quality. You'll be a core member of our product and engineering team dedicated to helping our clients replace time-consuming, manual processes to reach informed, real-time decisions about government markets, competitors, and agency relationships. We need a skilled and dedicated engineer to join our team to lead us in uncovering truth and meaning in data. You must be hands-on with a strong understanding of both data platforms and cloud infrastructure. You must also have strong interpersonal skills to work cross-functionally across internal teams as well as directly with end users and Air platform SMEs. In order to do this job well, you must be a curious and eager problem solver with a hunger for delivering high-quality data solutions. You have a passion for great work and nothing less than your best will do. You share our intolerance of mediocrity and produce simple solutions to complex problems. Knowing there are always multiple answers to a problem, you know how to engage in a constructive dialogue to find the best path forward. This role is a full-time position located out of our office in Pittsburgh, PA. This role may require up to 10% travel Scope of Responsibilities Design and implement platform architecture enhancements to support scalability and reliability Design, build, maintain, and support data platform solutions to enable efficient data consumption across functional teams, including AI/ML, Data Science, and Software Evaluate and integrate emerging technologies to improve the data platform Collaborate with teams across the organization to understand business needs and translate them into data solutions Mentor junior engineers by providing code reviews, guidance, and technical leadership through promotion of engineering best practices Qualifications U.S. Citizenship is required Bachelor's degree in Computer Science, Mathematics or a related technical field Required Skills: 5+ years of experience in building and maintaining scalable data platforms or services Strong proficiency in Python and SQL Proven experience in architecting and implementing data pipeline frameworks, including orchestration Solid understanding of data security and governance best practices Hands-on with big data frameworks, such as Apache Spark Knowledge of cloud-native data lake/lakehouse architectures Demonstrated ability to collaborate with stakeholders to deliver data platform solutions that support AI, machine learning, data science, and web applications Strong analytical skills and attention to detail Ability to work independently in a fast-paced environment A passion for solving complex problems and building sustainable, scalable systems Desired Skills: Current possession of a U.S. security clearance, or the ability to obtain one with our sponsorship Experience using Amazon Web Services Experience in or exposure to the nuances of a startup or other entrepreneurial environment Working knowledge of large (multiple terabytes) amounts of data We firmly believe that past performance is the best indicator of future performance. If you thrive while building solutions to complex problems, are a self-starter, and are passionate about making an impact in global security, we're eager to hear from you. We firmly believe that past performance is the best indicator of future performance. If you thrive while building solutions to complex problems, are a self-starter, and are passionate about making an impact in global security, we're eager to hear from you. Air is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans status or any other characteristic protected by law.