May 9, 2026
Data Engineer, TS/SCI
Senior • On-site
Reston, VA
Job Description
Seeking a Lead Data Engineer / Mission Data Pipeline to serve as a Subject Matter Expert (SME). You will work directly with government, technical, and industry stakeholders design, implement, and sustain data pipelines that ingest, transform, store, and distribute mission-relevant data across the DIFC2 software ecosystem and external mission partner systems/ You will support the development and delivery of software capabilities that improve data accessibility, mission integration, operational visibility, and cross-system interoperability across a highly dynamic mission environment.
The focuses on enabling reliable, automated, and secure data flows that support mission capabilities on operational timelines. This position requires collaboration with mission system stakeholders, platform engineers, and security teams to define data exchange artifacts, develop pipeline implementations, and ensure interfaces comply with Risk Management Framework (RMF) requirements. The engineer will leverage both low-code/no-code ETL technologies and custom software development to implement scalable data pipelines that support the data architecture.
Typical job responsibilities Include:
- Mission-Relevant Data Artifact Identification:
- Analyze mission workflows and system architecture to identify the data artifacts required to support operational capabilities.
- Define the structure, semantics, and lifecycle of mission data products within the data architecture.
- Ensure data artifacts align with architectural standards for interoperability, traceability, and reuse across mission components.
- Coordinate with External Stakeholders:
- Serve as the primary technical liaison for defining data exchange interfaces between mission systems.
- Collaborate with mission partners to define message schemas, data formats, transport mechanisms, and interface expectations.
- Document data interface specifications to support both system integration and operational sustainment.
- RMF & Security Documentation Support:
- Support the RMF process by defining technical details for machine-to-machine data interfaces.
- Produce or contribute to required RMF artifacts including:
- Data message descriptions
- Ports, Protocols, and Services Management (PPSM) entries
- System topology diagrams and interface documentation
- Coordinate with cybersecurity engineers to ensure pipeline implementations meet DoW security and compliance requirements.
- Data Pipeline Development & Sustainment:
- Design and implement automated data pipelines that ingest, transform, persist, and expose data artifacts for use by internal and external stakeholders.
- Utilize government-provided data platform infrastructure to build scalable ETL workflows.
- Develop custom transformation logic when required using appropriate programming languages and data processing frameworks.
- Expose processed data products through standardized service interfaces such as REST APIs or platform-native services.
- Data Pipeline Development & Sustainment:
- Implement data validation, normalization, and transformation logic to ensure accuracy and usability of integrated datasets.
- Troubleshoot data pipeline failures and implement monitoring, logging, and recovery mechanisms.
- Optimize pipeline performance to support mission timelines and large-scale data processing requirements.
- Deliverables & Key Projects
- Design and implementation of production data pipelines enabling automated ingestion and dissemination of mission data across systems.
- Technical documentation describing data artifacts, including purpose and operational relevance, source systems and ingestion methods, transformation logic, and destination systems and access interfaces
- Interface documentation supporting integration with external mission partners.
- RMF technical artifacts supporting security authorization of data interfaces.
- Continuous improvement and sustainment of the data pipeline ecosystem to support evolving mission needs.
Qualifications
- 9 years of experience and a Bachelor's degree in Computer Science, Software Engineering, Management Information Systems/Information Systems, or a related discipline; or a Master's degree and 7 years of experience; or a PhD/JD and 4 years of experience.
- 7+ years of experience supporting software engineering, systems integration, platform implementation, or application development efforts within DoD, IC, SAP/SCI, or other complex mission environments.
- Experience designing or implementing data pipelines, ETL workflows, or data integration architectures.
- Experience working with structured and semi-structured data formats (JSON, CSV, XML, etc.).
- Requires an active Top Secret clearance with the ability to obtain and maintain Sensitive Compartmented Information and Special Program access, as well as a willingness to consent to a polygraph examination.
You will wow us even more if you have experience will the following:
- Proficiency with low-code/no-code data orchestration technologies such as Apache NiFi, Apache Airflow, or similar ETL orchestration tools.
- Experience implementing data transformation logic using Python and data processing libraries such as Pandas and NumPy.
- Certification or demonstrated experience as a Palantir Foundry Data Engineer, including use of Code Repositories, Pipeline Builder, AIP Logic, and other platform-relevant technologies
- Experience designing pipelines that process and fuse large-scale datasets from multiple sources.
- Familiarity with DoW data architectures, mission system integration, or data interoperability frameworks.
- Ability to communicate effectively with engineers, architects, operators, and senior Government leadership while translating technical concepts into mission-relevant value.
- Proven ability to work in highly regulated, fast-moving environments where technical excellence, mission responsiveness, and stakeholder coordination are all critical to success.
Blue Sky Innovators, Inc. is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to sex, race, color, religion, national origin, disability, protected Veteran status, age, or any other characteristic protected by law. If you are a qualified job seeker with a disability or a disabled veteran, you have the right to request an accommodation if you are unable or limited in your ability to use or access http://www.blueskyinnovators.com as a result of your disability. To request an accommodation, please email us at careers@blueskyinnovators.com and provide your name and contact information. Please note: this is only for job seekers with disabilities requesting an accommodation.
Similar jobs you might like
Technology
SoftBlue
Data Engineer (Python & AWS)
Senior
Remote
Bydgoszcz, Poland
150 - 180 PLN
🏢 Summary: Senior Data Engineer role focused on designing and building scalable serverless data ingestion pipelines in AWS within the healthcare domain. The position emphasizes strong Python engineering, cloud architecture leadership, and implementation of modern data platforms and DevOps practices. The role involves driving technical excellence and delivering reliable, high-impact data solutions in an international environment. 🗂️ Requirements: 10+ years of experience in Python programming, Strong software engineering skills in data processing, Extensive experience with AWS Cloud and Serverless Architecture, Hands-on experience with AWS Lambda, S3, and Cognito, Experience building E2E automated tests for data pipelines, Practical knowledge of Data Mesh and Medallion Architecture, Experience with Infrastructure as Code using AWS CDK or Terraform, Experience with CI/CD pipelines using GitLab or GitHub Actions, Experience with ETL/ELT processes and dbt, Experience working with GraphQL, Minimum B2 level English proficiency 📃 Skills: Python, AWS, Lambda, S3, Cognito, Boto3, DataMesh, Medallion, CDK, Terraform, GitLab, GitHubActions, ETL, ELT, dbt, GraphQL, CI/CD 🏢 Description: We are looking for a highly skilled Data Engineer to join our client in the healthcare sector. Our requirements: Technical Expertise: Python Programming: 10+ years of experience with strong Software Engineering skills focused on data processing. AWS & Serverless: Extensive experience with AWS Cloud, specifically focusing on Serverless Architecture and services (including AWS Lambda , AWS S3 Tables , and AWS Cognito ). Automated Testing: Proven experience in developing End-to-End (E2E) automated tests to ensure pipeline reliability, utilizing tools such as Boto3 for AWS resource validation. Data Concepts: Practical knowledge of Data Mesh and Medallion Architecture , along with general data processing and analysis. DevOps & IaC: Hands-on experience with Infrastructure as Code ( AWS CDK or Terraform ) and CI/CD pipelines ( GitLab pipelines or GitHub Actions ). Modern Tooling: Experience with ETL/ELT solutions, dbt , and GraphQL . Communication & Soft Skills: English Language: Minimum B2 level , enabling smooth daily technical and business communication in a global environment. Collaboration: Excellent communication skills and the ability to thrive in a collaborative, international team. Standards: A strong commitment to high standards of ethics, quality (Clean Code), and reliable delivery. Nice to have: Experience with Snowflake and SQL . Knowledge of Data Vault 2.0 modeling. Experience with Databricks . Familiarity with the Microsoft ecosystem: C# / .Net, T-SQL, SQL Server , and Azure DevOps . Experience with Star Schema database modeling. Knowledge of Descriptive Statistics. Your responsibilites: Design and build scalable Data Ingestion pipelines within the AWS cloud ecosystem. Lead technical delivery and implementation of core platform components, ensuring architectural integrity across the entire data lifecycle. Collaborate with Engineering Managers and cross-functional teams across the globe and Poland. Drive technical excellence by improving team processes, architecture standards, and engineering best practices. Support and consult with stakeholders to ensure successful delivery of high-impact, data-driven solutions. Contribute to the growth and maturity of the team’s cloud and data engineering capabilities. We offer: Challenging role within the company that creates innovative solutions. Work in international environment on demanding projects. Remote work model. Subsidized private medical care, life insurance, multisport card. Integration meetings. Employee referral program. If you have a deep expertise in Python and AWS , and building scalable Serverless data architectures is where you truly excel, this is the perfect role for you!
Technology

Xometry
Staff Data Engineer
Senior
On-site
Waltham, MA
180,000 - 200,004 USD/yr
🏢 Summary: Senior individual contributor role leading enterprise-scale data architecture and real-time partner integrations, owning the design of scalable batch and streaming pipelines across systems. Responsible for building and operating the data plane behind a strategic DFM AI + IQE integration, enabling low-latency, bidirectional data flows between platforms. Sets engineering standards for data modeling, CI/CD, governance, and observability while collaborating cross-functionally. 🗂️ Requirements: Bachelor's degree in STEM or equivalent experience, 5+ years in data engineering with ownership of large-scale data systems, Deep expertise in Snowflake and cloud data warehouses, Expert-level SQL, Strong Python proficiency, Experience building and optimizing modern data pipelines (dbt, Airbyte, Airflow or similar), Experience designing enterprise data architecture across multiple systems and partner boundaries, Knowledge of batch and stream processing systems, Experience with highly scalable data stores, Experience writing database-heavy services or APIs, Strong understanding of CI/CD, automated testing, contract testing, schema evolution, Strong knowledge of AWS and cloud-native infrastructure, Enterprise or partner system integration experience (PLM, ERP, or SaaS), Experience with infrastructure as code frameworks, Experience with event-driven architectures and CDC pipelines 📃 Skills: Snowflake, SQL, Python, dbt, Airbyte, Airflow, Kafka, Spark, Kinesis, Apache, Iceberg, AWS, Terraform, CloudFormation, Teamcenter, BMIDE, APIs, CI/CD, CDC, Looker, Streamlit 🏢 Description: Xometry is looking for a Staff Data Engineer to join the Data Platform team. This is a senior individual contributor role with broad technical scope and high organizational impact. You will own data architecture decisions, lead the design of scalable pipelines and platforms, and set the engineering bar for how data systems are built and operated. A defining piece of this role is owning the data architecture behind the DFM AI + IQE integration with a strategic partner. You will serve as the data engineering lead for the digital thread connecting the platform to partner ecosystems including Solid Edge, NX, Designcenter, and Teamcenter, building pipelines, data contracts, and observability to move quotes, parts, manufacturability signals, and pricing data in real time. Responsibilities - Lead the design and implementation of enterprise-scale data architecture and engineering solutions across multiple systems and domains - Architect and build the data layer for embedded DFM AI + IQE integrations, including bidirectional pipelines and joint data models for parts, BOMs, quotes, and manufacturability signals - Design low-latency signal paths delivering DFM and pricing feedback into designer environments - Establish governance, lineage, and audit capabilities for partner-integrated data systems - Architect and optimize reliable batch and streaming pipelines for complex, high-volume, event-driven data flows - Own the full lifecycle of data engineering work from ingestion and transformation to delivery and observability - Define and enforce best practices for data modeling, CI/CD, testing, code quality, contract testing, and schema evolution - Solve complex cross-domain technical challenges aligned with business objectives - Develop multi-quarter technical roadmaps and execution plans - Collaborate with engineering, product, data science, business stakeholders, and partner engineering teams - Mentor engineers through design and code reviews - Evaluate and recommend tools, platforms, and architectural patterns Qualifications - Bachelor's degree in a STEM field (or equivalent experience) - At least 5 years of experience in data engineering with ownership of complex, large-scale systems - Deep expertise in Snowflake, including optimization and performance tuning - Expert-level SQL and strong Python proficiency - Experience with modern data tooling such as dbt, Airbyte, and Airflow - Experience designing enterprise data architectures spanning multiple systems and partner boundaries - Knowledge of batch and stream processing technologies (e.g., Kafka, Spark, Kinesis) and scalable data stores (e.g., Apache Iceberg) - Experience building database-heavy services or APIs with focus on testability and maintainability - Strong understanding of CI/CD, automated testing, contract testing, and schema evolution in data pipelines - Strong knowledge of AWS and cloud-native infrastructure - Enterprise integration experience with PLM, ERP, or large SaaS systems; Teamcenter experience is a strong plus - Familiarity with data visualization tools such as Looker or Streamlit - Experience with data governance, data quality frameworks, and observability tooling - Exposure to lakehouse or data mesh architectures - Experience with infrastructure as code (Terraform, CloudFormation) - Experience with event-driven architectures, CDC pipelines, and low-latency operational data flows Benefits - Estimated base salary range: $180,000–$200,000 annually plus commission, depending on experience and location - Competitive benefits package including 401(k) match - Medical, dental, and vision insurance - Life and disability insurance - Generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave - Employee assistance and wellbeing resources
Technology

Xometry
Staff Data Engineer
Senior
On-site
North Bethesda, MD
180,000 - 200,004 USD/yr
🏢 Summary: Senior individual contributor role responsible for designing and owning enterprise-scale data architecture and real-time data pipelines that power a strategic DFM AI + IQE partner integration. The position focuses on building scalable batch and streaming systems, defining data models, and ensuring governance, observability, and CI/CD standards across cross-system integrations. The engineer leads the digital data plane connecting internal platforms with external PLM ecosystems in a high-impact, cloud-native environment. 🗂️ Requirements: Bachelor’s degree in STEM or equivalent experience, Minimum 5 years of experience in data engineering, Deep expertise in Snowflake or similar cloud data warehouse, Expert-level SQL, Strong Python proficiency, Hands-on experience with modern data pipeline tools (dbt, Airbyte, Airflow or similar), Experience designing enterprise data architecture across multiple systems, Knowledge of batch and stream processing systems, Experience with highly scalable data stores, Experience with CI/CD, automated testing, contract testing, schema evolution, Strong knowledge of AWS data ecosystem, Experience integrating with enterprise or partner systems (e.g., PLM, ERP, SaaS) 📃 Skills: Snowflake, SQL, Python, dbt, Airbyte, Airflow, Kafka, Spark, Kinesis, Apache, Iceberg, AWS, Teamcenter, BMIDE, APIs, Looker, Streamlit, Terraform, CloudFormation, CI/CD, CDC 🏢 Description: Xometry is looking for a Staff Data Engineer to join the Data Platform team. This is a senior individual contributor role with broad technical scope and high organizational impact. You will own data architecture decisions, lead the design of scalable pipelines and platforms, and set the engineering bar for how data systems are built and operated. A defining piece of this role is owning the data architecture behind the DFM AI + IQE integration with a strategic partner. You will serve as the data engineering lead for the digital thread connecting the platform to partner ecosystems including Solid Edge, NX, Designcenter, and Teamcenter. You will build the pipelines, contracts, and observability that move quotes, parts, manufacturability signals, and pricing between systems in real time. Responsibilities Lead with technical depth – Design and drive the implementation of enterprise-scale data architecture and engineering solutions spanning multiple systems and domains. Own the partner integration data plane – Architect and build the data layer of the embedded DFM AI + IQE integration with Teamcenter and Designcenter. Own bidirectional pipelines, the joint data model for parts, BOMs, quotes, and manufacturability signals, low-latency feedback paths, and required governance, lineage, and audit controls. Build for scale – Architect and optimize reliable batch and streaming data pipelines, data models, and platforms handling complex, high-volume and event-driven data flows. Own the full lifecycle – Take end-to-end accountability from data acquisition and transformation through delivery, observability, and performance. Set the standard – Define and enforce best practices for data modeling, CI/CD, testing, code quality, contract testing, and schema evolution. Solve ambiguous problems – Navigate cross-domain technical challenges and deliver solutions meeting business and technical objectives. Develop multi-quarter roadmaps – Translate strategic priorities into technical plans and timelines. Collaborate broadly – Partner with engineering, product, data science, business stakeholders, and external partner engineering teams. Mentor and elevate – Guide engineers through design reviews, code reviews, and mentorship. Evaluate and adopt – Recommend tools, platforms, and architectural patterns within the data engineering ecosystem. Qualifications Bachelor's degree in a STEM field (or equivalent experience) and at least 5 years of experience in data engineering with ownership of large-scale data systems. Deep expertise with cloud data warehouses, preferably Snowflake, including optimization and performance tuning. Expert-level SQL and strong Python proficiency. Experience building and optimizing data pipelines and architectures using tools such as dbt, Airbyte, or Airflow. Experience planning and implementing enterprise data architecture across multiple systems and organizational boundaries. Working knowledge of queueing, batch and stream processing (Kafka, Spark, Kinesis) and scalable data stores (Apache Iceberg). Experience developing database-heavy services or APIs with focus on testability and maintainability. Strong understanding of CI/CD, automated testing, contract testing, and schema evolution in data pipelines. Strong knowledge of AWS data ecosystem and cloud-native infrastructure. Enterprise integration experience with PLM, ERP, or large SaaS systems; Teamcenter experience is a strong plus. Familiarity with data visualization tools such as Looker or Streamlit. Experience with data governance, data quality frameworks, and observability tooling. Exposure to lakehouse or data mesh architectures. Experience with infrastructure as code frameworks such as Terraform or CloudFormation. Experience with event-driven architecture, CDC pipelines, and low-latency operational data flows. Benefits Base salary range: $180,000–$200,000 annually plus commission, depending on experience and location. Competitive benefits package including 401(k) match, medical, dental, and vision insurance; life and disability insurance; generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave; employee assistance program and additional wellbeing resources.
Technology

Vulcan Elements
Data Engineer
Senior
On-site
Research Triangle Park, NC
🏢 Summary: Data Engineer role focused on designing and scaling data infrastructure, ETL pipelines, and Lakehouse architecture for a manufacturing environment supporting analytics and AI workloads. The position involves building operational data systems, ensuring data quality, and integrating industrial and manufacturing data sources. Candidates will collaborate cross-functionally and help establish scalable data architecture standards for future facility growth. 🗂️ Requirements: 8+ years of data engineering or data infrastructure experience, Experience designing data lakes or Lakehouse platforms, Experience building ETL/ELT pipelines, Strong data modeling expertise, Experience with relational databases, Strong SQL skills, Ability to document architecture and technical decisions, Experience collaborating with technical and non-technical stakeholders, U.S. Person status for export-controlled access 📃 Skills: SQL, PostgreSQL, SQLServer, ETL, ELT, Lakehouse, Python, Airflow, Prefect, dbt, InfluxDB, TimescaleDB, MQTT, DeltaLake, ApacheIceberg, AWS, Azure, GCP 🏢 Description: Vulcan Elements is manufacturing American rare-earth permanent magnets for a secure, resilient future. With a focus on national security and economic resiliency, we serve critical industries such as defense, aerospace, and automotive, powering a high-technology future. Vulcan Elements is building a team of ambitious professionals committed to Mission Focus, Technical Excellence, and Transparency. As the Data Engineer, you will design and build the data infrastructure that makes Vulcan's operational and business data useful — first at pilot scale, and then as the foundation for a 10,000 ton/year facility. You will work from architecture to implementation: evaluating and selecting platforms, designing data models and pipelines, and building the systems that collect, contextualize, and deliver data to the teams and tools that depend on it. You will collaborate closely with cross-functional stakeholders to translate operational requirements into a durable, scalable data architecture. As Vulcan grows, this role has the opportunity to expand into a team leadership position. Responsibilities Architecture & Platform Design - Design and own Vulcan's data architecture from operational data stores through ETL pipelines to the analytics and AI layer - Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, weighing scalability, compliance requirements, operational burden, and cost - Review, refine, and implement data architecture design documents, ensuring designs are technically sound and account for CUI and ITAR data handling requirements - Make and document key platform and design decisions with enough clarity that future team members can understand the reasoning and build on it - Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign - Apply sound engineering practices to everything you build: version control, testing, observability, and documentation, and hold those standards as the data team grows Data Pipeline & Integration - Design and build ETL pipelines that move data from operational data stores into the data Lakehouse with full contextual enrichment, making it ready for analytics and AI workloads - Build reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems - Collaborate across engineering, operations, and IT to understand data flows, dependencies, and integration requirements, and translate them into pipeline and architecture decisions - Identify and eliminate manual data workflows, replacing them with monitored, reliable pipelines - Diagnose and resolve data quality issues across the stack, and build monitoring into pipelines so problems surface early Data Modeling & Quality - Define data models that support operational queries, analytical workloads, and future AI and ML applications - Own data contextualization standards ensuring every data point carries the metadata needed to make it meaningful - Contribute to schema design and payload definitions for operational data stores, working toward consistency and legibility across the organization - Support the development of reporting and visibility tools that give operations and leadership clear insight into process and quality data - Write clear technical documentation for architecture decisions, data models, pipeline designs, and operational runbooks Responsibilities and tasks outlined are not exhaustive and may change as determined by the needs of the business. Qualifications - 8+ years of experience in data engineering, data infrastructure, or a closely related technical role with a track record of owning and delivering production systems - Demonstrated experience designing and building data lakes, Lakehouses, or analytical data stores; understands the tradeoffs between platforms and can make and defend platform selection decisions - Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data - Deep fluency with data modeling for both operational and analytical workloads; can design schemas that serve present needs without foreclosing future ones - Experience with relational databases (PostgreSQL, SQL Server, or similar); writes and debugs SQL confidently - Comfortable working in a fast-moving environment with a small team, making decisions with incomplete information and documenting them clearly for future colleagues - Strong communicator who can work across technical and non-technical stakeholders and translate between operational requirements and data architecture decisions - Must be a U.S. Person due to required access to U.S. export-controlled information or facilities Desired Skills - Experience with time-series databases (InfluxDB, TimescaleDB, or similar) common in industrial and IoT environments - Familiarity with industrial data concepts — historian data, process tags, OT/IT integration — and the data challenges specific to manufacturing environments - Experience working on or alongside a Unified Namespace or MQTT-based data architecture; understands how industrial messaging infrastructure relates to the data layer - Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar) - Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar) - Comfort with scripting and lightweight development (Python, SQL, or similar) for pipeline development and data quality tooling - Familiarity with cloud platforms (AWS, Azure, or GCP) and experience evaluating on-premises vs. cloud tradeoffs for data infrastructure - Experience working in a controlled information environment; familiarity with the handling requirements for Controlled Unclassified Information (CUI) or export-controlled technical data under ITAR or EAR - Experience in a manufacturing, industrial, or operations-heavy environment
Technology

Vulcan Elements
Data Engineer
Senior
On-site
Durham, NC
🏢 Summary: Data Engineer role focused on designing and scaling data infrastructure, ETL pipelines, and Lakehouse architecture for manufacturing operations supporting analytics and AI workloads. The position involves building reliable industrial data systems, defining data models, and collaborating across engineering and operations teams in a secure, compliance-driven environment. There is potential for future leadership responsibilities as the organization grows. 🗂️ Requirements: 8+ years of experience in data engineering or data infrastructure, Experience designing and building data lakes or Lakehouse platforms, Experience building ETL/ELT pipelines, Strong data modeling experience for operational and analytical workloads, Experience with relational databases, Strong SQL skills, Ability to work in fast-moving environments with small teams, Ability to communicate across technical and non-technical stakeholders, U.S. Person status for access to export-controlled information 📃 Skills: PostgreSQL, SQLServer, SQL, ETL, ELT, Lakehouse, Python, InfluxDB, TimescaleDB, MQTT, DeltaLake, Iceberg, Airflow, Prefect, dbt, AWS, Azure, GCP 🏢 Description: Vulcan Elements is manufacturing American rare-earth permanent magnets for a secure, resilient future. With a focus on national security and economic resiliency, the company serves critical industries such as defense, aerospace, and automotive. As the Data Engineer, you will design and build the data infrastructure that makes operational and business data useful — first at pilot scale, and then as the foundation for a 10,000 ton/year facility. You will work from architecture to implementation: evaluating and selecting platforms, designing data models and pipelines, and building the systems that collect, contextualize, and deliver data to the teams and tools that depend on it. You will collaborate closely with cross-functional stakeholders to translate operational requirements into a durable, scalable data architecture. As the organization grows, this role has the opportunity to expand into a team leadership position. Responsibilities Architecture & Platform Design - Design and own data architecture from operational data stores through ETL pipelines to the analytics and AI layer - Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, weighing scalability, compliance requirements, operational burden, and cost - Review, refine, and implement data architecture design documents, ensuring designs are technically sound and account for CUI and ITAR data handling requirements - Make and document key platform and design decisions with enough clarity that future team members can understand the reasoning and build on it - Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign - Apply sound engineering practices to everything you build: version control, testing, observability, and documentation, and hold those standards as the data team grows Data Pipeline & Integration - Design and build ETL pipelines that move data from operational data stores into the data Lakehouse with full contextual enrichment, making it ready for analytics and AI workloads - Build reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems - Collaborate across engineering, operations, and IT to understand data flows, dependencies, and integration requirements, and translate them into pipeline and architecture decisions - Identify and eliminate manual data workflows, replacing them with monitored, reliable pipelines - Diagnose and resolve data quality issues across the stack, and build monitoring into pipelines so problems surface early Data Modeling & Quality - Define data models that support operational queries, analytical workloads, and future AI and ML applications - Own data contextualization standards ensuring every data point carries the metadata needed to make it meaningful - Contribute to schema design and payload definitions for operational data stores, working toward consistency and legibility across the organization - Support the development of reporting and visibility tools that give operations and leadership clear insight into process and quality data - Write clear technical documentation for architecture decisions, data models, pipeline designs, and operational runbooks Responsibilities and tasks outlined are not exhaustive and may change as determined by business needs. Qualifications - 8+ years of experience in data engineering, data infrastructure, or a closely related technical role with a track record of owning and delivering production systems - Demonstrated experience designing and building data lakes, Lakehouses, or analytical data stores; understands the tradeoffs between platforms and can make and defend platform selection decisions - Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data - Deep fluency with data modeling for both operational and analytical workloads - Experience with relational databases (PostgreSQL, SQL Server, or similar) - Writes and debugs SQL confidently - Comfortable working in a fast-moving environment with a small team - Strong communicator able to work across technical and non-technical stakeholders - Must be a U.S. Person due to required access to U.S. export-controlled information or facilities Desired Skills - Experience with time-series databases (InfluxDB, TimescaleDB, or similar) - Familiarity with industrial data concepts including historian data, process tags, and OT/IT integration - Experience with Unified Namespace or MQTT-based data architecture - Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar) - Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar) - Comfort with scripting and lightweight development (Python, SQL, or similar) - Familiarity with cloud platforms (AWS, Azure, or GCP) - Experience working in controlled information environments with CUI, ITAR, or EAR requirements - Experience in manufacturing, industrial, or operations-heavy environments
Technology

Nimble Robotics
Data Engineer II/III
Mid
On-site
San Francisco, CA
140,004 - 200,004 USD/yr
🏢 Summary: Data Engineer role focused on building and scaling data infrastructure for advanced robotics systems. The position involves designing reliable batch and real-time pipelines, optimizing ETL/ELT processes, and implementing data governance across cloud platforms. You will collaborate cross-functionally to deliver high-quality, scalable data solutions supporting company-wide analytics and operations. 🗂️ Requirements: BS/MS/PhD in Computer Science, Mathematics, Computer Engineering or related field, or equivalent practical experience, 1–5 years of professional experience, Proficiency in Python and SQL, Hands-on experience with Rust, Go, Java, C#, or C++, Experience with Kafka, Spark, and lakehouse formats (Icebert, DeltaLake), Experience with AWS, GCP, or Azure, Strong debugging and problem-solving skills, Willingness to work extended hours and weekends if needed, Ability to work full time onsite in San Francisco 📃 Skills: Python, SQL, Rust, Go, Java, C#, C++, Kafka, Spark, Icebert, DeltaLake, AWS, GCP, Azure, Clickhouse, Flink, Databricks 🏢 Description: About the Role Help us advance our robotics moonshot by scaling our data engineering efforts. Drive design and development of data infrastructure across our products and internal tools. You will play a critical role in working with a cross-functional team to architect and build advanced robotic systems. Responsibilities - Design, build, and maintain scalable, reliable data pipelines that support both batch and real-time analytics. - Develop and optimize ETL/ELT processes to ingest, transform, and integrate data from multiple sources into data warehouses and lakehouse platforms. - Drive and apply best practices in data modeling, pipeline architecture, query optimization, data quality, and data engineering standards. - Collaborate closely with finance, bizops, and engineering teams to understand data needs and deliver solutions. - Provide data engineering support to ensure data accessibility and usability company-wide. - Implement and maintain data governance frameworks to support compliance with internal policies and external regulatory requirements. - Evaluate and adopt modern data engineering technologies, tools, and best practices to improve scalability, reliability, and operational efficiency. - Document data engineering processes, designs, and architectures. Qualifications - BS/MS/PhD in Computer Science, Mathematics, Computer Engineering, or a related field, or equivalent practical experience. - 1-5 years of professional experience. - Proficiency in writing production-grade code in Python and SQL. - Hands-on experience with one of the following languages: Rust, Go, Java, C#, C++. - Strong debugging skills and the ability to diagnose and resolve issues efficiently. - Experience with Kafka, Spark and common lakehouse formats like Icebert, DeltaLake. - Experience working with cloud platforms like AWS, GCP, Azure. Preferred Experience - Experience with Clickhouse, Apache Flink, Databricks. - Experience on financial reporting and understanding of accountability. Additional Requirements - Willing to work extended hours and weekends if needed. - This position is based full time in our San Francisco headquarters. Compensation The pay range for this position at the start of employment is expected to be between $140,000 - $200,000/year. The exact offer may vary depending on job-related knowledge, skills, and experience. In addition to cash compensation, this position will also receive generous equity. Benefits - Unlimited Flexible Time Off. - Health insurance (medical, dental, vision). - Paid parental leave. - Commuter benefits including fully paid parking spots. - Referral bonus. - 401k retirement plan. - Equity program.
Technology
TechTree
Lead Data Engineer
Senior
Remote
Krakow, Poland
270,000 - 406,000 PLN/yr
🏢 Summary: Lead Data Engineer role focused on driving architecture and leading a team to build scalable, secure ETL/ELT pipelines and analytics-ready data models on modern cloud platforms. The position combines hands-on engineering with technical leadership, ensuring high standards in governance, observability, and performance optimisation. You will shape data infrastructure that supports large-scale analytics across the organisation. 🗂️ Requirements: Proven experience leading data engineering or analytics engineering teams, Strong programming skills in SQL, Strong programming skills in Python, Hands-on experience with Airflow or Prefect in production, Deep practical experience with dbt, Experience with Snowflake or Databricks at scale, Strong knowledge of dimensional modelling and SCD strategies, Experience implementing data quality and governance frameworks, Experience with CI/CD and automated testing in data systems 📃 Skills: SQL, Python, Airflow, Prefect, dbt, Snowflake, Databricks, ETL, ELT, CI/CD, SCD, Dimensional, Git, IaC 🏢 Description: ABOUT THE COMPANY Our client is a global legal technology company that has been building software for the legal industry for over two decades. Our AI-powered cloud platform is used by leading law firms, Fortune 500 corporations, and government agencies worldwide to organise complex data, surface critical insights, and act on them — across litigation, investigations, regulatory inquiries, and data breach response. We're valued at $3.6 billion and invest over $170 million annually in R&D. We're making substantial investments in data lake technology and distributed systems to support future growth and advanced analytics. Our scale means the data problems here are genuinely hard — and the infrastructure you lead will have real consequence across the organisation. ABOUT THE ROLE We're looking for a Lead Data Engineer to combine deep technical expertise with hands-on team leadership, guiding a team of data engineers building and maintaining ETL/ELT pipelines, data models, and governance frameworks that power analytics and reporting across the organisation. This is a technical leadership role — you'll drive architectural decisions, mentor engineers, and ensure delivery of secure, reliable, and scalable data solutions. You'll collaborate closely with stakeholders to align technical work with business objectives, champion governance and observability standards, and foster a culture of continuous improvement. The expectation is that you're equally effective in an architecture review as you are pairing with an engineer on a tricky pipeline problem. WHAT YOU'LL WORK ON Team leadership and mentorship Lead and mentor a team of data engineers, promoting collaboration, knowledge sharing, and professional growth. Set the standard for engineering quality and hold the bar consistently. Architecture and pipeline design Drive architectural decisions for ETL/ELT pipelines, orchestration frameworks (Airflow/Prefect), and transformation layers (dbt). Facilitate architecture reviews and contribute to design decisions for scalable, fault-tolerant systems. Analytics data modelling Oversee design and implementation of analytics-ready data models — dimensional schemas, SCD strategies, and semantic layers — that internal teams can build on reliably. Engineering best practices Ensure adherence to clean code, modular design, CI/CD, automated testing, and code review standards across all data engineering work. Platform optimisation Manage performance tuning and cost optimisation for Snowflake, Databricks, and related cloud data platforms at scale. Governance and observability Champion governance, observability, and compliance frameworks across all data workflows — including data quality, lineage tracking, and multi-tenant environment controls. Stakeholder communication Communicate effectively with leadership and cross-functional teams to provide updates, resolve blockers, and ensure timely delivery aligned with business objectives. WHAT WE LOOK FOR Proven technical team leadership Demonstrated experience leading data engineering or analytics-focused development teams — mentoring engineers, driving architectural decisions, and owning delivery outcomes. SQL and Python Strong programming skills in both SQL and Python, applied to production data systems at scale. ETL/ELT orchestration Hands-on experience with orchestration tools — Airflow and/or Prefect — in production pipeline environments. dbt expertise Deep practical experience with dbt for transformation workflows and analytics modelling, including testing, documentation, and modular project design. Snowflake and Databricks Familiarity with Snowflake and/or Databricks for large-scale data processing, including performance tuning and cost management. Data modelling principles Solid understanding of data modelling principles, incremental strategies, and schema design for analytics — dimensional modelling, SCDs, and semantic layer design. Governance and data quality Knowledge of data quality frameworks, lineage tracking, and governance in multi-tenant environments. Software engineering practices Familiarity with CI/CD, automated testing, and infrastructure-as-code practices applied to data systems. Communication and stakeholder management Strong communication skills with the ability to operate confidently across technical teams and business stakeholders. THE TEAM You'll join a global engineering organisation working on a platform used by some of the world's largest legal teams. The culture is diverse, inclusive, and driven by high standards. Engineers here work on genuinely complex technical problems at scale — and are supported with the coaching, development, and tooling to keep growing. COMPENSATION & BENEFITS Salary 270,000 – 406,000 PLN per year, plus an annual performance bonus and long-term incentives. Health coverage Comprehensive health, dental, and vision plans. Parental leave Parental leave available for both primary and secondary caregivers. Flexible working Flexible work arrangements, hybrid model. Company breaks Two week-long company-wide breaks per year, plus additional time off. Training investment Dedicated training investment programme to support ongoing professional development.
Technology
Harvey Nash Technology
Senior Data Engineer (cloud&ai)
Senior
On-site
Warsaw, Poland
30,000 - 40,000 PLN
🏢 Summary: Design and scale high-throughput data pipelines on cloud platforms to support advanced analytics and AI-driven products. The role focuses on building distributed data architectures in AWS and Databricks, ensuring performance, governance, and data quality. You will collaborate with AI/ML teams to deliver scalable, production-grade data solutions. 🗂️ Requirements: 3+ years of data engineering experience, Strong Python programming skills, Experience with Spark or Scala, Experience building distributed data pipelines in cloud environments, Knowledge of data modeling and data warehousing principles, Bachelor’s or Master’s degree in Computer Science or Engineering 📃 Skills: Python, Spark, Scala, AWS, Glue, EMR, Fargate, StepFunctions, Databricks, SQL, APIs, GenAI, GraphDB 🏢 Description: Data Engineer – Cloud & AI Platforms We’re looking for a Data Engineer to design and scale high-throughput data pipelines supporting advanced analytics and AI-driven products. What You’ll Do Architect and maintain distributed data pipelines in Databricks and AWS (Glue, EMR, Fargate, Step Functions) Ingest and process large volumes of structured and unstructured data (internal, market, third-party, alternative sources) Collaborate with AI/ML and engineering teams to design scalable data architectures and APIs Optimize performance and cost using Spark and cloud-native best practices Implement data governance, privacy, lineage, and access controls Build automated validation, monitoring, and data quality frameworks Evaluate emerging GenAI and data tooling to enhance platform capabilities What You Bring 3+ years of experience in data engineering Strong Python and experience with Spark or Scala Proven experience building distributed pipelines in cloud environments Solid understanding of data modeling, architecture, and warehousing principles Innovative problem-solving mindset Bachelor’s or Master’s degree in Computer Science or Engineering Nice to have: Experience with graph databases.
Technology
emagine Polska
Backend Engineer - Java
Mid
Remote
Stockholm, Sweden
🏢 Summary: Hands-on data infrastructure engineering role focused on large-scale pipeline migrations and evolution of the company’s data processing stack. The position involves contributing to platform development across Flink and Lakehouse architectures while ensuring performance, reliability, and cost efficiency. High-impact role embedded in a data engineering team delivering production-grade data platforms. 🗂️ Requirements: Strong Java development experience, Experience with JVM-based data processing framework, Experience with Flink, Beam, Dataflow or Spark, Proficiency in SQL, Experience with BigQuery, Experience with cloud infrastructure, Experience with containerized applications, Knowledge of Kubernetes basics, Experience with Scala or Python for data pipelines, Experience working with production data engineering systems 📃 Skills: Java, Flink, Beam, Dataflow, Spark, SQL, BigQuery, Kubernetes, Scala, Python, JVM, DevOps, Lakehouse, Cloud 🏢 Description: The Data Infrastructure PA enables the company to solve complex and critical data engineering problems by providing platforms and tooling for the production, management, and consumption of high-quality data. We're looking for an engineer to support hands-on implementation and migration work as we evolve our data processing stack. This is a high impact and execution-focused engagement — you'll be contributing to company wide migration efforts and platform development. What You'll Work On You'll be embedded in a team in Data Infrastructure PA, contributing to hands-on engineering work. This includes large-scale pipeline migrations — validating performance and cost outcomes and helping move workloads to our evolving stack — as well as contributing to platform development across our Flink platform, Lakehouse architecture and beyond, as our priorities evolve. What We're Looking For You have solid, hands-on experience in backend engineering and are comfortable jumping into an existing platform codebase and making meaningful contributions quickly. Specifically: Strong Java development skills, with experience in data platform or data engineering contexts Practical experience with at least one JVM-based data processing framework — Flink experience is a plus; Beam, Dataflow, or Spark also relevant Comfortable with SQL and cloud data analytics platforms, particularly BigQuery DevOps is part of your day-to-day: you work with cloud infrastructure, containerised applications, and are familiar with Kubernetes basics Experience working with data engineering pipelines in Scala and/or Python You write quality code and understand what it means to ship reliably in a production environment You can work autonomously in an ambiguous environment and move quickly without waiting to be directed Nice to Have Prior experience with large-scale pipeline migrations Familiarity with cost optimisation in cloud data processing workloads Job Posting Start Date: 2026-05-18 Job Posting End Date: 2026-11-27
Technology
Team Up
🤖 Lead Data Engineer with AI (m/k) 🤖
Senior
Hybrid
Wroclaw, Poland
🏢 Summary: Lead Data Engineer role focused on driving AI initiatives and building scalable cloud-based data architectures in a global environment. The position involves technical ownership of data platforms, designing secure and high-performing systems, and leading data engineering efforts in a DevOps setting. The role emphasizes AI integration, cloud infrastructure, and enterprise-grade data governance. 🗂️ Requirements: Degree in Computer Science, AI, Data Science, Software Engineering or equivalent experience, 8+ years in software engineering, 5+ years of backend development with Python in production, Strong experience designing and scaling complex data systems, Hands-on experience with AI technologies, Hands-on experience with AWS or Azure, Strong knowledge of Python and SQL, Experience with APIs and data integration, Experience with automation tools, Knowledge of data governance practices, Understanding of data security and compliance standards, Proven experience leading and mentoring engineers 📃 Skills: Python, SQL, AWS, Azure, Java, AI, RAG, MCP, APIs, DevOps, Automation, Monitoring, Cloud, DataEngineering, DataPipelines, Governance, Security, Compliance, Backend 🏢 Description: We are looking for an experienced Lead Data Engineer to drive AI and cloud-based data solutions within a global technology organization. In this role, you will lead data initiatives, shape scalable architectures, and collaborate with both technical teams and senior stakeholders to deliver secure and high-performing systems. Key Responsibilities: Act as the main point of contact for data access and system-related topics with senior stakeholders Lead and mentor data engineers, promoting best practices and technical excellence Design, build, and maintain scalable cloud infrastructure and data pipelines Ensure data quality, security, compliance, and governance across the full lifecycle Develop secure and reliable cloud architectures (AWS/Azure) for AI and enterprise applications Implement monitoring, alerting, disaster recovery, and business continuity solutions Take technical ownership of applications within a DevOps environment Drive automation and self-service capabilities Support AI initiatives (e.g., AI Agents, RAG, MCP) with focus on quality and scalability Stay updated on emerging technologies and advise on strategic data direction Requirements: Degree in Computer Science, AI, Data Science, Software Engineering, or equivalent experience 8+ years in software engineering, including 5+ years of backend development with Python (production level) Strong experience designing and scaling complex data systems Hands-on experience with AI technologies and cloud platforms (AWS or Azure) Solid knowledge of Python, SQL (Java is a plus) Experience with APIs, data integration, automation tools, and data governance Strong understanding of data security and compliance standards Proven leadership and mentoring experience Excellent communication skills in English and Polish (min. B2); German is a plus What We Offer: Opportunity to work in a global, international environment Real impact on AI and cloud solutions in a large-scale organization Access to training platforms and professional development programs Hybrid work model with flexible hours (modern office in central Wroclaw) Comprehensive benefits package (medical & dental care, sports card, life insurance, mental health program) Cafeteria benefits platform with monthly points CSR initiatives, integration events, and employee passion clubs
