July 14, 2026
Staff Site Reliability Engineer (SRE)
Senior • On-site
165,000 - 189,996 USD/yr
Waltham, MA
Xometry is seeking a Staff Site Reliability Engineer to join our Site Reliability Engineering (SRE) Organization. In this role as a senior individual contributor, you will define the technical strategy and direction for the reliability and performance of our infrastructure and software systems. You will lead complex, cross-functional initiatives, mentor senior engineers, and influence key architectural and business decisions across our technology organization. You will utilize your deep technical expertise to architect and implement highly reliable and scalable infrastructure solutions that empower our technology organization to quickly and safely deliver features to customers at scale.
How You'll Contribute
- Lead the resolution of complex cross-cutting technical challenges and drive them to completion
- Define the long-term technical roadmap and strategy for core SRE platforms and services
- Architect, scope, and drive delivery of major engineering initiatives
- Set standards for clean, efficient, and well-documented code
- Collaborate across teams and communicate progress and outcomes to stakeholders
- Provide technical mentorship and promote operational excellence, observability, and infrastructure security best practices
- Develop, configure, and maintain AWS platforms, networking, and Kubernetes clusters
- Maintain observability and monitoring tools including Coralogix and Sentry
- Develop and maintain CI/CD tools including GitHub Actions runners and ArgoCD
What You'll Bring
- 8+ years of professional experience in infrastructure management or backend software development
- Deep expertise in Python, Javascript, or Unix Shell
- Extensive AWS experience across multiple accounts and regions
- Strong knowledge of Terraform, Kubernetes, CI/CD pipelines, and Docker/containerization
- Experience leading technical projects with cross-functional impact
- Ability to operate in escalation scenarios for complex production issues
- Excellent communication and collaboration skills
Benefits
- Base salary range of $165,000 - $190,000 annually plus bonus
- 401(k) match
- Medical, dental, and vision insurance
- Life and disability insurance
- Paid time off including vacation, sick leave, and holidays
- Maternity and bonding leave
- EAP and wellbeing resources
Similar jobs you might like
Technology
EPAM Systems
Senior Site Reliability Engineer (SRE)
Senior
Remote
🏢 Summary: The offer is for a Site Reliability Engineer responsible for ensuring high reliability, scalability, and performance of cloud-based systems. The role focuses on implementing SRE practices, automating infrastructure, managing incidents, and enhancing monitoring and CI/CD processes. You will collaborate with cross-functional teams to optimize operations and maintain service excellence. 🗂️ Requirements: Bachelor’s degree in Computer Science, Engineering, or related field, 3+ years of experience in Site Reliability Engineering or similar role, Experience with cloud platforms (AWS, GCP, or Azure), Hands-on experience with SRE practices (SLO, SLI, error budgets, postmortems, toil reduction, capacity planning, incident management), Proficiency in Python or other scripting/programming language, Experience with monitoring tools, Experience with CI/CD tools, Experience with infrastructure as code, Experience with configuration management, Knowledge of Kubernetes and Docker, English proficiency B2 or higher 📃 Skills: AWS, GCP, Azure, Python, Kubernetes, Docker, CI/CD, Terraform, Ansible, Monitoring, SLO, SLI, Git, Bash 🏢 Description: We are seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our team. In this critical role, you will collaborate closely with software developers and operations teams to ensure high reliability, scalability, and efficiency of our systems, with a strong focus on meeting and exceeding customer expectations. Your expertise will be crucial in deploying, maintaining, and automating our infrastructure and application environments to ensure seamless user experiences. Your proactive involvement will be key to enhancing system reliability, optimizing resource utilization, and ensuring continuous improvement in our operational practices. Your responsibilities will include defining and tracking Service Level Objectives (SLOs), managing error budgets, and reducing toil through automation. You will play a pivotal role in driving the success of technology initiatives, maximizing their impact across the organization, and ensuring that solutions consistently meet the high standards our customers expect. Responsibilities Collaborate with development, security, quality, and operation teams to implement SRE practices and ensure system reliability Define and support required level of reliability, availability, and performance for services and applications Design and deliver Cloud-based solutions tailored to client needs Troubleshoot, mitigate, and support fixing of the infrastructure and application issues in a timely manner Implement a monitoring system for the infrastructure and application reliability Communicate technical concepts clearly to both engineering teams and management stakeholders Requirements Bachelor’s degree in Computer Science, Engineering, or a related field 3+ years of hands-on experience in Site Reliability Engineering or related roles Proven experience in any cloud (AWS/GCP/Azure) Experience with implementing SRE practices such as SLO/SLI, Error budgets, Postmortems, Reducing Toil, capacity planning, and Incident Management Python or other scripting/programming language Strong background in monitoring tools Proficiency in CI/CD tools, infrastructure as code, and configuration management Solid knowledge of container orchestration technologies (Kubernetes, Docker) English language proficiency at an Upper-Intermediate level (B2) or higher Nice to have Expertise in deployment and management of LLMs, including technologies like RAG Certification in Kubernetes, AWS/GCP/Azure, or similar technologies Proven experience in DevOps Knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance We offer/Benefits We gather like-minded people: Engineering community of industry professionals Friendly team and enjoyable working environment Flexible schedule and opportunity to work remotely within Poland Chance to work abroad for up to 60 days annually Business-driven relocation opportunities We provide growth opportunities: Outstanding career roadmap Leadership development, career advising, soft skills, and well-being programs Certification (GCP, Azure, AWS) Unlimited access to LinkedIn Learning, Get Abstract, Cloud Guru English classes We cover it all: Stable income (Employment Contract or B2B) Participation in the Employee Stock Purchase Plan Benefits package (health insurance, multisport, shopping vouchers) Strategically located offices featuring entertainment and relaxation zones, table tennis and football, free snacks, fantastic coffee, and more Referral bonuses Corporate, social and well-being events Please, note: The set of bonuses might vary based on the role you apply for – specifics will be discussed with our recruiter during the general interview. We will reach out to selected candidates exclusively. EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
Technology
Link Group
Site Reliability Engineer
Mid
Hybrid
Warsaw, Poland
🏢 Summary: Hands-on Site Reliability Engineer role focused on building and scaling reliability practices across cloud and on-prem environments. The position involves improving performance, scalability, and resilience of production systems through automation, observability, and Kubernetes-based infrastructure. You will drive SRE standards and collaborate with engineering teams to enhance system stability and fault tolerance. 🗂️ Requirements: 4+ years experience in SRE, DevOps or similar roles, Strong experience with distributed systems, Strong experience with Kubernetes, Experience with AWS cloud, Hands-on automation experience with Python, Bash or Go, Solid understanding of CI/CD practices, Experience with observability and monitoring tools, Experience managing production systems 📃 Skills: Kubernetes, AWS, Python, Bash, Go, Prometheus, Grafana, CI/CD, SRE, DevOps 🏢 Description: We’re looking for a Site Reliability Engineer (SRE) to help build and scale reliability practices across our engineering organization. This is a hands-on role where you’ll work across cloud and on-prem environments, improving the performance, scalability, and resilience of critical production systems. 🔧 What you’ll be doing: • Driving SRE best practices, standards, and ways of working • Building and scaling observability & monitoring solutions (e.g. Prometheus, Grafana) • Working with Kubernetes-based infrastructure to ensure reliability and efficiency • Automating deployments, incident response, and recovery processes • Collaborating closely with engineering teams to improve system stability and fault tolerance • Contributing to a strong reliability culture (SLOs, post-mortems, continuous improvement) ✅ What we’re looking for: • 4+ years of experience in SRE / DevOps / similar roles • Strong experience with distributed systems, Kubernetes, and cloud (AWS preferred) • Hands-on approach to automation (Python, Bash, or Go) • Solid understanding of CI/CD and modern software delivery • Proactive mindset and strong ownership of production systems Name and surname*
Technology

Xometry
Staff Data Engineer
Senior
On-site
Waltham, MA
180,000 - 200,004 USD/yr
🏢 Summary: Senior individual contributor role leading enterprise-scale data architecture and real-time partner integrations, owning the design of scalable batch and streaming pipelines across systems. Responsible for building and operating the data plane behind a strategic DFM AI + IQE integration, enabling low-latency, bidirectional data flows between platforms. Sets engineering standards for data modeling, CI/CD, governance, and observability while collaborating cross-functionally. 🗂️ Requirements: Bachelor's degree in STEM or equivalent experience, 5+ years in data engineering with ownership of large-scale data systems, Deep expertise in Snowflake and cloud data warehouses, Expert-level SQL, Strong Python proficiency, Experience building and optimizing modern data pipelines (dbt, Airbyte, Airflow or similar), Experience designing enterprise data architecture across multiple systems and partner boundaries, Knowledge of batch and stream processing systems, Experience with highly scalable data stores, Experience writing database-heavy services or APIs, Strong understanding of CI/CD, automated testing, contract testing, schema evolution, Strong knowledge of AWS and cloud-native infrastructure, Enterprise or partner system integration experience (PLM, ERP, or SaaS), Experience with infrastructure as code frameworks, Experience with event-driven architectures and CDC pipelines 📃 Skills: Snowflake, SQL, Python, dbt, Airbyte, Airflow, Kafka, Spark, Kinesis, Apache, Iceberg, AWS, Terraform, CloudFormation, Teamcenter, BMIDE, APIs, CI/CD, CDC, Looker, Streamlit 🏢 Description: Xometry is looking for a Staff Data Engineer to join the Data Platform team. This is a senior individual contributor role with broad technical scope and high organizational impact. You will own data architecture decisions, lead the design of scalable pipelines and platforms, and set the engineering bar for how data systems are built and operated. A defining piece of this role is owning the data architecture behind the DFM AI + IQE integration with a strategic partner. You will serve as the data engineering lead for the digital thread connecting the platform to partner ecosystems including Solid Edge, NX, Designcenter, and Teamcenter, building pipelines, data contracts, and observability to move quotes, parts, manufacturability signals, and pricing data in real time. Responsibilities - Lead the design and implementation of enterprise-scale data architecture and engineering solutions across multiple systems and domains - Architect and build the data layer for embedded DFM AI + IQE integrations, including bidirectional pipelines and joint data models for parts, BOMs, quotes, and manufacturability signals - Design low-latency signal paths delivering DFM and pricing feedback into designer environments - Establish governance, lineage, and audit capabilities for partner-integrated data systems - Architect and optimize reliable batch and streaming pipelines for complex, high-volume, event-driven data flows - Own the full lifecycle of data engineering work from ingestion and transformation to delivery and observability - Define and enforce best practices for data modeling, CI/CD, testing, code quality, contract testing, and schema evolution - Solve complex cross-domain technical challenges aligned with business objectives - Develop multi-quarter technical roadmaps and execution plans - Collaborate with engineering, product, data science, business stakeholders, and partner engineering teams - Mentor engineers through design and code reviews - Evaluate and recommend tools, platforms, and architectural patterns Qualifications - Bachelor's degree in a STEM field (or equivalent experience) - At least 5 years of experience in data engineering with ownership of complex, large-scale systems - Deep expertise in Snowflake, including optimization and performance tuning - Expert-level SQL and strong Python proficiency - Experience with modern data tooling such as dbt, Airbyte, and Airflow - Experience designing enterprise data architectures spanning multiple systems and partner boundaries - Knowledge of batch and stream processing technologies (e.g., Kafka, Spark, Kinesis) and scalable data stores (e.g., Apache Iceberg) - Experience building database-heavy services or APIs with focus on testability and maintainability - Strong understanding of CI/CD, automated testing, contract testing, and schema evolution in data pipelines - Strong knowledge of AWS and cloud-native infrastructure - Enterprise integration experience with PLM, ERP, or large SaaS systems; Teamcenter experience is a strong plus - Familiarity with data visualization tools such as Looker or Streamlit - Experience with data governance, data quality frameworks, and observability tooling - Exposure to lakehouse or data mesh architectures - Experience with infrastructure as code (Terraform, CloudFormation) - Experience with event-driven architectures, CDC pipelines, and low-latency operational data flows Benefits - Estimated base salary range: $180,000–$200,000 annually plus commission, depending on experience and location - Competitive benefits package including 401(k) match - Medical, dental, and vision insurance - Life and disability insurance - Generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave - Employee assistance and wellbeing resources
Technology
Link Group
Senior Site Reliability Engineer
Senior
Hybrid
Warsaw, Poland
170 - 230 PLN
🏢 Summary: The role focuses on ensuring reliability, scalability, and performance of large-scale cloud-based applications by building and maintaining resilient infrastructure. You will manage AWS cloud environments, Kubernetes clusters, and CI/CD pipelines while implementing monitoring, automation, and incident response processes. The position emphasizes Infrastructure-as-Code, observability, and continuous reliability improvements. 🗂️ Requirements: 5+ years experience in SRE, DevOps or similar role, Strong experience with AWS cloud services, Experience with Infrastructure-as-Code tools, Hands-on experience with Kubernetes, Proficiency with Docker, Experience with CI/CD pipelines, Solid knowledge of PostgreSQL or Amazon RDS, Strong SQL knowledge, Knowledge of networking concepts (VPC, DNS, troubleshooting), Strong Linux/Unix administration skills, Experience with observability tools, Experience with automation in infrastructure, Experience with incident management 📃 Skills: AWS, Terraform, Pulumi, Kubernetes, EKS, Docker, GitHub, PostgreSQL, RDS, SQL, VPC, DNS, Linux, Unix, Prometheus, Grafana, Datadog, Dynatrace, CI/CD 🏢 Description: We are looking for an experienced Site Reliability Engineer to ensure the reliability, scalability, and performance of large-scale cloud-based web applications. You will work closely with software development, cloud operations, and platform teams to build and maintain resilient infrastructure and improve system stability. Key Responsibilities: Design and maintain monitoring, alerting, and incident response systems to ensure high availability Collaborate closely with engineering, product, and architecture teams Build and manage cloud infrastructure using Infrastructure-as-Code (e.g., Terraform, Pulumi) on AWS Operate and optimize Kubernetes environments (e.g., EKS) Develop and maintain containerized applications using Docker Improve CI/CD pipelines and drive automation across deployment processes Implement and manage observability tools (logging, metrics, tracing) Participate in incident management, postmortems, and reliability improvements Support capacity planning, disaster recovery, and system scaling Contribute to security, compliance, and operational best practices Develop automation and AI-driven solutions for monitoring and incident prevention Requirements: 5+ years of experience in SRE, DevOps, or similar roles Strong experience with AWS cloud services and Infrastructure-as-Code tools Hands-on experience with Kubernetes and containerized environments Proficiency in Docker and CI/CD pipelines (e.g., GitHub Actions) Solid understanding of databases (e.g., PostgreSQL, Amazon RDS) and SQL Knowledge of networking concepts (VPC, DNS, troubleshooting tools like dig/traceroute) Strong Linux/Unix administration skills Experience with observability tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) Familiarity with automation and AI-based solutions in infrastructure Strong problem-solving and incident management skills
Technology

Xometry
Staff Data Engineer
Senior
On-site
North Bethesda, MD
180,000 - 200,004 USD/yr
🏢 Summary: Senior individual contributor role responsible for designing and owning enterprise-scale data architecture and real-time data pipelines that power a strategic DFM AI + IQE partner integration. The position focuses on building scalable batch and streaming systems, defining data models, and ensuring governance, observability, and CI/CD standards across cross-system integrations. The engineer leads the digital data plane connecting internal platforms with external PLM ecosystems in a high-impact, cloud-native environment. 🗂️ Requirements: Bachelor’s degree in STEM or equivalent experience, Minimum 5 years of experience in data engineering, Deep expertise in Snowflake or similar cloud data warehouse, Expert-level SQL, Strong Python proficiency, Hands-on experience with modern data pipeline tools (dbt, Airbyte, Airflow or similar), Experience designing enterprise data architecture across multiple systems, Knowledge of batch and stream processing systems, Experience with highly scalable data stores, Experience with CI/CD, automated testing, contract testing, schema evolution, Strong knowledge of AWS data ecosystem, Experience integrating with enterprise or partner systems (e.g., PLM, ERP, SaaS) 📃 Skills: Snowflake, SQL, Python, dbt, Airbyte, Airflow, Kafka, Spark, Kinesis, Apache, Iceberg, AWS, Teamcenter, BMIDE, APIs, Looker, Streamlit, Terraform, CloudFormation, CI/CD, CDC 🏢 Description: Xometry is looking for a Staff Data Engineer to join the Data Platform team. This is a senior individual contributor role with broad technical scope and high organizational impact. You will own data architecture decisions, lead the design of scalable pipelines and platforms, and set the engineering bar for how data systems are built and operated. A defining piece of this role is owning the data architecture behind the DFM AI + IQE integration with a strategic partner. You will serve as the data engineering lead for the digital thread connecting the platform to partner ecosystems including Solid Edge, NX, Designcenter, and Teamcenter. You will build the pipelines, contracts, and observability that move quotes, parts, manufacturability signals, and pricing between systems in real time. Responsibilities Lead with technical depth – Design and drive the implementation of enterprise-scale data architecture and engineering solutions spanning multiple systems and domains. Own the partner integration data plane – Architect and build the data layer of the embedded DFM AI + IQE integration with Teamcenter and Designcenter. Own bidirectional pipelines, the joint data model for parts, BOMs, quotes, and manufacturability signals, low-latency feedback paths, and required governance, lineage, and audit controls. Build for scale – Architect and optimize reliable batch and streaming data pipelines, data models, and platforms handling complex, high-volume and event-driven data flows. Own the full lifecycle – Take end-to-end accountability from data acquisition and transformation through delivery, observability, and performance. Set the standard – Define and enforce best practices for data modeling, CI/CD, testing, code quality, contract testing, and schema evolution. Solve ambiguous problems – Navigate cross-domain technical challenges and deliver solutions meeting business and technical objectives. Develop multi-quarter roadmaps – Translate strategic priorities into technical plans and timelines. Collaborate broadly – Partner with engineering, product, data science, business stakeholders, and external partner engineering teams. Mentor and elevate – Guide engineers through design reviews, code reviews, and mentorship. Evaluate and adopt – Recommend tools, platforms, and architectural patterns within the data engineering ecosystem. Qualifications Bachelor's degree in a STEM field (or equivalent experience) and at least 5 years of experience in data engineering with ownership of large-scale data systems. Deep expertise with cloud data warehouses, preferably Snowflake, including optimization and performance tuning. Expert-level SQL and strong Python proficiency. Experience building and optimizing data pipelines and architectures using tools such as dbt, Airbyte, or Airflow. Experience planning and implementing enterprise data architecture across multiple systems and organizational boundaries. Working knowledge of queueing, batch and stream processing (Kafka, Spark, Kinesis) and scalable data stores (Apache Iceberg). Experience developing database-heavy services or APIs with focus on testability and maintainability. Strong understanding of CI/CD, automated testing, contract testing, and schema evolution in data pipelines. Strong knowledge of AWS data ecosystem and cloud-native infrastructure. Enterprise integration experience with PLM, ERP, or large SaaS systems; Teamcenter experience is a strong plus. Familiarity with data visualization tools such as Looker or Streamlit. Experience with data governance, data quality frameworks, and observability tooling. Exposure to lakehouse or data mesh architectures. Experience with infrastructure as code frameworks such as Terraform or CloudFormation. Experience with event-driven architecture, CDC pipelines, and low-latency operational data flows. Benefits Base salary range: $180,000–$200,000 annually plus commission, depending on experience and location. Competitive benefits package including 401(k) match, medical, dental, and vision insurance; life and disability insurance; generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave; employee assistance program and additional wellbeing resources.
Technology
emagine Polska
Senior DevOps / SRE (Platform Reliability Engineer) - French fluent
Senior
Remote
Lisbon, Portugal
🏢 Summary: Senior DevOps / SRE role focused on ensuring reliability, scalability, security, and performance of a cloud-native AWS platform. The position centers on infrastructure automation, CI/CD, Kubernetes operations, observability, and implementing SRE best practices to support highly available production systems. You will lead incident management, optimize cloud costs, and drive continuous improvement of platform resilience. 🗂️ Requirements: 5+ years in DevOps/SRE/Cloud/Platform Engineering, Strong Linux administration and troubleshooting, Production experience with Kubernetes, Experience with CI/CD tools, Expertise in Infrastructure as Code, Hands-on experience with AWS, Strong networking fundamentals, Experience with monitoring and logging tools, Scripting skills (Bash or Python) 📃 Skills: AWS, Kubernetes, Docker, Helm, Terraform, Ansible, CloudFormation, Linux, GitLab, Jenkins, GitHub, Azure, Prometheus, Grafana, ELK, Datadog, Splunk, Bash, Python, TCP/IP, DNS 🏢 Description: We are looking for a Senior DevOps / Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and security of our platform and cloud infrastructure. You will play a key role in building and operating cloud-native systems, improving observability, automating operations, implementing SRE best practices (SLOs/SLIs), and supporting development teams to deliver highly available services. Key Responsibilities Design, implement, and maintain highly available and scalable infrastructure on AWS. Own and improve the reliability of production systems using SRE principles (SLO, SLI, error budgets). Build and manage CI/CD pipelines to support fast and safe software delivery. Develop and maintain Infrastructure as Code (IaC) using Terraform, Ansible, CloudFormation, etc. Manage and optimize container orchestration platforms (Kubernetes, Docker, Helm). Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK, Datadog, Splunk). Lead incident response, perform root cause analysis, and write postmortems to drive continuous improvement. Improve system performance, capacity planning, scaling strategies, and disaster recovery processes. Collaborate closely with development teams to improve deployment strategies and system resilience. Implement security best practices (IAM, secret management, vulnerability scanning, patching). Define operational standards, runbooks, documentation, and best practices for platform reliability. Participate in on-call rotation and provide senior-level support for critical production issues. Key Responsibilities (5 Main Missions) The DevOps / SRE lead will be responsible for the stability and evolution of the platform. Your role is structured around five main areas: Mission 1: AWS Infrastructure Management (Build & Run) Mission 2: CI/CD and Deployment Automation Mission 3: Monitoring, Observability, and Alerting: Global Monitoring , Log Management , Application Monitoring , Business Analytics Mission 4: Incident Management, Resilience, and Security Mission 5: FinOps and AWS Cost Optimization Key Requirements 5+ years of experience in DevOps / SRE / Cloud Infrastructure / Platform Engineering. Strong expertise in Linux systems administration and troubleshooting. Proven experience with Kubernetes in production environments. Strong experience with CI/CD tools (GitLab CI, Jenkins, GitHub Actions, Azure DevOps). Solid knowledge of Infrastructure as Code (Terraform highly preferred). Experience with AWS cloud platforms. Strong understanding of networking fundamentals (TCP/IP, DNS, load balancing, reverse proxies). Experience with observability tools: monitoring, metrics, logging, tracing. Strong scripting skills (Bash, Python, or similar). French advanced level. Nice to Have Experience with additional cloud platforms (Azure, GCP). Strong understanding of networking fundamentals.
Technology

Xometry
Staff Cyber Resilience Engineer
Senior
On-site
Waltham, MA
204,996 - 233,004 USD/yr
🏢 Summary: Staff Cyber Resilience Engineer role focused on designing and automating secure AWS recovery environments, ransomware resilience, and infrastructure recovery workflows. The position leads Infrastructure as Code adoption, disaster recovery testing, and cyber resilience engineering across cloud environments using Terraform, AWS, and automation tooling. Candidates will drive recovery architecture, restoration automation, and high-availability design in collaboration with engineering and SRE teams. 🗂️ Requirements: 8+ years in complex cloud environments, 3+ years of AWS experience, Strong Terraform expertise, Experience with Infrastructure as Code, Hands-on knowledge of Secure Vault patterns, Advanced shell scripting skills, Proficiency in Python or Go, Experience with CI/CD tooling, Ability to automate end-to-end restoration workflows 📃 Skills: AWS, GCP, Azure, Terraform, EKS, Kubernetes, Python, Go, Shell, CI/CD, Scalr, GitHub, IAM, VPC, API, CLI 🏢 Description: We’re looking for a Staff Cyber Resilience Engineer to lead defense against ransomware, destructive wipes, and large-scale data loss. This is a hands-on technical leadership role focused on designing and engineering an Isolated Recovery Environment, advancing Infrastructure as Code practices, and ensuring rapid restoration of AWS operations after compromise. How You'll Contribute: Own Our Recovery Architecture - Design and build an Isolated Recovery Environment — a hardened AWS account with immutable vaults that prevent attackers from reaching critical data - Threat model cloud-native attack patterns including IAM privilege escalation, backup deletion, ransomware persistence, and lateral movement across accounts - Validate and continuously improve backup configurations to ensure recoverability Standardize and Automate Infrastructure - Lead the transition to 100% Infrastructure as Code with Terraform for all assets including VPCs, IAM roles, and security groups - Build automated recovery workflows to tear down compromised environments and bootstrap hardened replacements from verified code and clean data - Write and maintain executable recovery playbooks with tested and versioned API calls and CLI commands Validate, Test, and Lead Exercises - Develop automated scripts using Python or Go to validate recovered data integrity - Lead hands-on recovery drills simulating total environment loss and recovery into a clean secondary account - Drive after-action reviews and continuous improvements Drive Engineering Standards - Act as the resilience authority for engineering teams - Shape high-availability architecture decisions and influence design reviews - Partner with Site Reliability Engineering on multi-region deployments and resilient architecture - Champion Infrastructure as Code and immutable infrastructure practices across teams What You'll Bring: - 8+ years of experience in complex cloud environments including AWS, GCP, or Azure - At least 3 years of AWS experience - Strong Terraform skills with modular environment design expertise - Familiarity with Secure Vault patterns in AWS - Advanced shell scripting skills and proficiency in Python or Go - Experience with CI/CD tooling such as Scalr or GitHub Actions - Proven ability to automate end-to-end restoration workflows Preferred Qualifications: - Experience leading recovery efforts after cyber attacks or destructive incidents - Experience with chaos engineering tools - Familiarity with NIST SP 800-34 or similar contingency planning frameworks - AWS Security Specialty certification or equivalent expertise Benefits: - Base salary range of $205,000–$233,000 annually plus bonus - 401(k) match - Medical, dental, and vision insurance - Life and disability insurance - Generous paid time off including vacation, sick leave, holidays, maternity, and bonding leave - Employee assistance and wellbeing resources Xometry is an equal opportunity employer. For U.S.-based roles, participation in E-Verify is required after acceptance of an offer.
Technology

Relativity
Senior Engineer - Site Reliability Engineering
Senior
Remote
Krakow, Poland
208,000 - 312,000 PLN/yr
🏢 Summary: Remote Senior Software Engineer – SRE role focused on building and maintaining highly available, scalable, and observable cloud-native systems. The position emphasizes automation, CI/CD improvements, incident management, and implementation of reliability best practices across SaaS platforms. The engineer collaborates cross-functionally to enhance system resilience, performance, and operational excellence. 🗂️ Requirements: 5+ years in Software Engineering, SRE, or Cloud Infrastructure roles, Experience with DevOps tools and practices, Proficiency in Python, Go, Java, C#, or .Net, Experience with at least two: GitHub, Azure DevOps, GitLab, Jenkins, Hands-on experience with observability tools, Strong experience with CI/CD pipelines and automation, Experience with cloud-native distributed systems, Experience in high-availability SaaS environments, Knowledge of SLOs, SLIs, and error budgets, Experience with redundancy and disaster recovery, Participation in on-call rotations 📃 Skills: Python, Go, Java, C#, .Net, GitHub, Azure, GitLab, Jenkins, Prometheus, Grafana, OpenTelemetry, CI/CD, DevOps, SLO, SLI, SaaS, Automation, Cloud, Agile 🏢 Description: Posting Type Remote Job Overview As the Senior Software Engineer – SRE you will focus on implementing and maintaining reliability solutions across the platform. This role emphasizes hands-on engineering work, automation, and operational excellence. The Senior Software Engineer will work closely with other engineers to ensure systems are highly available, observable, and resilient. As a member of the engineering team, the Senior Software Engineer will work closely with Infrastructure, Engineering, and Product teams to develop highly resilient, observable, and automated solutions that enhance system availability and efficiency. The ideal candidate will bring deep technical expertise, strong problem-solving skills, and a passion for reliability engineering. Job Description and Requirements Job Responsibilities Implement, and advocate for best-in-class reliability, observability, and scalability practices across the platform. Develop automated solutions for system reliability, capacity planning, and incident response to minimize manual intervention. Participate in improving Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to enhance system reliability. Contribute to CI/CD pipeline improvements and DevOps practices. Support root cause analysis (RCA) investigations, drive corrective actions, and advocate for a blameless postmortem culture. Participate in on-call rotations to ensure 24/7 availability of critical systems. Influence and mentor engineering teams on SRE principles, DevOps culture, and best practices. Stay ahead of industry trends, adopting new tools, frameworks, and methodologies to continually improve system reliability. Preferred Qualifications 5+ years of experience in software engineering, site reliability engineering, or cloud infrastructure roles. Experience with DevOps tooling and practices. Proficient in building service-oriented architectures and cloud-native distributed systems. Proficiency in programming languages such as Python, Go, Java, or C# or .Net. In-depth technical understanding and experience with at least two of the following DevOps platforms: GitHub, Azure DevOps, GitLab, or Jenkins. Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry or others). Strong background in CI/CD pipelines, automation, and DevOps practices. Experience working in global, high-availability SaaS environments. Experience implementing redundancy and disaster recovery scenarios. Excellent teamwork and cross-group collaboration skills. Ability to collaborate with both technical and business professionals. Hands-on experience with Agile Project Development Methodologies. Experience delivering complex technical solutions. Excellent problem-solving, analytical, and communication skills. Nice to have: Experience with Chaos Engineering and/or AI Ops . Competencies and Skills Automation-First Mindset – Commitment to reducing toil through scripting and automation. Reliability Engineering – Expertise in SLOs, SLIs, error budgets, and high-availability architectures. Incident Management & Postmortems – Experience in handling production incidents and driving continuous improvement. Observability & Monitoring – Deep understanding of logging, monitoring, and alerting best practices. Practical knowledge of data structures and modern data engines. Collaboration & Communication – Ability to work across teams, influence stakeholders, and advocate for reliability improvements. Mentorship & Coaching – Passion for mentoring engineers and building an SRE culture within the organization. Additional Information This role offers a unique opportunity to shape the future of SRE in a cutting-edge SaaS company, ensuring the reliability and scalability of mission-critical applications for customers worldwide. If you are passionate about solving complex reliability challenges and driving technical excellence, we’d love to hear from you! Relativity is a diverse workplace with different skills and life experiences—and we love and celebrate those differences. We believe that employees are happiest when they're empowered to be their full, authentic selves, regardless how you identify. Benefit Highlights: Comprehensive health, dental, and vision plans Parental leave for primary and secondary caregivers Flexible work arrangements Two, week-long company breaks per year Additional time off Long-term incentive program Training investment program All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law. Relativity is committed to competitive, fair, and equitable compensation practices. This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives. The expected salary range for this role is between following values: 208 000 and 312 000PLN The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position. Required Skills: Automation, Data Analysis, Database Management, Network Architecture, Performance Optimizations, Problem Solving, Project Management, Software Development, System Designs, Technical Leadership
Technology
Yard Corporate
Site Reliability Engineer (SRE)
Senior
Hybrid
Warsaw, Poland
40,000 - 55,000 PLN
🏢 Summary: Senior Site Reliability Engineer role focused on building and standardizing SRE practices across a hybrid AWS and on-prem infrastructure. The position centers on ensuring scalability, resilience, and high availability of high-frequency, data-intensive platforms through observability, automation, and Kubernetes optimization. You will define SLOs, enhance monitoring architecture, and drive reliability culture across engineering teams. 🗂️ Requirements: 5+ years experience in SRE, DevOps, or Infrastructure Engineering supporting distributed production systems, Bachelor’s degree in Computer Science, Computer Engineering, or related field (or equivalent experience), Deep expertise in Grafana, Prometheus, Loki, and Tempo (OpenTelemetry), Strong production experience with Docker and Kubernetes, Experience managing hybrid infrastructure (AWS and on-premises), Proficiency in at least one language: Python, Go, or Bash, Hands-on experience with CI/CD pipelines and Infrastructure-as-Code, Experience defining and managing SLOs and SLAs, Willingness to participate in on-call rotation 📃 Skills: AWS, Kubernetes, Docker, Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Python, Go, Bash, CI/CD, IaC, Git, Hypervisors 🏢 Description: About the Client Our client is a premier, global investment management firm operating at the intersection of finance and technology. Known for their sophisticated, data-intensive systems, they build and maintain high-performance platforms that process massive volumes of market and operational data. To support their expanding footprint, they are looking for a senior-level Site Reliability Engineer (SRE) who will take ownership of shaping, standardizing, and scaling their SRE frameworks and reliability culture from the ground up. The Role In this role, you will serve as a foundational force for SRE practices, partnering directly with Cloud, Infrastructure, and Software Engineering squads. You will work across a hybrid infrastructure (combining advanced AWS cloud environments and physical on-premises servers) to guarantee the scalability, resilience, and maximum uptime of critical, high-frequency transactional platforms. Core Responsibilities SRE Evangelism: Design, implement, and champion core reliability principles, helping technology teams adopt sustainable scaling practices. Observability Architecture: Implement, scale, and maintain end-to-end monitoring, telemetry, and distributed tracing systems utilizing Prometheus, Grafana, Loki, and Tempo (OpenTelemetry framework). Kubernetes Optimization: Establish best-practice configurations for containerized workloads, ensuring applications running on Kubernetes are highly resilient, cost-effective, and performant. Incident Management & Culture: Participate in a balanced, shared on-call rotation (averaging one week per month). Automation & Engineering: Build custom tooling and CI/CD pipelines to automate routine tasks, system health checks, and rapid disaster recovery workflows. SLO/SLA Definition: Partner with product and engineering teams to define, monitor, and enforce Service Level Objectives (SLOs) and Error Budgets. What We Look For Experience: 5+ years of hands-on experience in a dedicated SRE, DevOps, or Infrastructure Engineering role supporting complex, distributed production systems. Education: A Bachelor’s degree in Computer Science, Computer Engineering, or a related technical discipline (or equivalent practical experience). Observability Expertise: Deep, subject-matter knowledge of modern monitoring stacks, specifically Grafana, Prometheus, Loki, and Tempo (OTel). Orchestration & Containers: Strong, production-grade expertise in containerization (Docker) and orchestration (Kubernetes). Hybrid Infrastructure: Experience navigating hybrid models—managing both cloud services (AWS preferred) and physical on-premise hardware resources. Scripting/Coding: Proficiency in writing clean, maintainable code in at least one scripting or programming language (e.g., Python, Bash, or Go) to build reliable automation. Methodologies: Solid grounding in CI/CD concepts, infrastructure-as-code (IaC), and agile development processes. Soft Skills: Excellent verbal and written communication skills, with a proven ability to convey complex infrastructure and reliability concepts to both technical and non-technical stakeholders. What We Offer Stable Employment: Full-time employment contract ( Umowa o Pracę - UoP ). Tax Optimization: Eligibility for creative tax-deductible costs ( KUP - Koszty Uzyskania Przychodu). Financial Reward: Highly competitive base salary accompanied by a generous annual performance bonus . Comprehensive Health: Premium private medical care package that fully includes dental coverage (stomatologia) . Wellness & Lifestyle: MultiSport card to keep you active and healthy. Daily Perks: Pre-funded lunch card for your daily meals. Tech Stack at a Glance Cloud & Virtualization: AWS, Kubernetes, Docker, On-Premises Hypervisors Observability: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (OTel) Languages: Python, Go, Bash CI/CD & Automation: Git-based pipelines, Configuration Management, IaC
Technology

Relativity
Senior Engineer - Site Reliability Engineering
Senior
Remote
Krakow, Poland
208,000 - 312,000 PLN/yr
🏢 Summary: Senior Software Engineer – SRE role focused on building and maintaining highly available, observable, and resilient cloud-native systems. The position emphasizes automation, CI/CD improvements, incident management, and implementation of reliability best practices across a SaaS platform. You will collaborate cross-functionally to enhance scalability, performance, and operational excellence. 🗂️ Requirements: 5+ years in Software Engineering, SRE, or Cloud Infrastructure, Experience with DevOps tools and practices, Proficiency in Python, Go, Java, or C#/.NET, Experience with at least two: GitHub, Azure DevOps, GitLab, Jenkins, Hands-on experience with observability tools, Strong experience with CI/CD pipelines and automation, Experience with cloud-native distributed systems, Experience in high-availability SaaS environments, Knowledge of SLOs, SLIs, and error budgets, Experience with incident management and root cause analysis, Experience implementing redundancy and disaster recovery, Experience with Agile methodologies 📃 Skills: Python, Go, Java, C#, DotNet, GitHub, Azure, GitLab, Jenkins, Prometheus, Grafana, OpenTelemetry, CI/CD, DevOps, SLO, SLI, SaaS, Automation, Cloud, DistributedSystems 🏢 Description: Job Overview As the Senior Software Engineer – SRE you will focus on implementing and maintaining reliability solutions across the platform. This role emphasizes hands-on engineering work, automation, and operational excellence. The Senior Software Engineer will work closely with other engineers to ensure systems are highly available, observable, and resilient. As a member of the engineering team, the Senior Software Engineer will work closely with Infrastructure, Engineering, and Product teams to develop highly resilient, observable, and automated solutions that enhance system availability and efficiency. The ideal candidate will bring deep technical expertise, strong problem-solving skills, and a passion for reliability engineering. Job Description and Requirements Job Responsibilities Implement, and advocate for best-in-class reliability, observability, and scalability practices across the platform. Develop automated solutions for system reliability, capacity planning, and incident response to minimize manual intervention. Participate in improving Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to enhance system reliability. Contribute to CI/CD pipeline improvements and DevOps practices. Support root cause analysis (RCA) investigations, drive corrective actions, and advocate for a blameless postmortem culture. Participate in on-call rotations to ensure 24/7 availability of critical systems. Influence and mentor engineering teams on SRE principles, DevOps culture, and best practices. Stay ahead of industry trends, adopting new tools, frameworks, and methodologies to continually improve system reliability. Preferred Qualifications 5+ years of experience in software engineering, site reliability engineering, or cloud infrastructure roles. Experience with DevOps tooling and practices. Proficient in building service-oriented architectures and cloud-native distributed systems. Proficiency in programming languages such as Python, Go, Java, or C# or .Net. In-depth technical understanding and experience with at least two of the following DevOps platforms: GitHub, Azure DevOps, GitLab, or Jenkins. Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry or others). Strong background in CI/CD pipelines, automation, and DevOps practices. Experience working in global, high-availability SaaS environments. Experience implementing redundancy and disaster recovery scenarios. Excellent teamwork and cross-group collaboration skills. Ability to collaborate with both technical and business professionals. Hands-on experience with Agile Project Development Methodologies. Experience delivering complex technical solutions. Excellent problem-solving, analytical, and communication skills. Nice to have: Experience with Chaos Engineering and/or AI Ops . Competencies and Skills Automation-First Mindset – Commitment to reducing toil through scripting and automation. Reliability Engineering – Expertise in SLOs, SLIs, error budgets, and high-availability architectures. Incident Management & Postmortems – Experience in handling production incidents and driving continuous improvement. Observability & Monitoring – Deep understanding of logging, monitoring, and alerting best practices. Practical knowledge of data structures and modern data engines. Collaboration & Communication – Ability to work across teams, influence stakeholders, and advocate for reliability improvements. Mentorship & Coaching – Passion for mentoring engineers and building an SRE culture within the organization. Additional Information This role offers a unique opportunity to shape the future of SRE in a cutting-edge SaaS company, ensuring the reliability and scalability of mission-critical applications for customers worldwide. If you are passionate about solving complex reliability challenges and driving technical excellence, we’d love to hear from you! Relativity is a diverse workplace with different skills and life experiences—and we love and celebrate those differences. We believe that employees are happiest when they're empowered to be their full, authentic selves, regardless how you identify. Benefit Highlights: Comprehensive health, dental, and vision plans Parental leave for primary and secondary caregivers Flexible work arrangements Two, week-long company breaks per year Additional time off Long-term incentive program Training investment program All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law. Relativity is committed to competitive, fair, and equitable compensation practices. This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives. The expected salary range for this role is between following values: 208 000 and 312 000PLN The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position. Required Skills: Automation, Data Analysis, Database Management, Network Architecture, Performance Optimizations, Problem Solving, Project Management, Software Development, System Designs, Technical Leadership