April 29, 2026
Senior DevOps Engineer
Senior • On-site
Seattle, WA
About Hashgraph:
Hashgraph is a fast-growing software company committed to supporting, developing and servicing Hedera, an open source, proof-of-stake platform. Hedera is EVM-compatible and has been specifically built to meet the needs of enterprise and web3 applications, which require speed, security, stability and sustainability. Hedera's public network is governed by industry-leading organizations, spanning 11 sectors and 14 regions who oversee the development and direction of the decentralized platform.
The role:
We are hiring a Senior DevOps Engineer (Node Operations) to ensure the reliability, security, and operational excellence of Hedera consensus node environments. This role exists to reduce operational toil, strengthen infrastructure automation, and improve release and preproduction readiness across a globally distributed network. Without this role, we risk increased availability incidents, slower recovery times, and delays in delivering against the product roadmap.
In this role, you will:
- Operate and improve Hedera consensus node environments across testnet, previewnet, and preproduction
- Design and implement automation-first workflows for release and preproduction environments
- Build and maintain Infrastructure-as-Code (Terraform) on GCP
- Improve change management, release safety, and operational predictability
- Participate in on-call rotation, incident response, and RCA, driving corrective actions into automation
- Partner with internal engineering teams and external stakeholders, including Hedera Governing Council members, to support operational requirements
What success looks like in 6-12 months:
- Operational toil is significantly reduced through durable automation and standardization
- Node environments are more reliable, with fewer incidents and faster recovery times
- Release and preproduction workflows are predictable, repeatable, and automated
- Infrastructure changes are consistent, testable, and auditable through IaC best practices
What you bring:
Core capabilities:
- Strong systems reliability mindset with experience in incident response and RCA
- Proven ability to automate operational workflows and reduce manual toil
- Clear communicator with the ability to work across engineering, security, and external partners
- Deep ownership mentality with a bias toward preventative engineering over reactive fixes
- Strong Linux and networking troubleshooting in production environments
Functional expertise:
- Infrastructure-as-Code with Terraform (module design, state management)
- Configuration management with Ansible
- CI/CD automation (Jenkins or equivalent pipeline tooling)
- Experience operating distributed systems or production infrastructure at scale
- Familiarity with Kubernetes fundamentals
Nice to haves:
- Observability stacks (e.g., Grafana, Loki, Tempo, Mimir)
- Programming/scripting (Go, Python, Bash)
- GitHub/GitHub Actions experience
Similar jobs you might like
Technology

Hashgraph
Senior Site Reliability Engineer - Azure
Senior
On-site
San Francisco, CA
🏢 Summary: Senior Site Reliability Engineer (Azure) role focused on building and scaling secure, production-grade Azure infrastructure for a private distributed ledger network. The position centers on designing infrastructure from scratch, implementing infrastructure as code, and ensuring high availability, reliability, and enterprise readiness. The engineer will support complex deployments and operational excellence in a distributed systems environment. 🗂️ Requirements: Proven experience designing and building production-grade systems on Azure, Strong expertise in Azure cloud services: networking, compute, identity, security, storage, Hands-on experience with Terraform for infrastructure as code at production scale, Programming experience in Go or Python, Experience with distributed systems and high-availability architectures, Experience with CI/CD and infrastructure automation tooling, Ability to design greenfield infrastructure environments 📃 Skills: Azure, Terraform, Go, Python, CI/CD, Kubernetes, Prometheus, Grafana, Argo, Spacelift 🏢 Description: About Hashgraph: Hashgraph is a fast-growing software company committed to supporting, developing and servicing Hedera, an open source, proof-of-stake platform. Hedera is EVM-compatible and has been specifically built to meet the needs of enterprise and web3 applications, which require speed, security, stability and sustainability. Hedera's public network is governed by industry-leading organizations, spanning 11 sectors and 14 regions who oversee the development and direction of the decentralized platform.The role:We are hiring a Senior Site Reliability Engineer (Azure) to build and scale the Azure infrastructure foundation for HashSphere, a new private DLT network harnessing Hedera's institutional grade technology, being built by a passionate team of industry leaders.This role exists to ensure that our platform can operate as a secure, scalable, and production-ready system in Azure, supporting complex enterprise use cases and high reliability expectations.The impact you'll have: In this role, you will: Design and build secure, scalable Azure infrastructure from first principles for a production-grade distributed system Develop and own Terraform-based infrastructure as code, enabling repeatable and automated deployments Translate product and customer requirements into technical architecture and execution plans Build and enhance platform services, APIs, and integrations that extend HashSphere capabilities Partner across engineering, security, and product teams to deliver enterprise-ready infrastructure solutions Contribute to operational excellence, including reliability, observability, and incident response Support customer deployments and production environments through Tier 2 infrastructure support What success looks like in 6-12 months: Azure is a production-ready deployment environment for HashSphere Customer deployments are repeatable, scalable, and secure Azure achieves feature parity with other supported cloud environments What you bring: Core capabilities: Proven experience designing and building production-grade systems on Azure Ability to take ambiguous requirements to structured technical solutions to delivered systems Strong technical communication skills across engineering and non-technical stakeholders High ownership mindset with a bias for action and accountability Collaborative approach with a focus on building durable, scalable solutions Functional expertise: Azure cloud services (networking, compute, identity, security, storage) Terraform (infrastructure as code at production scale) Programming experience in Go and/or Python Experience building greenfield infrastructure environments Distributed systems, high-availability architectures, or platform engineering CI/CD and automation tooling for infrastructure lifecycle management Nice to haves: Kubernetes and container orchestration Observability tooling (Prometheus, Grafana) Workflow/orchestration platforms (Argo, Spacelift, or similar)
Technology

Hashgraph
Senior Site Reliability Engineer - Azure
Senior
Remote
🏢 Summary: Senior Site Reliability Engineer (Azure) role focused on designing and building secure, scalable Azure infrastructure for a production-grade distributed ledger platform. The position centers on infrastructure as code, automation, and high-availability architecture to support enterprise deployments. You will ensure reliability, scalability, and operational excellence for customer-facing environments. 🗂️ Requirements: Proven experience designing and operating production-grade systems on Azure, Strong hands-on experience with Terraform for infrastructure as code, Programming experience in Go or Python, Experience with distributed systems or high-availability architectures, Experience building greenfield infrastructure environments, Experience with CI/CD and infrastructure automation tooling, Ability to translate requirements into technical architecture and implementation, Experience supporting production environments and incident response 📃 Skills: Azure, Terraform, Go, Python, DistributedSystems, HighAvailability, CI/CD, Networking, Compute, Identity, Security, Storage, Automation, APIs, IncidentResponse, Infrastructure 🏢 Description: About Hashgraph: Hashgraph is a fast-growing software company committed to supporting, developing and servicing Hedera, an open source, proof-of-stake platform. Hedera is EVM-compatible and has been specifically built to meet the needs of enterprise and web3 applications, which require speed, security, stability and sustainability. Hedera's public network is governed by industry-leading organizations, spanning 11 sectors and 14 regions who oversee the development and direction of the decentralized platform.The role:We are hiring a Senior Site Reliability Engineer (Azure) to build and scale the Azure infrastructure foundation for HashSphere, a new private DLT network harnessing Hedera's institutional grade technology, being built by a passionate team of industry leaders.This role exists to ensure that our platform can operate as a secure, scalable, and production-ready system in Azure, supporting complex enterprise use cases and high reliability expectations.The impact you'll have: In this role, you will: Design and build secure, scalable Azure infrastructure from first principles for a production-grade distributed system Develop and own Terraform-based infrastructure as code, enabling repeatable and automated deployments Translate product and customer requirements into technical architecture and execution plans Build and enhance platform services, APIs, and integrations that extend HashSphere capabilities Partner across engineering, security, and product teams to deliver enterprise-ready infrastructure solutions Contribute to operational excellence, including reliability, observability, and incident response Support customer deployments and production environments through Tier 2 infrastructure support What success looks like in 6-12 months: Azure is a production-ready deployment environment for HashSphere Customer deployments are repeatable, scalable, and secure Azure achieves feature parity with other supported cloud environments What you bring: Core capabilities: Proven experience designing and building production-grade systems on Azure Ability to take ambiguous requirements to structured technical solutions to delivered systems Strong technical communication skills across engineering and non-technical stakeholders High ownership mindset with a bias for action and accountability Collaborative approach with a focus on building durable, scalable solutions Functional expertise: Azure cloud services (networking, compute, identity, security, storage) Terraform (infrastructure as code at production scale) Programming experience in Go and/or Python Experience building greenfield infrastructure environments Distributed systems, high-availability architectures, or platform engineering CI/CD and automation tooling for infrastructure lifecycle management Nice to haves: Kubernetes and container orchestration Observability tooling (Prometheus, Grafana) Workflow/orchestration platforms (Argo, Spacelift, or similar)
Technology

Hashgraph
Staff Software Engineer - CI/CD & Release Engineering
Senior
On-site
New York City, NY
🏢 Summary: Staff Software Engineer role focused on architecting and scaling CI/CD and release engineering systems to support complex, multi-product software delivery. The position centers on building reliable, automated, and observable release infrastructure using cloud-native and Kubernetes-based environments. It aims to improve deployment speed, safety, and operational excellence across internal and open-source platforms. 🗂️ Requirements: 7+ years building and operating production-grade CI/CD or engineering infrastructure, Deep experience with CI/CD systems (GitHub Actions or GitLab CI), Strong experience with Kubernetes and Docker, Strong programming skills in Kotlin, Java, Go, Python, or Bash, Experience building automated build, test, release, and deployment pipelines, Experience with Gradle, Maven, or NPM, Experience operating infrastructure in AWS or Google Cloud Platform, Experience implementing observability and monitoring solutions, Strong understanding of release management and deployment reliability practices 📃 Skills: Kubernetes, Docker, GitHubActions, GitLabCI, Kotlin, Java, Go, Python, Bash, Gradle, Maven, NPM, AWS, GCP, Grafana 🏢 Description: About Hashgraph: Hashgraph is a fast-growing software company committed to supporting, developing and servicing Hedera, an open source, proof-of-stake platform. Hedera is EVM-compatible and has been specifically built to meet the needs of enterprise and web3 applications, which require speed, security, stability and sustainability. Hedera's public network is governed by industry-leading organizations, spanning 11 sectors and 14 regions who oversee the development and direction of the decentralized platform.The role: We are hiring a Staff Software Engineer - CI/CD & Release Engineering to architect and scale the systems that power software delivery across Hashgraph's internal and open-source platforms. This role exists because our engineering organization is operating increasingly complex distributed systems across multiple products, environments, and release streams. We need a senior technical builder who can design reliable, scalable, and developer-friendly release infrastructure that enables engineering teams to ship quickly, safely, and repeatedly at production scale. You will help establish the foundation for how software is built, tested, released, deployed, and observed across the company — improving engineering velocity while strengthening reliability and operational confidence. In this role, you will: Architect and evolve scalable CI/CD systems that support complex multi-product release workflows across internal and open-source platforms Design and build developer tooling, deployment automation, and release orchestration systems using technologies such as GitHub Actions, Kubernetes, and cloud-native infrastructure Lead the engineering strategy for build pipelines, artifact management, release governance, and deployment reliability Improve developer experience by reducing friction in build, test, deployment, and operational workflows Build and maintain highly reliable Kubernetes-based infrastructure powering release automation and engineering productivity systems Drive observability and operational excellence across release systems through metrics, monitoring, and performance instrumentation Partner closely with Platform Engineering, DevOps, Security, Program Management, and Product Engineering teams to align release infrastructure with business and technical priorities Serve as a senior technical leader and multiplier within the organization through architecture guidance, mentorship, and operational rigor What success looks like in 6-12 months: CI/CD systems are highly reliable, scalable, and capable of supporting rapid multi-product releases with minimal operational overhead Engineering teams ship software faster with improved deployment confidence and reduced release friction Build and release infrastructure is standardized, observable, and repeatable across products and environments Deployment automation significantly reduces manual operational effort and release risk Release engineering becomes a strategic enabler for engineering throughput, platform reliability, and product delivery velocity What you bring: Core capabilities: Strong systems thinking with the ability to design scalable engineering infrastructure from first principles Deep ownership mentality with a bias toward automation, reliability, and operational excellence Strong communication and cross-functional collaboration skills across engineering and business stakeholders Ability to balance strategic architectural thinking with hands-on execution Proven ability to operate effectively in fast-moving, high-autonomy engineering environments Functional expertise: 7+ years building and operating production-grade software engineering infrastructure or CI/CD platforms Deep experience designing and operating CI/CD systems using tools such as GitHub Actions and/or GitLab CI Strong experience with Kubernetes, Docker, and cloud-native infrastructure patterns Strong programming and automation experience in one or more of: Kotlin, Java, Go, Python, or Bash Experience building scalable build, test, release, and deployment automation systems Experience with modern build systems and dependency management tools such as Gradle, Maven, or NPM Experience operating engineering infrastructure within AWS and/or Google Cloud Platform environments Experience designing observability and monitoring solutions using tools such as Grafana and related telemetry systems Strong understanding of release management strategy, deployment safety, and operational reliability practices Experience writing high-quality technical documentation, engineering standards, and operational processes Nice to haves: Experience supporting large-scale open-source software development workflows Experience with distributed systems, blockchain, or decentralized infrastructure technologies Experience optimizing build performance, release throughput, or engineering productivity at scale Expertise in Golang and advanced shell automation
Technology
ALTER GPU CENTER
DevOps Engineer
Mid
Remote
Łódź, Poland
🏢 Summary: Hands-on DevOps Engineer role focused on building and operating automation, deployment, and reliability standards for large-scale GPU infrastructure supporting AI training and inference. The position involves Infrastructure as Code, CI/CD, observability, security, and low-level automation across bare-metal servers, networking, storage, and Kubernetes-based platforms. The role emphasizes reliability, scalability, and automation in complex, high-performance environments. 🗂️ Requirements: 4–7 years in DevOps, SRE, or Platform Engineering, Experience with infrastructure automation in production environments, Hands-on experience with Terraform or Ansible, Experience building and maintaining CI/CD pipelines, Knowledge of GitOps practices, Understanding of infrastructure security and vulnerability management, Experience with security tools (e.g., Snyk, CrowdStrike), Practical experience with Kubernetes, Experience with GPU technologies (e.g., NVIDIA GPU Operator, MIG), Scripting or programming skills in Python, Go, or Bash, Experience with bare-metal provisioning or low-level infrastructure automation, Knowledge of observability tools (Prometheus, Grafana, Loki, OpenTelemetry) 📃 Skills: Terraform, Ansible, Kubernetes, Python, Go, Bash, Prometheus, Grafana, Loki, OpenTelemetry, Snyk, CrowdStrike, NVIDIA, MIG, CI/CD, GitOps 🏢 Description: About the role We are looking for a DevOps Engineer to help build and operate automation, deployment, and reliability standards for large-scale GPU infrastructure used for AI training and inference workloads. In this role, you will work on software-defined infrastructure supporting GPU clusters, high-performance networking, storage platforms, and internal AI services. This is a hands-on position for someone who is comfortable working close to infrastructure, improving operational processes, and building reliable automation in a complex technical environment. Responsibilities Design, implement, and maintain Infrastructure as Code solutions for provisioning and managing bare-metal GPU servers, networking, storage, and cluster orchestration components Build and improve CI/CD pipelines for infrastructure, platform services, and internal tooling Develop and maintain monitoring, logging, alerting, and observability solutions for large-scale GPU environments Support reliability initiatives by defining and tracking SLIs/SLOs , automating incident response, and contributing to post-incident analysis Automate operational tasks such as cluster scaling, firmware and BIOS updates, hardware validation, diagnostics, and capacity planning Work closely with Infrastructure, Networking, Facilities, and AI/ML teams to ensure stable and scalable platform operations Support DevSecOps practices, including infrastructure hardening, vulnerability management, and compliance automation Identify repetitive manual work and replace it with efficient automation Evaluate new tools and solutions related to GPU infrastructure, orchestration, and cloud-native operations Requirements 4–7 years of experience in DevOps, SRE, Platform Engineering , or a similar role Strong practical experience with infrastructure automation in complex production environments Good hands-on knowledge of Terraform, Ansible , or similar Infrastructure as Code tools Experience building and maintaining CI/CD pipelines and working with GitOps practices Good understanding of infrastructure security, vulnerability management, and security best practices Experience with security tools such as Snyk, CrowdStrike , or similar solutions Practical experience with Kubernetes Experience working with GPU-related technologies such as NVIDIA GPU Operator, device plugins, MIG, or time-slicing Good scripting or programming skills in Python, Go, or Bash Experience with bare-metal provisioning, low-level infrastructure automation, or data center operations Good knowledge of observability tools such as Prometheus, Grafana, Loki, and OpenTelemetry Ability to work independently, prioritize tasks, and communicate effectively with technical teams English proficiency at least at a communicative level is required, as you will be working in an international team Nice to have Experience in AI infrastructure, HPC environments, hyperscale infrastructure, or data center operations Familiarity with orchestration and scheduling tools such as Slurm, Ray, Run:ai, KServe , or Kubernetes-based schedulers Experience integrating telemetry from power, cooling, or environmental systems Experience building internal platforms or self-service tools for engineering teams Understanding of compliance and audit requirements in security-sensitive environments What we offer Benefits package Opportunity to work on advanced infrastructure supporting large-scale AI workloads Real impact on the reliability and scalability of next-generation compute environments Collaboration with experienced engineers across infrastructure, platform, and AI domains A fast-moving environment with space for ownership, technical input, and professional growth
Technology
Grid Dynamics Poland
Senior DevOps Engineer
Senior
Hybrid
Warsaw, Poland
🏢 Summary: Senior DevOps Engineer role focused on ensuring production stability and driving automation across multi-cloud environments. The position bridges infrastructure and application support, emphasizing CI/CD optimization, incident management, and elimination of manual operational tasks. Responsibilities include managing deployments, patching, and building automation tools to enhance system reliability. 🗂️ Requirements: Commercial experience with AWS, Azure, or GCP, Strong proficiency in Python, Strong proficiency in Java, Hands-on experience designing and maintaining CI/CD pipelines, Experience with Jenkins, GitHub Actions, GitLab CI, or Azure DevOps, Practical knowledge of container ecosystems, Experience in incident management and production monitoring, Experience executing system deployments and infrastructure patching 📃 Skills: AWS, Azure, GCP, Python, Java, Jenkins, GitHubActions, GitLabCI, AzureDevOps, CI/CD, Containers, Monitoring, Automation, DevOps 🏢 Description: We are looking for a highly skilled Senior DevOps Engineer to take ownership of production stability and drive the aggressive automation of support operations. In this role, you will bridge the gap between infrastructure and application support, ensuring the reliability of multi-cloud environments while eliminating manual operational toil. Essential functions Production Operations & Support: Act as the first line of defense for system reliability. Monitor system health, manage incidents, and troubleshoot complex production issues. Support Automation: Write robust code and scripts to automate routine support activities, system patching, and deployment processes. CI/CD Management: Build, maintain, and optimize deployment pipelines across various tools (Jenkins, GitHub Actions, GitLab CI, Azure DevOps). Deployments & Patching: Execute seamless software deployments and infrastructure patching with minimal downtime. Qualifications Cloud Platforms: Commercial experience managing infrastructure in AWS, Azure, or GCP . Programming Skills: Strong proficiency in both Python and Java (essential for building automation tools and supporting the core application stack). CI/CD Mastery: Solid hands-on experience designing and maintaining pipelines (Jenkins, GitHub Actions, GitLab CI, or Azure DevOps). Containerization: Practical working knowledge of container ecosystems. Ops & Reliability: Deep understanding of incident management, production monitoring, and executing system deployments. We offer Opportunity to work on bleeding-edge projects Work with a highly motivated and dedicated team Competitive salary Flexible schedule Benefits package - medical insurance, sports Corporate social events Professional development opportunities Well-equipped office About us Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI , supported by profound expertise and ongoing investment in data , analytics , cloud & DevOps , application modernization and customer experience . Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.
Technology
ALTER GPU CENTER
Lead DevOps Engineer
Senior
Remote
Łódź, Poland
🏢 Summary: Technical leadership role combining hands-on DevOps/SRE engineering with team management to build and operate large-scale GPU infrastructure for AI workloads. Focused on infrastructure automation, reliability, observability, and high-performance networking across complex production environments. Responsible for shaping IaC standards, CI/CD, and operational excellence for software-defined, GPU-based platforms. 🗂️ Requirements: 8+ years in DevOps, SRE, or Platform Engineering, 3+ years in technical leadership role, Experience with large-scale infrastructure automation, Proficiency in Infrastructure as Code tools, Experience with GitOps and CI/CD, Hands-on experience with Kubernetes, Experience with GPU technologies, Scripting or programming in Python, Go, or Bash, Experience with bare-metal provisioning, Knowledge of observability and monitoring tools, Understanding of distributed systems reliability, Experience with high-performance networking technologies, Ability to lead technical discussions and mentor engineers, English proficiency at communicative level 📃 Skills: Terraform, Ansible, Pulumi, Crossplane, GitOps, Kubernetes, NVIDIA, MIG, Python, Go, Bash, Prometheus, Grafana, Loki, OpenTelemetry, RDMA, InfiniBand, RoCE, CI/CD 🏢 Description: About the role We are looking for a Lead DevOps Engineer to provide technical leadership for DevOps and Site Reliability Engineering practices supporting large-scale GPU infrastructure used for AI training and inference workloads. This role combines hands-on engineering with team leadership. You will be responsible for shaping automation standards, improving platform reliability, and leading a team working on software-defined infrastructure, high-performance networking, observability, and operational excellence across complex production environments. Responsibilities Lead, mentor, and support a team of DevOps and SRE engineers working across the full lifecycle of GPU infrastructure platforms Design and implement Infrastructure as Code solutions for provisioning and managing bare-metal GPU servers, networking, storage, and cluster orchestration components Build and improve CI/CD pipelines for infrastructure, platform services, and internal tooling Develop and maintain monitoring, logging, alerting, and observability solutions for large-scale GPU environments Define and track SLIs/SLOs , improve incident response processes, and contribute to post-incident reviews and long-term reliability improvements Work closely with Infrastructure, Networking, Facilities, and AI/ML teams to ensure stable and scalable platform operations Automate operational processes such as cluster scaling, firmware and BIOS updates, hardware diagnostics, and capacity planning Support DevSecOps practices, including infrastructure hardening, vulnerability management, and compliance automation Identify operational inefficiencies and reduce repetitive manual work through automation Evaluate and introduce new tools and solutions related to GPU infrastructure, orchestration, and cloud-native operations Requirements 8+ years of experience in DevOps, SRE, Platform Engineering , or a similar area At least 3 years of experience in a technical lead, lead engineer, or team leadership role Strong practical experience with infrastructure automation in large-scale or complex production environments Very good knowledge of Terraform, Ansible, Pulumi, Crossplane , or similar Infrastructure as Code tools Experience with GitOps , configuration management, and CI/CD practices Hands-on experience with Kubernetes Experience working with GPU-related technologies such as NVIDIA GPU Operator, device plugins, MIG, or time-slicing Good scripting or programming skills in Python, Go, or Bash Experience with bare-metal provisioning, infrastructure automation, or data center environments Good knowledge of observability tools such as Prometheus, Grafana, Loki, and OpenTelemetry Good understanding of distributed systems reliability and production incident management Experience with high-performance networking technologies such as RDMA, InfiniBand, or RoCE will be a strong advantage Ability to lead technical discussions, support team development, and communicate effectively with both technical and business stakeholders English proficiency at least at a communicative level is required, as you will be working in an international team Nice to have Experience in AI infrastructure, HPC environments, hyperscale infrastructure, or data center operations Familiarity with orchestration and scheduling tools such as Slurm, Ray, Run:ai, KServe , or Kubernetes-based schedulers Experience integrating telemetry from power, cooling, or environmental systems Experience building internal platforms or self-service tools for engineering or research teams Understanding of security, compliance, and audit requirements in regulated or security-sensitive environments What we offer Benefits package Opportunity to shape the DevOps and SRE foundation for advanced GPU infrastructure supporting AI workloads Real impact on the scalability, reliability, and operational standards of next-generation compute environments Collaboration with experienced engineers across infrastructure, platform, and AI domains A dynamic environment with space for ownership, technical leadership, and professional growth
Technology
DCG
Senior DevOps Engineer (AI & Platform Operations)
Senior
Hybrid
Warsaw, Poland
120 - 125 PLN
🏢 Summary: The role focuses on ensuring stable and reliable AI and platform operations in production environments, with ownership of incident management, monitoring, and deployment oversight. The engineer works closely with development teams to diagnose issues, improve observability, and enhance service continuity through automation and operational excellence. This position emphasizes production support, RCA, and continuous improvement rather than building infrastructure from scratch. 🗂️ Requirements: 5+ years in IT operations or production support roles, Experience owning incidents end-to-end including RCA, Minimum 2 years working within ITIL framework, Experience in Agile delivery environments, Proficiency with log analysis and alerting tools, Experience with observability and monitoring tools, Hands-on experience supporting services on Kubernetes, Experience with CI/CD pipelines and deployment troubleshooting, Experience with relational databases and query analysis, Ability to trace issues across Java-based application stacks 📃 Skills: ITIL, Splunk, Apica, Sysdig, Prometheus, Grafana, Kubernetes, Jenkins, Oracle, DB2, Spring, Hibernate, Kafka, XML, JSON, Java, J2EE, Datastage, Bash, Python, Ansible 🏢 Description: As a recruitment company, DCG understands that every business is powered by experienced professionals. Our management style and partnership approach enable us to meet your needs and provide continuous support. Due to our ongoing growth and the large number of recruitment projects we undertake for our partners, we are currently looking for: Senior DevOps Engineer (AI & Platform Operations) Responsibilities: Incident & Problem Management: Own the RCA process for production incidents — diagnose, resolve, and put preventive measures in place so issues don't recur Production Monitoring & Support: Continuously monitor service health, detect anomalies early, and act before they become incidents Deployment Execution: Trigger and oversee release deployments through existing CI/CD pipelines; troubleshoot failed deployments and coordinate rollbacks when needed Environment Oversight: Keep Pre-Production and Production environments stable and aligned — not building them from scratch, but ensuring they behave as expected day to day Runbook & Knowledge Management: Document operational procedures, known issues, and resolution steps to build a reliable knowledge base for the team Cross-team Collaboration: Work shoulder-to-shoulder with development and platform teams to triage issues, clarify operational requirements, and close the feedback loop between prod and dev Identify recurring pain points and propose automation or tooling to reduce toil Improve observability coverage — dashboards, alerts, log queries — to catch issues faster Contribute to service continuity initiatives and disaster recovery drills Requirements: 5+ years in IT operations, application support (2nd/3rd line), or a similar production-facing role Proven track record of owning incidents end-to-end — from alert to RCA to prevention 2+ years working within an ITIL framework (incident, problem, change management) Experience working in Agile delivery environments alongside development teams Excellent English communication skills — able to explain technical issues clearly to both engineers and non-technical stakeholders Proficiency with log analysis and alerting tools: Splunk, Apica, Sysdig Observability tooling: Prometheus, Grafana — reading dashboards, tuning alerts Comfortable operating services running on Kubernetes (checking pod health, reading logs, triggering restarts — not cluster administration) Familiarity with Jenkins pipelines to execute and troubleshoot deployments Relational databases (Oracle, DB2) — querying, interpreting execution plans, identifying data-related incidents Working knowledge of Spring/Hibernate application behavior, Kafka message flows, XML/JSON payloads — enough to trace an issue through the stack Nice to have: Java/J2EE development background (helps enormously when reading stack traces and working with dev teams) IBM Datastage operational experience Scripting (Bash, Python) for automation of repetitive operational tasks Ansible for applying configuration changes in controlled operational scenarios Offer: Private medical care Co-financing for the sports card Constant support of dedicated consultant Employee referral program
Technology
Link Group
Senior Devops Engineer
Senior
Hybrid
Warsaw, Poland
28,000 - 38,000 PLN
🏢 Summary: Senior DevOps Engineer role focused on owning and evolving cloud-native infrastructure and CI/CD platforms that support large-scale data processing systems. The position combines hands-on engineering and strategic impact to ensure scalable, secure, and reliable production environments. You will design, automate, and optimize platform services enabling efficient delivery of data-driven applications. 🗂️ Requirements: 5+ years in DevOps, SRE, or infrastructure engineering, Experience supporting distributed production systems, Hands-on experience with public cloud platforms, Strong knowledge of containerization and orchestration, Experience with infrastructure as code, Strong scripting or programming skills, Experience building and maintaining CI/CD pipelines, Knowledge of observability practices and tools, Strong troubleshooting and incident response skills in Linux environments 📃 Skills: AWS, Docker, Kubernetes, Terraform, Python, Bash, CI/CD, Linux, Monitoring, Logging, Alerting 🏢 Description: Senior DevOps Engineer We are looking for an experienced engineer to take ownership of our infrastructure and platform ecosystem, supporting large-scale data processing systems and enabling efficient, reliable software delivery. This role combines hands-on engineering with strategic impact — you will design, build, and evolve the platform that underpins data pipelines and production services, ensuring scalability, security, and operational excellence across environments. Key Responsibilities Own and evolve CI/CD and automation platforms to support fast and reliable delivery of data-driven applications Design and manage cloud-native infrastructure supporting high-volume data ingestion, processing, and serving Build and maintain infrastructure as code to ensure consistency and scalability across environments Manage containerized environments and orchestration platforms to deliver resilient and scalable services Implement observability solutions (monitoring, logging, alerting) to ensure full system visibility and reliability Automate deployment processes, configuration management, and system recovery workflows Collaborate with engineering, data, and compliance teams to deliver secure and production-ready solutions Drive incident management practices and continuous improvement initiatives Contribute to platform strategy, tooling decisions, and mentoring within the team Requirements 5+ years of experience in DevOps, SRE, or infrastructure engineering roles Strong experience supporting production systems in distributed environments Hands-on experience with public cloud platforms (AWS or similar) Solid knowledge of containerization and orchestration technologies (Docker, Kubernetes) Experience with infrastructure as code tools (e.g., Terraform) Strong scripting/programming skills (Python, Bash, or similar) Experience building and maintaining CI/CD pipelines and automation tooling Knowledge of observability practices and tools Strong troubleshooting and incident response skills in Linux environments Excellent communication skills and ability to work cross-functionally Nice to Have Experience working with large-scale data platforms Exposure to regulated environments or compliance requirements Experience contributing to platform or engineering standards
Technology
Link Group
Senior Devops Engineer
Senior
Hybrid
Warsaw, Poland
28,000 - 38,000 PLN
🏢 Summary: Senior DevOps Engineer role focused on owning and evolving cloud-native infrastructure and CI/CD platforms supporting large-scale data processing systems. The position combines hands-on engineering with strategic platform development to ensure scalable, secure, and reliable production environments. You will design, automate, and maintain infrastructure and observability solutions across distributed systems. 🗂️ Requirements: 5+ years in DevOps, SRE, or infrastructure engineering, Experience supporting production systems in distributed environments, Hands-on experience with public cloud platforms (AWS or similar), Strong knowledge of Docker and Kubernetes, Experience with infrastructure as code tools (Terraform), Strong scripting/programming skills (Python or Bash), Experience building and maintaining CI/CD pipelines, Knowledge of observability, monitoring, and logging tools, Strong troubleshooting and incident response skills in Linux environments 📃 Skills: AWS, Docker, Kubernetes, Terraform, Python, Bash, Linux, CICD, Observability, Automation, Infrastructure, Cloud 🏢 Description: Senior DevOps Engineer We are looking for an experienced engineer to take ownership of our infrastructure and platform ecosystem, supporting large-scale data processing systems and enabling efficient, reliable software delivery. This role combines hands-on engineering with strategic impact — you will design, build, and evolve the platform that underpins data pipelines and production services, ensuring scalability, security, and operational excellence across environments. Key Responsibilities Own and evolve CI/CD and automation platforms to support fast and reliable delivery of data-driven applications Design and manage cloud-native infrastructure supporting high-volume data ingestion, processing, and serving Build and maintain infrastructure as code to ensure consistency and scalability across environments Manage containerized environments and orchestration platforms to deliver resilient and scalable services Implement observability solutions (monitoring, logging, alerting) to ensure full system visibility and reliability Automate deployment processes, configuration management, and system recovery workflows Collaborate with engineering, data, and compliance teams to deliver secure and production-ready solutions Drive incident management practices and continuous improvement initiatives Contribute to platform strategy, tooling decisions, and mentoring within the team Requirements 5+ years of experience in DevOps, SRE, or infrastructure engineering roles Strong experience supporting production systems in distributed environments Hands-on experience with public cloud platforms (AWS or similar) Solid knowledge of containerization and orchestration technologies (Docker, Kubernetes) Experience with infrastructure as code tools (e.g., Terraform) Strong scripting/programming skills (Python, Bash, or similar) Experience building and maintaining CI/CD pipelines and automation tooling Knowledge of observability practices and tools Strong troubleshooting and incident response skills in Linux environments Excellent communication skills and ability to work cross-functionally Nice to Have Experience working with large-scale data platforms Exposure to regulated environments or compliance requirements Experience contributing to platform or engineering standards
Technology
Fundacja Szkoła w Chmurze
DevOps Engineer
Mid
Remote
Warsaw, Poland
🏢 Summary: Offer for an experienced DevOps Engineer to design, automate and maintain large-scale hybrid cloud infrastructure using modern DevOps tools. The role focuses on CI/CD automation, Kubernetes orchestration, infrastructure as code and system monitoring. You will ensure high availability, performance and reliability in a complex project environment. 🗂️ Requirements: 3–5 years of DevOps experience, Strong knowledge of Kubernetes, Strong knowledge of Helm, Experience with GitLab and CI/CD pipelines, Strong Linux knowledge, Knowledge of networking, Experience with Prometheus and Grafana, Experience with ELK stack, Practical experience with Python, Knowledge of PostgreSQL, Experience with Terraform, Experience with AWS infrastructure 📃 Skills: Kubernetes, Helm, GitLab, CICD, Linux, Networking, Prometheus, Grafana, ELK, Python, Django, PostgreSQL, Terraform, AWS 🏢 Description: Do naszego zespołu poszukujemy doświadczonego DevOps Engineera, który dołączy do realizacji złożonego, wymagającego projektu o dużej skali. Jeśli lubisz pracę z nowoczesnym stackiem technologicznym, automatyzację procesów i rozwiązywanie realnych problemów infrastrukturalnych - zapraszamy! Zakres obowiązków: projektowanie, rozwój i utrzymanie infrastruktury w środowisku chmurowym (hybrydowym); automatyzacja procesów CI/CD oraz rozwój pipeline’ów; zarządzanie i orkiestracja środowisk kontenerowych (Kubernetes, Helm); monitorowanie i optymalizacja systemów (Prometheus, Grafana, ELK); współpraca z zespołami developerskimi przy wdrażaniu i utrzymaniu aplikacji; diagnozowanie problemów wydajnościowych i zapewnienie wysokiej dostępności systemów. Wymagania: 3–5 lat doświadczenia na stanowisku DevOps lub pokrewnym; bardzo dobra znajomość Kubernetes oraz Helm; doświadczenie z GitLab i GitLab CI/CD; solidna znajomość systemów Linux oraz zagadnień sieciowych; praktyczne doświadczenie z narzędziami monitoringu (Prometheus, Grafana) oraz ELK Stack; umiejętność pracy z Pythonem (mile widziane Django); znajomość baz danych PostgreSQL; doświadczenie z Terraform oraz infrastrukturą w AWS; umiejętność pracy w złożonym środowisku projektowym. Oferujemy: udział w ambitnym i technicznie wymagającym projekcie; dużą samodzielność i realny wpływ na architekturę systemu; bardzo elastyczny model współpracy - możliwa praca od 0,5 etatu do pełnego zaangażowania, w zależności od miesiąca i aktualnych potrzeb projektu; elastyczną formę współpracy; możliwość pracy z nowoczesnym stackiem technologicznym; przyjazną atmosferę i wsparcie zespołu. Proces rekrutacji: Przegląd zgłoszeń. 10-minutowy video screening na Google Meet z wybranymi kandydatami - zamiast screeningu telefonicznego wolimy tę formę, bo wierzymy, że wtedy możemy lepiej poczuć, czy do siebie pasujemy! Rozmowa z Liderką teamu oraz krótki test kompetencji logicznych.