June 16, 2026

Site Reliability Engineer

Senior • Remote

Lisbon, Portugal

Role Overview

We are looking for a skilled and proactive Observability Engineer to implement, automate, and support enterprise-grade observability and monitoring solutions across cloud and application platforms. The ideal candidate should have strong AWS infrastructure knowledge, hands-on automation skills, and experience building reliable monitoring and alerting ecosystems for modern distributed applications.

The role involves working closely with Platform Engineering, Data Engineering, and Application teams to develop observability solutions and bring operational visibility, reliability, incident detection, and platform performance.

Main Responsibilities

·        Design, implement, and maintain observability solutions for cloud-native and distributed systems.

·        Build monitoring, logging, alerting, and dashboarding solutions across infrastructure and applications.

·        Develop automation scripts and tooling using Python.

·        Implement and maintain Infrastructure as Code (IaC) using Terraform.

·        Build and support CI/CD pipelines using Jenkins and Git-based workflows.

·        Configure and optimize monitoring for AWS services, Kubernetes workloads, APIs, databases, and applications.

·        Create actionable alerts and operational dashboards to improve incident response and system reliability.

·        Work with engineering teams to onboard applications into observability platforms.

·        Support troubleshooting, root cause analysis, and performance optimization initiatives.

·        Ensure observability standards, governance, and best practices are followed across projects.

Key Requirements

·        Strong hands-on experience with Amazon Web Services (AWS).

·        Solid Python development/scripting experience.

·        Strong experience with Terraform.

·        Experience building and maintaining CI/CD pipelines using Jenkins.

·        Elasticsearch / ELK Stack experience and building queries.

·        Worked with Data Platforms monitoring is preferred.

·        Experience with Linux systems and shell scripting.

·        Understanding of monitoring, logging, and alerting concepts.

·        Experience working in Agile/DevOps environments.

Nice to Have Skills

Experience with any of the following is highly desirable:

·        Snowflake

·        Databricks

·        dbt

·        Matillion

·        Grafana

·        New Relic

·        Datadog

·        Prometheus

·        Elasticsearch / ELK Stack experience

NOTES: We are looking for an Engineer who loves to build. This is a highly technical role—90% of the job is hands-on coding in python and terraform.

Similar jobs you might like

Technology

Transition Technologies PSC

AWS DevOps Engineer

Mid

Remote

Lodz, Poland

🏢 Summary: The offer is for a DevOps Engineer responsible for designing, implementing, and maintaining scalable AWS-based cloud infrastructure within a Platform Engineering team. The role focuses on automation, CI/CD pipeline development, container orchestration, and infrastructure as code. It includes managing Kubernetes clusters, optimizing cloud environments, and ensuring reliable deployment and monitoring processes. 🗂️ Requirements: 3+ years experience with AWS services (EC2, S3, RDS, Lambda, VPC, IAM, CloudFormation), 2+ years experience with CI/CD tools (GitHub Actions, GitLab CI/CD, Jenkins), Experience with containerization and orchestration (Docker, Kubernetes, EKS, ECR, Helm), Experience with Infrastructure as Code tools (Terraform, CloudFormation), Experience with configuration management (Ansible), Scripting skills in Bash or Python, Experience managing cloud environments (dev, staging, production), Experience implementing deployment strategies (blue-green, canary, rollback), Experience with monitoring, alerting, and log aggregation 📃 Skills: AWS, EC2, S3, RDS, Lambda, VPC, IAM, CloudFormation, Terraform, Ansible, Docker, Kubernetes, EKS, ECR, Helm, GitHubActions, GitLabCI, Jenkins, Bash, Python 🏢 Description: We are looking for a DevOps Engineer to join our Platform Engineering team. In this role, you will be responsible for designing, implementing, and maintaining cloud infrastructure based on AWS. The ideal candidate has a strong background in automation, containerization, and CI/CD pipelines. Responsibilities: Design and implement scalable AWS infrastructure Manage development, staging, and production environments Monitor and optimize AWS costs Implement disaster recovery strategies Develop and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins Automate deployment processes Implement blue-green deployments Manage rollback strategies and canary deployments Manage Kubernetes clusters (EKS) Optimize Docker images and perform security scanning Develop and maintain Helm charts Automate infrastructure provisioning with Terraform or CloudFormation Manage configuration with Ansible Set up monitoring and alerting systems Aggregate and analyze logs Requirements: 3+ years AWS experience : EC2, S3, RDS, Lambda, VPC, IAM, CloudFormation 2+ years DevOps tools : GitHub Actions, GitLab CI/CD, Jenkins Containerization & orchestration : Docker, Kubernetes, ECR, Helm IaC & automation : Terraform, CloudFormation, Ansible, Bash, Python What can we offer: Flexible forms of employment and working hours (CoE or B2B) An interesting, challenging job in the dynamically developing Capital Group company; Work on innovative projects using modern technologies; Direct impact on shaping the image of the Capital Group’s companies on the market; Possibility to develop competences in a wide range; Attractive salary; Stability of employment and a friendly work atmosphere; Cool benefits, among others integration meetings, internal company competitions, fruit Tuesdays, sweet Thursdays and much more.

Technology

Link Group

Network Observability & Automation Engineer

Senior

Hybrid

Warsaw, Poland

37,000 - 47,000 PLN

🏢 Summary: Engineering role focused on building and scaling observability and automation solutions for a global, high-performance trading network. The position involves designing telemetry, monitoring, and alerting systems while driving automation to improve reliability and operational efficiency. You will troubleshoot complex network environments and develop infrastructure tooling using Python and Infrastructure-as-Code practices. 🗂️ Requirements: Hands-on experience with network observability and automation, Practical experience with Prometheus, Grafana, Telegraf, Experience with gNMI, YANG, CiscoTelemetry, Experience with Zabbix, LibreNMS, SNMP environments, Strong Python scripting skills, Experience with Ansible, Jinja, Git, Experience with Terraform and Infrastructure-as-Code, Familiarity with NetBox, NETCONF, RESTCONF, Experience with AWS cloud networking, Strong knowledge of BGP, OSPF, PIM, IGMP, Proven troubleshooting skills in large-scale network environments 📃 Skills: Python, Ansible, Jinja, Git, Terraform, NetBox, NETCONF, RESTCONF, Prometheus, Grafana, Telegraf, Zabbix, LibreNMS, SNMP, gNMI, YANG, CiscoTelemetry, AWS, BGP, OSPF, PIM, IGMP, Telemetry, Observability, Automation 🏢 Description: Join a team building and evolving observability and automation capabilities for a global, high-performance multi-vendor trading network. This role is ideal for an engineer passionate about scalable monitoring, telemetry, automation, and operational excellence in complex network environments. What you’ll do: Build and enhance network observability platforms and automation frameworks Design scalable telemetry, monitoring, and alerting solutions across a global infrastructure Drive automation initiatives to improve network reliability, visibility, and operational efficiency Develop tooling and infrastructure integrations using Python and Infrastructure-as-Code principles Collaborate with technology and investment teams to deliver resilient business-critical solutions Troubleshoot complex network and observability issues in large-scale environments What we’re looking for: Strong hands-on experience with network observability and automation Practical expertise with modern observability stack: Prometheus, Grafana, Telegraf, telemetry technologies (gNMI, YANG, Cisco Telemetry) Experience with monitoring platforms such as Zabbix, LibreNMS, and SNMP-based environments Strong scripting and automation skills using Python, Ansible, Jinja, and Git Familiarity with automation and infrastructure tooling including Terraform, NetBox, NETCONF, and RESTCONF Experience supporting cloud networking environments (AWS preferred) Good understanding of networking fundamentals and protocols such as BGP, OSPF, PIM, and IGMP Proven troubleshooting skills within complex, large-scale infrastructure environments Who you are: Passionate about observability, automation, and scalable infrastructure Proactive, self-driven, and eager to continuously improve systems and processes Comfortable operating in fast-paced, mission-critical environments Strong communicator and collaborative problem solver

Technology

DCV Technologies

Senior AWS DevOps Engineer

Senior

Remote

Warsaw, Poland

🏢 Summary: Senior AWS DevOps Engineer role focused on designing, automating, and maintaining scalable cloud infrastructure and CI/CD pipelines for distributed applications. The position involves managing Kubernetes environments, implementing Infrastructure as Code, and improving reliability, security, and observability across AWS platforms. You will collaborate with engineering teams to enhance deployment processes and ensure high availability of production systems. 🗂️ Requirements: 5+ years DevOps or Cloud Engineering experience, Strong hands-on experience with AWS services, Experience with Kubernetes and containerized environments, Strong experience with Terraform and Infrastructure as Code, Experience building and maintaining CI/CD pipelines, Strong Linux administration skills, Scripting skills in Bash or Python, Experience with monitoring and logging tools, Knowledge of networking, security, and IAM, Experience with Git-based workflows 📃 Skills: AWS, Kubernetes, EKS, Docker, Terraform, CloudFormation, CICD, Linux, Bash, Python, CloudWatch, Prometheus, Grafana, ELK, Datadog, Git, IAM, Jenkins, GitHub, GitLab, AzureDevOps, Helm, ArgoCD, Serverless, Networking, Security 🏢 Description: About the Role We are looking for a Senior AWS DevOps Engineer to join a fast-paced international engineering team. The ideal candidate will be responsible for designing, implementing, automating, and maintaining cloud infrastructure and CI/CD ecosystems for highly scalable distributed applications. You will work closely with development, security, and infrastructure teams to improve deployment processes, system reliability, scalability, and operational efficiency across cloud environments. Key Responsibilities Design, build, and maintain scalable AWS cloud infrastructure Implement and optimize CI/CD pipelines for automated deployments Manage Kubernetes clusters and containerized workloads Develop Infrastructure as Code (IaC) solutions using Terraform or CloudFormation Improve system observability, monitoring, logging, and alerting Automate operational and deployment processes Ensure high availability, performance, and security of production systems Collaborate with development teams to support cloud-native application delivery Participate in troubleshooting, incident response, and root cause analysis Support platform reliability, scalability, and disaster recovery initiatives Required Skills & Experience 5+ years of experience in DevOps / Cloud Engineering roles Strong hands-on experience with AWS services Experience with Kubernetes (EKS preferred) and Docker Strong experience with Terraform and Infrastructure as Code Experience building and maintaining CI/CD pipelines Strong Linux administration and scripting skills (Bash/Python) Experience with monitoring and logging tools such as CloudWatch, Prometheus, Grafana, ELK, Datadog, etc. Experience with Git-based workflows and automation tools Good understanding of networking, security, IAM, and cloud best practices Experience working in Agile/Scrum environments Strong troubleshooting and communication skills Nice to Have AWS Certifications Experience with GitHub Actions, GitLab CI, Jenkins, or Azure DevOps Experience with serverless architectures Exposure to multi-cloud environments (AWS/Azure) Knowledge of security and compliance frameworks Experience with Helm, ArgoCD, or service mesh technologies

Technology

Link Group

Platform Engineer

Mid

Remote

Warsaw, Poland

100 - 130 PLN

🏢 Summary: The offer is for a Platform Engineer to design, build, and support custom web applications and cloud infrastructure across AWS, Azure, CoreWeave, and on-prem environments. The role focuses on Infrastructure-as-Code, API integration, and ensuring system reliability and performance for technical users. It combines platform engineering, DevOps, and full-stack collaboration in enterprise environments. 🗂️ Requirements: 2+ years in platform engineering or DevOps-focused role, Experience with Linux administration, Proficiency in Python, Experience with Docker and Kubernetes, Experience with Terraform for Infrastructure-as-Code, Knowledge of modern JavaScript frameworks, Experience with AWS or Azure, Ability to design and maintain cloud infrastructure, Experience working with APIs and datasets 📃 Skills: Linux, Python, Docker, Kubernetes, Terraform, AWS, Azure, CoreWeave, JavaScript, React, Node.js, AWSCDK, APIs, DevOps 🏢 Description: THE WORK As a Platform Engineer, you will design, build, and support custom web applications running across enterprise environments such as AWS, Azure, CoreWeave, and on-prem infrastructure. You’ll support users in working with APIs and datasets, helping them configure, optimize, and operate these systems effectively. The role also includes designing and maintaining cloud infrastructure using Infrastructure-as-Code tools like Terraform and AWS CDK. You’ll collaborate with highly technical users, so clear communication and strong Linux skills are essential for troubleshooting, performance tuning, and ensuring system reliability. HERE’S WHAT YOU NEED Around 2+ years of experience in platform engineering or a similar development/DevOps-focused role Solid experience working with modern infrastructure and development stacks, including Linux, Python, Docker/Kubernetes, DevOps practices, and modern JavaScript frameworks (e.g. React or Node.js) Hands-on experience with Infrastructure-as-Code tools such as Terraform Strong problem-solving skills and ability to communicate complex technical topics clearly BONUS POINTS Experience with AI / LLMs or High-Performance Computing (HPC) Background in Data Engineering or Data Science Strong Linux expertise, including deeper OS-level understanding Experience with AWS services Building or supporting full-stack applications using cloud APIs

Technology

Link Group

DevOps Cloud engineer

Senior

Remote

Krakow, Poland

120 - 150 PLN

🏢 Summary: The role involves designing, implementing, and maintaining scalable cloud infrastructure and CI/CD pipelines, with a strong focus on Google Cloud Platform and Azure. The DevOps Engineer will automate deployments, manage infrastructure as code, and ensure reliability, security, and performance of cloud environments. The position requires close collaboration with development teams to improve deployment processes and platform stability. 🗂️ Requirements: Proven experience as DevOps Engineer or similar cloud role, Expert-level experience with Google Cloud Platform, Hands-on experience with Microsoft Azure, Experience with CI/CD tools (Azure DevOps, GitHub, GitHub Actions), Experience with Infrastructure as Code using Terraform, Scripting and automation skills using Python, Knowledge of cloud architecture and deployment best practices, Strong troubleshooting skills, English proficiency 📃 Skills: GCP, Azure, Terraform, Python, AzureDevOps, GitHub, GitHubActions, CI/CD, IaC 🏢 Description: About the Role We are looking for a skilled DevOps Engineer to join our technology team and help design, implement, and maintain scalable cloud infrastructure and modern CI/CD pipelines. In this role, you will work closely with development and platform teams to automate processes, improve deployment efficiency, and ensure the reliability and security of cloud-based systems. The ideal candidate has strong experience with cloud platforms, infrastructure as code, and modern DevOps practices, with particular expertise in Google Cloud Platform and CI/CD tooling. Key Responsibilities Design, implement, and maintain CI/CD pipelines using Azure DevOps , GitHub , and GitHub Actions . Build and manage scalable cloud infrastructure on Google Cloud Platform and Microsoft Azure . Develop and maintain infrastructure using Infrastructure as Code practices with Terraform . Automate operational and deployment processes using Python . Monitor and optimize cloud environments for performance, reliability, and cost efficiency. Collaborate with development teams to improve application deployment processes and platform reliability. Implement best practices in security, access control, and cloud governance. Troubleshoot infrastructure and deployment issues across environments. Required Skills & Experience Proven experience as a DevOps Engineer or in a similar cloud/infrastructure role. Strong hands-on experience with Google Cloud Platform at a principal or expert level . Solid experience with Microsoft Azure . Hands-on experience with CI/CD tools such as Azure DevOps , GitHub , and GitHub Actions . Strong knowledge of Infrastructure as Code tools, particularly Terraform . Good scripting and automation skills using Python . Experience with cloud architecture, automation, and deployment best practices. Strong troubleshooting and problem-solving skills. Ability to work collaboratively in cross-functional teams. Good communication skills and proficiency in English.

Technology

Link Group

DevOps Engineer with Data background

Mid

Remote

Warsaw, Poland

130 - 150 PLN

🏢 Summary: DevOps Engineer role focused on supporting and scaling modern data platforms by combining DevOps practices with data engineering expertise. The position involves building and maintaining CI/CD pipelines, managing cloud-based data infrastructure, and ensuring reliable, automated data workflows. You will collaborate with data teams to optimize performance, scalability, and security of data solutions. 🗂️ Requirements: 4–5 years experience in DevOps or similar role, Hands-on experience with Databricks, Strong experience with AWS, Experience building and maintaining CI/CD pipelines with Jenkins, Background in data engineering or data platforms, Basic to intermediate Python knowledge, Understanding of ETL/ELT processes, Knowledge of distributed data systems 📃 Skills: Jenkins, Databricks, AWS, Python, CICD, ETL, ELT, DevOps, DataEngineering, Cloud, DistributedSystems 🏢 Description: About the Role We are looking for a DevOps Engineer with a strong data background to support and scale modern data platforms. In this role, you will work at the intersection of DevOps and Data Engineering , ensuring reliable, automated, and efficient data pipelines and environments. You will be responsible for maintaining and improving CI/CD processes, supporting cloud-based data platforms, and collaborating with data engineers and analytics teams to enable smooth data operations. Key Responsibilities Design, implement, and maintain CI/CD pipelines using Jenkins . Manage and optimize data platforms based on Databricks . Build and maintain cloud infrastructure on Amazon Web Services . Support deployment, monitoring, and troubleshooting of data pipelines and workflows. Collaborate closely with data engineers to improve reliability and performance of data solutions. Automate infrastructure and operational processes. Ensure best practices in security, scalability, and performance across the platform. Required Skills & Experience 4–5 years of experience in a DevOps Engineer or similar role. Hands-on experience with Databricks . Strong experience with Amazon Web Services . Experience building and maintaining CI/CD pipelines using Jenkins . Background in data engineering or experience working with data platforms. Basic to intermediate knowledge of Python . Understanding of data pipelines, ETL/ELT processes, and distributed data systems. Strong problem-solving skills and ability to troubleshoot complex issues. Good communication skills and ability to work in cross-functional teams.

Technology

SoftBlue

Data Engineer (Python & AWS)

Senior

Remote

Bydgoszcz, Poland

150 - 180 PLN

🏢 Summary: Senior Data Engineer role focused on designing and building scalable serverless data ingestion pipelines in AWS within the healthcare domain. The position emphasizes strong Python engineering, cloud architecture leadership, and implementation of modern data platforms and DevOps practices. The role involves driving technical excellence and delivering reliable, high-impact data solutions in an international environment. 🗂️ Requirements: 10+ years of experience in Python programming, Strong software engineering skills in data processing, Extensive experience with AWS Cloud and Serverless Architecture, Hands-on experience with AWS Lambda, S3, and Cognito, Experience building E2E automated tests for data pipelines, Practical knowledge of Data Mesh and Medallion Architecture, Experience with Infrastructure as Code using AWS CDK or Terraform, Experience with CI/CD pipelines using GitLab or GitHub Actions, Experience with ETL/ELT processes and dbt, Experience working with GraphQL, Minimum B2 level English proficiency 📃 Skills: Python, AWS, Lambda, S3, Cognito, Boto3, DataMesh, Medallion, CDK, Terraform, GitLab, GitHubActions, ETL, ELT, dbt, GraphQL, CI/CD 🏢 Description: We are looking for a highly skilled Data Engineer to join our client in the healthcare sector. Our requirements: Technical Expertise: Python Programming: 10+ years of experience with strong Software Engineering skills focused on data processing. AWS & Serverless: Extensive experience with AWS Cloud, specifically focusing on Serverless Architecture and services (including AWS Lambda , AWS S3 Tables , and AWS Cognito ). Automated Testing: Proven experience in developing End-to-End (E2E) automated tests to ensure pipeline reliability, utilizing tools such as Boto3 for AWS resource validation. Data Concepts: Practical knowledge of Data Mesh and Medallion Architecture , along with general data processing and analysis. DevOps & IaC: Hands-on experience with Infrastructure as Code ( AWS CDK or Terraform ) and CI/CD pipelines ( GitLab pipelines or GitHub Actions ). Modern Tooling: Experience with ETL/ELT solutions, dbt , and GraphQL . Communication & Soft Skills: English Language: Minimum B2 level , enabling smooth daily technical and business communication in a global environment. Collaboration: Excellent communication skills and the ability to thrive in a collaborative, international team. Standards: A strong commitment to high standards of ethics, quality (Clean Code), and reliable delivery. Nice to have: Experience with Snowflake and SQL . Knowledge of Data Vault 2.0 modeling. Experience with Databricks . Familiarity with the Microsoft ecosystem: C# / .Net, T-SQL, SQL Server , and Azure DevOps . Experience with Star Schema database modeling. Knowledge of Descriptive Statistics. Your responsibilites: Design and build scalable Data Ingestion pipelines within the AWS cloud ecosystem. Lead technical delivery and implementation of core platform components, ensuring architectural integrity across the entire data lifecycle. Collaborate with Engineering Managers and cross-functional teams across the globe and Poland. Drive technical excellence by improving team processes, architecture standards, and engineering best practices. Support and consult with stakeholders to ensure successful delivery of high-impact, data-driven solutions. Contribute to the growth and maturity of the team’s cloud and data engineering capabilities. We offer: Challenging role within the company that creates innovative solutions. Work in international environment on demanding projects. Remote work model. Subsidized private medical care, life insurance, multisport card. Integration meetings. Employee referral program. If you have a deep expertise in Python and AWS , and building scalable Serverless data architectures is where you truly excel, this is the perfect role for you!

Technology

Link Group

Senior Site Reliability Engineer

Senior

Hybrid

Warsaw, Poland

170 - 230 PLN

🏢 Summary: The role focuses on ensuring reliability, scalability, and performance of large-scale cloud-based applications by building and maintaining resilient infrastructure. You will manage AWS cloud environments, Kubernetes clusters, and CI/CD pipelines while implementing monitoring, automation, and incident response processes. The position emphasizes Infrastructure-as-Code, observability, and continuous reliability improvements. 🗂️ Requirements: 5+ years experience in SRE, DevOps or similar role, Strong experience with AWS cloud services, Experience with Infrastructure-as-Code tools, Hands-on experience with Kubernetes, Proficiency with Docker, Experience with CI/CD pipelines, Solid knowledge of PostgreSQL or Amazon RDS, Strong SQL knowledge, Knowledge of networking concepts (VPC, DNS, troubleshooting), Strong Linux/Unix administration skills, Experience with observability tools, Experience with automation in infrastructure, Experience with incident management 📃 Skills: AWS, Terraform, Pulumi, Kubernetes, EKS, Docker, GitHub, PostgreSQL, RDS, SQL, VPC, DNS, Linux, Unix, Prometheus, Grafana, Datadog, Dynatrace, CI/CD 🏢 Description: We are looking for an experienced Site Reliability Engineer to ensure the reliability, scalability, and performance of large-scale cloud-based web applications. You will work closely with software development, cloud operations, and platform teams to build and maintain resilient infrastructure and improve system stability. Key Responsibilities: Design and maintain monitoring, alerting, and incident response systems to ensure high availability Collaborate closely with engineering, product, and architecture teams Build and manage cloud infrastructure using Infrastructure-as-Code (e.g., Terraform, Pulumi) on AWS Operate and optimize Kubernetes environments (e.g., EKS) Develop and maintain containerized applications using Docker Improve CI/CD pipelines and drive automation across deployment processes Implement and manage observability tools (logging, metrics, tracing) Participate in incident management, postmortems, and reliability improvements Support capacity planning, disaster recovery, and system scaling Contribute to security, compliance, and operational best practices Develop automation and AI-driven solutions for monitoring and incident prevention Requirements: 5+ years of experience in SRE, DevOps, or similar roles Strong experience with AWS cloud services and Infrastructure-as-Code tools Hands-on experience with Kubernetes and containerized environments Proficiency in Docker and CI/CD pipelines (e.g., GitHub Actions) Solid understanding of databases (e.g., PostgreSQL, Amazon RDS) and SQL Knowledge of networking concepts (VPC, DNS, troubleshooting tools like dig/traceroute) Strong Linux/Unix administration skills Experience with observability tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) Familiarity with automation and AI-based solutions in infrastructure Strong problem-solving and incident management skills

Technology

Toro Performance Sp. z o.o.

Senior Cloud Engineer

Senior

Remote

🏢 Summary: Senior Cloud Engineer role focused on designing and implementing scalable, secure AWS infrastructure for an IoT and machine learning project processing telemetry data. The position involves infrastructure as code, container orchestration, CI/CD automation, and cloud security management. The engineer will manage EKS/ECS environments and ensure high availability, scalability, and compliance with security best practices. 🗂️ Requirements: Minimum 5 years experience as Cloud Engineer or DevOps, Higher education in Computer Science or related field or equivalent experience, English level B2 or higher, Strong knowledge of AWS services and architecture, Proficiency in Terraform and AWS CloudFormation or AWS CDK, Experience with Kubernetes and Amazon ECS, Experience building CI/CD pipelines with Jenkins and GitHub Actions, Scripting skills in Python or Bash or Nodejs, Knowledge of infrastructure security best practices, Experience managing and troubleshooting EKS and ECS environments 📃 Skills: AWS, Terraform, CloudFormation, CDK, Kubernetes, ECS, EKS, Jenkins, GitHub, Python, Bash, Nodejs, Kinesis, Glue, EMR, Redshift, Athena, Sagemaker, EventBridge, RDS, Aurora, PostgreSQL, S3, Iceberg, Parquet, Flink, Datadog, Domo, OvalEdge, OpsGenie, Java, Scala, Kotlin, Docker, Argo 🏢 Description: Poszukujemy Senior Cloud Engineera, który wesprze projekt łączący technologię IoT, uczenie maszynowe i dane telemetryczne dla rozwoju przyszłościowych, inteligentnych rozwiązań dla całego świata. Obowiązki: Zaprojektowanie i wdrożenie infrastruktury AWS oraz środowisk kontenerowych prowadzące do zwiększenia dostępności i skalowalności. Rozwój i utrzymanie infrastruktury przy użyciu Terraform, CloudFormation i AWS CDK. Monitoring infrastruktury. Tworzenie i utrzymanie CI/CD korzystając zarówno z Jenkins, jak i GitHub Actions. Utrzymanie bezpiecznych środowisk kontenerowych. Skuteczne wdrażanie najlepszych praktyk w zakresie bezpieczeństwa, prowadzące do bezpiecznego środowiska infrastrukturalnego. Zarządzanie środowiskami Amazon EKS i Amazon ECS, w tym wdrażanie, skalowanie i rozwiązywanie problemów. Wdrażanie zasad bezpieczeństwa, aby zapewnić zgodność z najlepszymi praktykami. Przeprowadzanie regularnych audytów bezpieczeństwa i oceny podatności. Reagowanie na incydenty związane z bezpieczeństwem i naprawianie ich. Współpraca z zespołami i prowadzenie dokumentacji. Wymagania: Minimum 5 lat doświadczenia na stanowisku Cloud Engineera, DevOps lub innym podobnym stanowisku. Wykształcenie wyższe (min. licencjat z informatyki (lub inny kierunek pokrewny) lub równoważne doświadczenie zawodowe). Język angielski na poziomie min. B2 (zespół międzynarodowy) Znajomość usług i rozwiązań AWS. Biegła znajomość Terraform, AWS CloudFormation i/lub AWS CDK. Doświadczenie w Kubernetes i Amazon ECS. Praktyczne doświadczenie w automatyzacji CI/CD przy użyciu technologii takich jak Jenkins i GitHub Actions. Umiejętność pisania skryptów (np. Python, Bash, Nodejs) i znajomość narzędzi do automatyzacji. Znajomość najlepszych praktyk w zakresie bezpieczeństwa infrastruktury. Doskonałe umiejętności rozwiązywania problemów i współpracy. Certyfikaty AWS i Kubernetes będą dodatkowym atutem. Mile widziana znajomość: systemów kontroli wersji (BitBucket, GitHub, GitLab). Mile widziane doświadczenie z Docker, Kubernetes i Argo Tech Stack: AWS: Kinesis, Glue ETL, EMR, Redshift, Athena, Sagemaker, EventBridge, EKS, RDS Aurora, PostgreSQL AWS S3 + Iceberg + Paruqet Flink Datadog Domo BI OvalEdge OpsGenie GitHub & GitHub Actions Python, Java, Scala, Kotlin

Technology

emagine Polska

Observability Specialist

Senior

Hybrid

Warsaw, Poland

🏢 Summary: The offer is for an Observability Specialist responsible for designing, implementing, and maintaining a scalable telemetry and monitoring infrastructure in cloud-native environments. The role focuses on Kubernetes observability, Elastic Stack management, and performance optimization using modern telemetry standards. It involves driving SRE practices and ensuring high system reliability through advanced monitoring and AIOps solutions. 🗂️ Requirements: Experience monitoring Kubernetes (OpenShift) environments, Hands-on implementation of OpenTelemetry for logs, traces, and metrics, Strong expertise in ELK stack deployment and maintenance, Proficiency in automating Elastic environments using Ansible, Experience with Application Performance Monitoring for code-level analysis, Knowledge of shard optimization, mapping, and Index Lifecycle Management, Experience defining and monitoring SLOs and managing Error Budgets, Integration of observability solutions with major cloud providers 📃 Skills: Kubernetes, OpenShift, OpenTelemetry, Elasticsearch, Logstash, Kibana, Ansible, ElasticAPM, AIOps, SRE, ILM, Sharding, Mapping, Cloud 🏢 Description: Introduction & Summary We are seeking an experienced Observability Specialist dedicated to ensuring the reliability and performance of our systems. This role involves collaborating with enterprise architects and IT professionals to design, implement, and oversee a scalable telemetry infrastructure. The ideal candidate will possess deep expertise in ELK or similiar technologies and modern telemetry standards. Main Responsibilities As our Observability Engineer, your core duties will include: Architectural Collaboration: Partner with system architects and local engineering teams in Denmark to design resilient monitoring solutions. Monitor Kubernetes environments with OpenTelemetry (OTel) standards for logs, traces, and metrics. Manage centralized data collection and automate Elastic deployments using Ansible. Utilize Elastic APM for identifying code-level bottlenecks and resolving latency issues. Implement AIOps configurations for proactive anomaly detection and automated root-cause analysis. Drive Site Reliability Engineering (SRE) methodologies across teams. Elastic Stack Management: Deploy, scale, and maintain Elasticsearch, Logstash, and Kibana (ELK) environments. Key Requirements Cloud-Native Observability: Strong skills in monitoring Kubernetes (Openshift) environments and integrating with major cloud providers. APM & Distributed Tracing: Expertise in Application Performance Monitoring (APM) to identify code-level bottlenecks and latency issues. OpenTelemetry (OTel): Hands-on experience implementing OpenTelemetry (or similiar) standards for logs, traces, and metrics to ensure vendor-neutral telemetry. Infrastructure as Code (IaC): Proficiency in automating Elastic environments with Ansible. Performance Engineering: Expert-level knowledge of shard optimization, mapping, and Index Lifecycle Management (ILM) to balance high performance with cost control. SRE Methodology: Experience defining and monitoring Service Level Objectives (SLOs) and managing Error Budgets. Strong communication skills for collaboration with IT teams. NIce to Have: Elastic Stack Mastery: Deep expertise in architecting and managing Elasticsearch, Logstash, and Kibana (ELK) at scale. Data Ingestion & Fleet: Proven experience deploying Elastic Agent and Fleet for centralized agent management and data collection. AIOps & Machine Learning: Ability to configure Elastic ML models for proactive anomaly detection and automated root cause analysis. Other Details This is position based in Warsaw, flexible Hybrid model, focused on leading-edge observability solutions in a dynamic and collaborative environment.