June 30, 2026

Sr. Software Engineer (Data Center Automation)

Senior • On-site

Memphis, TN

About the Role

We are seeking a highly skilled Sr. Software Engineer to manage and enhance reliability across a multi-data center environment. This role focuses on automating processes, building robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The position bridges software engineering principles with physical data center realities to deliver resilient, scalable systems with near-zero downtime.

The primary objective is to mitigate downtime and minimize end-user impact from scheduled and unscheduled maintenance through proactive automation, observability, and integrated software-physical reliability strategies.

Responsibilities

  • Design, develop, and deploy scalable services (primarily in Python and Rust) to automate monitoring, alerting, incident response, and infrastructure provisioning.
  • Implement and maintain observability solutions including metrics, logging, tracing, dashboards, and alerting systems.
  • Collaborate with software, network, site, and facility operations teams to automate fault tolerance, disaster recovery, capacity planning, and environmental risk mitigation.
  • Troubleshoot complex data center issues including hardware failures, software bugs, and network problems while adhering to SLAs and error budgets.
  • Optimize Linux-based systems through kernel tuning, container orchestration (e.g., Kubernetes), and automation scripting.
  • Analyze and troubleshoot large-scale network topologies across multi-data center environments.
  • Participate in on-call rotations, incident response, and blameless postmortems.
  • Mentor junior engineers and promote documentation and automation best practices.

Basic Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience).
  • 3+ years of experience in SRE, Infrastructure, DevOps, or Systems Engineering in distributed production environments.
  • Strong production programming experience in Python; familiarity with Rust or other systems languages (Go, C++).
  • Experience with Linux systems administration, performance tuning, and automation.
  • Knowledge of containerization and orchestration tools such as Docker and Kubernetes.
  • Experience implementing observability tools (e.g., Prometheus, Grafana) including metrics, logging, tracing, and alerting.
  • Understanding of networking fundamentals including TCP/IP, routing, redundancy, and DNS.
  • Experience with incident response, on-call rotations, and reliability best practices (SLAs, error budgets).

Preferred Skills and Experience

  • 5+ years of SRE or infrastructure experience in hyperscale or AI/ML environments.
  • Large-scale Kubernetes operations and automation experience.
  • Proficiency in Rust for systems programming.
  • Experience integrating software reliability with physical data center infrastructure.
  • Experience building automated remediation and disaster recovery systems.
  • Background optimizing Linux systems for AI workloads or GPU clusters.
  • Experience with bare-metal provisioning and multi-site failover mechanisms.
  • Mentoring and strong documentation skills.

Similar jobs you might like

Technology

xAI

Sr. Software Engineer (Data Center Automation)

Senior

On-site

Palo Alto, CA

🏢 Summary: Senior Software Engineer role focused on building automation and observability solutions to enhance reliability across multi-data center AI infrastructure. The position combines strong programming skills with hands-on data center and Linux systems expertise to minimize downtime and optimize performance. It involves developing scalable services, improving monitoring and incident response, and collaborating across infrastructure and facility teams. 🗂️ Requirements: Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering or related field (or equivalent experience), 3+ years experience in SRE, Infrastructure, DevOps, or Systems Engineering in large-scale production environments, Strong production experience in Python, Solid Linux systems administration and kernel-level knowledge, Experience with containerization and orchestration (Docker, Kubernetes or similar), Experience implementing observability solutions (metrics, logging, tracing, monitoring, alerting), Understanding of networking fundamentals (TCP/IP, routing, DNS, redundancy), Experience troubleshooting distributed systems, hardware and network issues, Experience with on-call rotations and incident response practices (SLAs, error budgets), Ability to collaborate with cross-functional technical teams 📃 Skills: Python, Rust, Linux, Kubernetes, Docker, Prometheus, Grafana, TCP/IP, DNS, Scripting, Automation, Observability, Monitoring, Tracing, Networking 🏢 Description: ABOUT xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are seeking a highly skilled Sr. Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment. This role focuses on automating processes, building and implementing robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The ideal candidate will combine strong coding abilities with hands-on data center experience to build scalable reliability services, optimize system performance, and minimize downtime—including close partnership with facility operations to address physical infrastructure impacts. In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities. By prioritizing automation and observability, team members in this role can reduce mean time to recovery (MTTR) by up to 50% through proactive monitoring and automated remediation. The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers. This is achieved through proactive automation, robust observability, and integrated software-physical reliability strategies. RESPONSIBILITIES: - Design, develop, and deploy scalable code and services (primarily in Python and Rust) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning. - Implement and maintain observability tools and practices, such as metrics collection, logging, tracing, and dashboards, to provide real-time insights into system health across multiple data centers. - Collaborate with cross-functional teams to identify reliability bottlenecks and automate solutions for fault tolerance, disaster recovery, capacity planning, and physical/environmental risk mitigation. - Troubleshoot and resolve complex issues in data center environments, including hardware failures, environmental anomalies, software bugs, and network-related problems, while adhering to reliability principles like error budgets and SLAs. - Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration, and scripting for automation. - Understand network topologies and concepts in large-scale, multi-data center environments to troubleshoot connectivity, routing, redundancy, and performance issues. - Participate in on-call rotations, post-incident reviews (blameless postmortems), and continuous improvement initiatives. - Mentor junior team members and document processes to foster a culture of automation and knowledge sharing. BASIC QUALIFICATIONS: - Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience). - 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering in large-scale, distributed, or production environments. - Strong programming skills with proven production experience in Python; experience with Rust or another systems-level language (e.g., Go, C++) is essential. - Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments. - Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes. - Experience implementing observability solutions, including metrics, logging, tracing, monitoring tools, alerting, and dashboards. - Familiarity with troubleshooting complex issues in distributed systems, including software bugs, hardware failures, network problems, and environmental factors. - Understanding of networking fundamentals (TCP/IP, routing, redundancy, DNS) in large-scale or multi-site environments. - Experience participating in on-call rotations, incident response, post-incident reviews, and reliability practices such as error budgets or SLAs. - Ability to collaborate effectively with cross-functional teams. PREFERRED SKILLS AND EXPERIENCE: - 5+ years of experience in SRE or infrastructure roles in hyperscale, cloud, or AI/ML training environments with multi-data center setups. - Hands-on experience operating or scaling Kubernetes clusters at large scale, including automation for provisioning and high availability. - Proficiency in Rust for systems programming and performance-critical components. - Experience integrating software reliability tools with physical data center infrastructure (power, cooling, environmental monitoring). - Experience building automated remediation, fault tolerance, disaster recovery, capacity planning, or predictive failure detection systems. - Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments. - Experience with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms. - Mentoring experience and strong documentation skills. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

Relativity

Senior Engineer - Site Reliability Engineering

Senior

Remote

Krakow, Poland

208,000 - 312,000 PLN/yr

🏢 Summary: Senior Software Engineer – SRE role focused on building and maintaining highly available, observable, and resilient cloud-native systems. The position emphasizes automation, CI/CD improvements, incident management, and implementation of reliability best practices across a SaaS platform. You will collaborate cross-functionally to enhance scalability, performance, and operational excellence. 🗂️ Requirements: 5+ years in Software Engineering, SRE, or Cloud Infrastructure, Experience with DevOps tools and practices, Proficiency in Python, Go, Java, or C#/.NET, Experience with at least two: GitHub, Azure DevOps, GitLab, Jenkins, Hands-on experience with observability tools, Strong experience with CI/CD pipelines and automation, Experience with cloud-native distributed systems, Experience in high-availability SaaS environments, Knowledge of SLOs, SLIs, and error budgets, Experience with incident management and root cause analysis, Experience implementing redundancy and disaster recovery, Experience with Agile methodologies 📃 Skills: Python, Go, Java, C#, DotNet, GitHub, Azure, GitLab, Jenkins, Prometheus, Grafana, OpenTelemetry, CI/CD, DevOps, SLO, SLI, SaaS, Automation, Cloud, DistributedSystems 🏢 Description: Job Overview As the Senior Software Engineer – SRE you will focus on implementing and maintaining reliability solutions across the platform. This role emphasizes hands-on engineering work, automation, and operational excellence. The Senior Software Engineer will work closely with other engineers to ensure systems are highly available, observable, and resilient. As a member of the engineering team, the Senior Software Engineer will work closely with Infrastructure, Engineering, and Product teams to develop highly resilient, observable, and automated solutions that enhance system availability and efficiency. The ideal candidate will bring deep technical expertise, strong problem-solving skills, and a passion for reliability engineering. Job Description and Requirements Job Responsibilities Implement, and advocate for best-in-class reliability, observability, and scalability practices across the platform. Develop automated solutions for system reliability, capacity planning, and incident response to minimize manual intervention. Participate in improving Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to enhance system reliability. Contribute to CI/CD pipeline improvements and DevOps practices. Support root cause analysis (RCA) investigations, drive corrective actions, and advocate for a blameless postmortem culture. Participate in on-call rotations to ensure 24/7 availability of critical systems. Influence and mentor engineering teams on SRE principles, DevOps culture, and best practices. Stay ahead of industry trends, adopting new tools, frameworks, and methodologies to continually improve system reliability. Preferred Qualifications 5+ years of experience in software engineering, site reliability engineering, or cloud infrastructure roles. Experience with DevOps tooling and practices. Proficient in building service-oriented architectures and cloud-native distributed systems. Proficiency in programming languages such as Python, Go, Java, or C# or .Net. In-depth technical understanding and experience with at least two of the following DevOps platforms: GitHub, Azure DevOps, GitLab, or Jenkins. Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry or others). Strong background in CI/CD pipelines, automation, and DevOps practices. Experience working in global, high-availability SaaS environments. Experience implementing redundancy and disaster recovery scenarios. Excellent teamwork and cross-group collaboration skills. Ability to collaborate with both technical and business professionals. Hands-on experience with Agile Project Development Methodologies. Experience delivering complex technical solutions. Excellent problem-solving, analytical, and communication skills. Nice to have: Experience with Chaos Engineering and/or AI Ops . Competencies and Skills Automation-First Mindset – Commitment to reducing toil through scripting and automation. Reliability Engineering – Expertise in SLOs, SLIs, error budgets, and high-availability architectures. Incident Management & Postmortems – Experience in handling production incidents and driving continuous improvement. Observability & Monitoring – Deep understanding of logging, monitoring, and alerting best practices. Practical knowledge of data structures and modern data engines. Collaboration & Communication – Ability to work across teams, influence stakeholders, and advocate for reliability improvements. Mentorship & Coaching – Passion for mentoring engineers and building an SRE culture within the organization. Additional Information This role offers a unique opportunity to shape the future of SRE in a cutting-edge SaaS company, ensuring the reliability and scalability of mission-critical applications for customers worldwide. If you are passionate about solving complex reliability challenges and driving technical excellence, we’d love to hear from you! Relativity is a diverse workplace with different skills and life experiences—and we love and celebrate those differences. We believe that employees are happiest when they're empowered to be their full, authentic selves, regardless how you identify. Benefit Highlights: Comprehensive health, dental, and vision plans Parental leave for primary and secondary caregivers Flexible work arrangements Two, week-long company breaks per year Additional time off Long-term incentive program Training investment program All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law. Relativity is committed to competitive, fair, and equitable compensation practices. This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives. The expected salary range for this role is between following values: 208 000 and 312 000PLN The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position. Required Skills: Automation, Data Analysis, Database Management, Network Architecture, Performance Optimizations, Problem Solving, Project Management, Software Development, System Designs, Technical Leadership

Technology

Relativity

Senior Engineer - Site Reliability Engineering

Senior

Remote

Krakow, Poland

208,000 - 312,000 PLN/yr

🏢 Summary: Remote Senior Software Engineer – SRE role focused on building and maintaining highly available, scalable, and observable cloud-native systems. The position emphasizes automation, CI/CD improvements, incident management, and implementation of reliability best practices across SaaS platforms. The engineer collaborates cross-functionally to enhance system resilience, performance, and operational excellence. 🗂️ Requirements: 5+ years in Software Engineering, SRE, or Cloud Infrastructure roles, Experience with DevOps tools and practices, Proficiency in Python, Go, Java, C#, or .Net, Experience with at least two: GitHub, Azure DevOps, GitLab, Jenkins, Hands-on experience with observability tools, Strong experience with CI/CD pipelines and automation, Experience with cloud-native distributed systems, Experience in high-availability SaaS environments, Knowledge of SLOs, SLIs, and error budgets, Experience with redundancy and disaster recovery, Participation in on-call rotations 📃 Skills: Python, Go, Java, C#, .Net, GitHub, Azure, GitLab, Jenkins, Prometheus, Grafana, OpenTelemetry, CI/CD, DevOps, SLO, SLI, SaaS, Automation, Cloud, Agile 🏢 Description: Posting Type Remote Job Overview As the Senior Software Engineer – SRE you will focus on implementing and maintaining reliability solutions across the platform. This role emphasizes hands-on engineering work, automation, and operational excellence. The Senior Software Engineer will work closely with other engineers to ensure systems are highly available, observable, and resilient. As a member of the engineering team, the Senior Software Engineer will work closely with Infrastructure, Engineering, and Product teams to develop highly resilient, observable, and automated solutions that enhance system availability and efficiency. The ideal candidate will bring deep technical expertise, strong problem-solving skills, and a passion for reliability engineering. Job Description and Requirements Job Responsibilities Implement, and advocate for best-in-class reliability, observability, and scalability practices across the platform. Develop automated solutions for system reliability, capacity planning, and incident response to minimize manual intervention. Participate in improving Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to enhance system reliability. Contribute to CI/CD pipeline improvements and DevOps practices. Support root cause analysis (RCA) investigations, drive corrective actions, and advocate for a blameless postmortem culture. Participate in on-call rotations to ensure 24/7 availability of critical systems. Influence and mentor engineering teams on SRE principles, DevOps culture, and best practices. Stay ahead of industry trends, adopting new tools, frameworks, and methodologies to continually improve system reliability. Preferred Qualifications 5+ years of experience in software engineering, site reliability engineering, or cloud infrastructure roles. Experience with DevOps tooling and practices. Proficient in building service-oriented architectures and cloud-native distributed systems. Proficiency in programming languages such as Python, Go, Java, or C# or .Net. In-depth technical understanding and experience with at least two of the following DevOps platforms: GitHub, Azure DevOps, GitLab, or Jenkins. Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry or others). Strong background in CI/CD pipelines, automation, and DevOps practices. Experience working in global, high-availability SaaS environments. Experience implementing redundancy and disaster recovery scenarios. Excellent teamwork and cross-group collaboration skills. Ability to collaborate with both technical and business professionals. Hands-on experience with Agile Project Development Methodologies. Experience delivering complex technical solutions. Excellent problem-solving, analytical, and communication skills. Nice to have: Experience with Chaos Engineering and/or AI Ops . Competencies and Skills Automation-First Mindset – Commitment to reducing toil through scripting and automation. Reliability Engineering – Expertise in SLOs, SLIs, error budgets, and high-availability architectures. Incident Management & Postmortems – Experience in handling production incidents and driving continuous improvement. Observability & Monitoring – Deep understanding of logging, monitoring, and alerting best practices. Practical knowledge of data structures and modern data engines. Collaboration & Communication – Ability to work across teams, influence stakeholders, and advocate for reliability improvements. Mentorship & Coaching – Passion for mentoring engineers and building an SRE culture within the organization. Additional Information This role offers a unique opportunity to shape the future of SRE in a cutting-edge SaaS company, ensuring the reliability and scalability of mission-critical applications for customers worldwide. If you are passionate about solving complex reliability challenges and driving technical excellence, we’d love to hear from you! Relativity is a diverse workplace with different skills and life experiences—and we love and celebrate those differences. We believe that employees are happiest when they're empowered to be their full, authentic selves, regardless how you identify. Benefit Highlights: Comprehensive health, dental, and vision plans Parental leave for primary and secondary caregivers Flexible work arrangements Two, week-long company breaks per year Additional time off Long-term incentive program Training investment program All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law. Relativity is committed to competitive, fair, and equitable compensation practices. This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives. The expected salary range for this role is between following values: 208 000 and 312 000PLN The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position. Required Skills: Automation, Data Analysis, Database Management, Network Architecture, Performance Optimizations, Problem Solving, Project Management, Software Development, System Designs, Technical Leadership

Technology

emagine Polska

Senior Software Engineer (Java // Python)

Senior

Hybrid

Lisbon, Portugal

🏢 Summary: Senior Software Engineer responsible for defining technical standards and architecture, leading backend and frontend development, and ensuring secure, scalable, and observable solutions. The role drives CI/CD, enterprise integrations, and production support while aligning engineering practices with product and operations goals. 🗂️ Requirements: Proven experience defining technical standards and architecture, Experience with CI/CD orchestration, Experience with observability implementation, Experience implementing on-premises, hybrid, and cloud solutions, Experience leading backend and frontend architectural decisions, Experience maintaining CI/CD pipelines, Experience coordinating enterprise integrations, Knowledge of DDD and clean architecture, Expertise in Java and Spring Boot, Expertise in Python and FastAPI, Expertise in React or Angular, Experience with L2/L3 production support, Experience with Spring Security and OAuth 📃 Skills: Java, SpringBoot, Python, FastAPI, React, Angular, CICD, Observability, ServiceNow, Jira, Boomi, DDD, OAuth, SpringSecurity 🏢 Description: To strengthen the team, we are seeking a person for the role of Senior Software Engineer, responsible for defining technical standards and architecture, guiding the development team, ensuring quality practices, security, and CI/CD, and collaborating with product and operations in delivering scalable, business-aligned solutions. Main Responsibilities Define technical standards, architecture, and development practices, ensuring quality, security, performance, and resilience in production environments. Orchestrate end-to-end delivery with CI/CD and observability while promoting responsible autonomy among teams aligned with product and operational objectives. Implement on-premises, hybrid, and cloud solutions with security best practices, observability, and integration with existing systems. Lead architectural decisions in backend and frontend (Java/Spring Boot, Python/FastAPI, React/Angular), ensuring scalability, consistent testing and code reviews. Implement and maintain end-to-end CI/CD pipelines and observability (logs, metrics, distributed tracing). Coordinate enterprise integrations (ServiceNow, Jira, CRMs) and iPaaS (e.g., Boomi), including mapping, validation, and compliance. Apply practices such as DDD, clean architecture, application security, and secrets/IAM management throughout the development lifecycle. Provide technical guidance to the team, elevate quality standards, and align practices with product and operations. Lead L2/L3 support in production: incident triage and resolution, prevention escalation (on call), SLA management, and conducting post-mortems focused on root causes and corrective actions. Operate and evolve applications in on-premises, hybrid, or cloud environments: VMs, networks, VPNs, certificates, Nginx/reverse proxy, load balancing, and enhancing the security of exposed services. Key Requirements Proven ability to define technical standards, architecture, and development practices. Experience with CI/CD orchestration and observability practices. Knowledge of implementing on-premises, hybrid, and cloud solutions with security best practices. Leadership in architectural decisions for backend and frontend technologies. Experience with maintaining CI/CD pipelines and end-to-end observability. Ability to coordinate enterprise integrations and ensure compliance. Mastery of practices such as DDD and clean architecture. Technical expertise in Java/Spring Boot, Python/FastAPI, React/Angular. Experience in providing L2/L3 technical support and leading incident resolution. Expertise in security practices, including Spring Security and OAuth. Nice to Have Experience with Docker, Kubernetes, and cloud environments (Azure/AWS/GCP). Background in data management with PostgreSQL, MongoDB, and Redis. Familiarity with AI models, particularly RAG. Knowledge in security access management and policies based on the principle of least privilege.

Technology

Link Group

Senior C# Engineer

Senior

Hybrid

Warsaw, Poland

35,000 - 45,000 PLN

🏢 Summary: Senior Software Engineer role focused on building and scaling a high-performance, AI-supported trade management platform using Rust and C#. The position involves designing distributed, event-driven systems with strong emphasis on performance, scalability, and reliability. The engineer will own solutions end-to-end and optimize cloud-based backend systems. 🗂️ Requirements: 5+ years of software engineering experience, Strong system design knowledge, Strong understanding of concurrency, Strong understanding of performance optimization, Hands-on Rust experience in production, Familiarity with C#/.NET, Experience with distributed systems, Experience with cloud platforms (preferably AWS), Practical use of AI tools in development workflows 📃 Skills: Rust, C#, NET, AWS, Kafka, RabbitMQ, AI, DistributedSystems, Concurrency, PerformanceOptimization 🏢 Description: Senior Software Engineer C# (AI-driven systems) We are looking for an experienced Software Engineer who combines strong system design skills with a modern, AI-supported development approach. This role focuses on building and scaling a high-performance trade management platform, primarily in Rust, with some exposure to C#/.NET. You will work on distributed, event-driven systems processing large volumes of data, with a strong focus on performance, scalability, and reliability. Key Responsibilities Develop and optimize high-performance backend systems Design and maintain scalable distributed services Use AI tools to accelerate development, testing, and debugging Own solutions end-to-end, from design to production Improve system performance and cloud efficiency Requirements 5+ years of experience in software engineering Strong fundamentals in system design, concurrency, and performance Hands-on experience with Rust (production level) Familiarity with C#/.NET Experience with distributed systems and cloud (preferably AWS) Practical use of AI tools in development workflows Nice to have Experience with messaging systems (Kafka, RabbitMQ, etc.) Exposure to financial systems or trading environments

Technology

Spire Global

Senior Backend Software Engineer

Senior

Hybrid

Boulder, CO

🏢 Summary: Senior Software Engineer role focused on backend and platform engineering to build reliable, secure systems supporting satellite and ground infrastructure. The position involves developing backend services, improving observability, modernizing container and CI/CD environments, and collaborating with infrastructure and cybersecurity teams. Ideal for engineers experienced in complex, high-reliability production systems and AWS-based cloud environments. 🗂️ Requirements: 5+ years professional software engineering experience, Strong Python skills, Experience with at least one compiled language (Rust, C++, Go, Java, or C), Strong backend engineering experience in production systems, Experience with AWS or GCP environments, Experience with containers and CI/CD pipelines, Proficiency in Linux environments, Strong understanding of systems engineering, performance, and reliability, Experience collaborating with infrastructure and/or cybersecurity teams, Strong communication and documentation skills 📃 Skills: Python, Rust, C++, Go, Java, C, AWS, GCP, Linux, Docker, Kubernetes, Terraform, PostgreSQL, Grafana, Databricks, Elasticsearch, GitHubActions, CI/CD, Observability, Containers 🏢 Description: About the Role Spire is looking for a Senior Software Engineer to help design, build, and improve the reliable and secure systems that support satellite and ground infrastructure. This is a hands-on engineering role focused on backend systems, platform engineering, observability, and infrastructure-aware software development. You'll work closely with software, infrastructure, and cybersecurity teams to implement scalable solutions that support operational reliability and secure engineering practices. We are not looking for a dedicated cybersecurity specialist. Instead, we're seeking an experienced software engineer who has worked alongside security teams and has experience implementing technical requirements related to security, reliability, and compliance within production systems. This role is best suited to engineers who enjoy solving technically challenging problems and have experience working on complex systems in industries such as aerospace, scientific computing, telecommunications, biotech, robotics, or other high-reliability environments. What You'll Do: - Design, develop, and maintain backend services and platform tooling - Collaborate with cybersecurity and infrastructure teams to implement secure engineering requirements - Contribute to modernising services, container environments, and CI/CD workflows - Improve telemetry, monitoring, logging, and observability across production systems - Support cloud-based infrastructure and deployment workflows in AWS environments - Participate in code reviews, architecture discussions, and engineering best practices - Work across infrastructure, backend systems, and operational tooling in a collaborative engineering environment Who You Are: Required Qualifications: - 5+ years of professional software engineering experience - Strong Python development skills - Experience with at least one compiled language such as Rust, C++, Go, Java, or C - Strong backend engineering experience building and maintaining production systems - Experience working with AWS-based environments (GCP experience also considered) - Experience with containers, CI/CD pipelines, and modern software delivery practices - Comfortable working in Linux-based development environments - Strong understanding of systems engineering, performance, reliability, and operational concerns - Experience collaborating with infrastructure and/or cybersecurity teams to implement technical requirements - Strong communication and documentation skills Preferred Qualifications: - Experience working in aerospace, scientific computing, biotech, telecommunications, robotics, or other technically complex industries - Experience in smaller engineering organisations where engineers wear multiple hats across development and operations - Familiarity with Kubernetes, Terraform, and infrastructure-as-code approaches - Experience with PostgreSQL or other relational databases - Experience building observability or telemetry solutions using tools such as Grafana, Databricks, or Elasticsearch - Familiarity with GitHub Actions or modern CI/CD tooling - Exposure to secure software development or DevSecOps practices - Advanced academic research background (PhD/Postdoc) in engineering, physics, computer science, or related scientific disciplines What We're Looking For: We value engineers who are: - Curious and adaptable - Pragmatic problem-solvers - Interested in complex systems and infrastructure - Open to learning new technologies and domains - Collaborative and low-ego - Comfortable operating in fast-moving engineering environments This role is primarily backend and platform focused. We are generally looking for engineers with stronger backend and systems experience rather than heavily frontend-focused backgrounds. Spire operates a hybrid work model, and this position will require you to work a minimum of three days per week in the office. Access to US export-controlled software and/or technology may be required for this role. If needed, Spire will arrange the necessary licenses—this is not something candidates need to have before applying. Salary Range $130,500—$171,000 USD Global Perks - Name Your Satellite Program (NYSP) Launch Attendance - Generous Time Off Policy - Education Assistance Program - Employee Assistance Program (EAP) - Employee Stock Purchase Program (ESPP) - Family Leave - Fitness Reimbursement - Employee Referral Program - Healthy snacks & beverages in every office

Technology

Link Group

Senior Software Engineer C# (AI-driven systems) | Senior Rust Engineer (AI-driven systems)

Senior

Hybrid

Warsaw, Poland

35,000 - 45,000 PLN

🏢 Summary: Senior Software Engineer role focused on building and scaling a high-performance, AI-supported trade management platform using C#/.NET in distributed, event-driven systems. The position involves designing scalable backend services, optimizing performance, and collaborating in a multi-language environment including Rust. Strong emphasis is placed on scalability, reliability, and cloud efficiency in production systems. 🗂️ Requirements: 5+ years of software engineering experience, Strong C#/.NET experience in production systems, Strong system design fundamentals, Strong knowledge of concurrency, Experience with performance optimization, Experience with distributed systems, Experience with cloud platforms (AWS preferred), Willingness to work with Rust, Practical use of AI tools in development workflows 📃 Skills: C#, .NET, Rust, AWS, Kafka, RabbitMQ, AI, Concurrency, DistributedSystems, Cloud, SystemDesign 🏢 Description: Senior Software Engineer C# (AI-driven systems) We are looking for an experienced Software Engineer with strong expertise in C#/.NET and a solid understanding of system design, who is open to working in a modern, multi-language environment that includes Rust. This role focuses on building and scaling a high-performance trade management platform processing large volumes of data in distributed, event-driven systems. You will work in a highly performance-driven environment with a strong emphasis on scalability, reliability, and engineering excellence. Key Responsibilities Develop and optimize high-performance backend systems in C#/.NET Design and maintain scalable distributed services Collaborate with Rust engineers and contribute to cross-language system development Use AI tools to accelerate development, testing, and debugging Own solutions end-to-end, from design through to production Improve system performance, reliability, and cloud efficiency Requirements 5+ years of experience in software engineering Strong C#/.NET development experience (production systems) Strong fundamentals in system design, concurrency, and performance Willingness to work with Rust in a production environment Experience with distributed systems and cloud platforms (preferably AWS) Practical use of AI tools in development workflows Nice to have Experience with messaging systems (Kafka, RabbitMQ, etc.) Exposure to financial systems or trading environments 2. Senior Rust Engineer (AI-driven systems) We are looking for an experienced Rust Engineer who combines strong systems engineering skills with a modern, AI-supported development approach, and is open to working in a multi-language environment that includes C#/.NET. This role focuses on building and scaling a high-performance trade management platform processing large volumes of data in distributed, event-driven architectures. The environment is highly performance-focused, with strong requirements around scalability, reliability, and low-latency processing. Key Responsibilities Develop and optimize high-performance backend systems in Rust Design and maintain scalable distributed services Collaborate with C# engineers and contribute to cross-stack system design Use AI tools to accelerate development, testing, and debugging Own solutions end-to-end, from architecture to production Improve system performance, reliability, and cloud efficiency Requirements 5+ years of experience in software engineering Strong Rust experience in production systems Strong fundamentals in system design, concurrency, and performance Willingness to work with C#/.NET in a hybrid environment Experience with distributed systems and cloud platforms (preferably AWS) Practical use of AI tools in development workflows Nice to have Experience with messaging systems (Kafka, RabbitMQ, etc.) Exposure to financial systems or trading environments

Technology

Grid Dynamics Poland

Senior Site Reliability Engineer (SRE)

Senior

Hybrid

Warsaw, Poland

100 - 128 PLN

🏢 Summary: Senior Site Reliability Engineer role focused on ensuring reliability, performance, and resilience of enterprise products by bridging infrastructure and software engineering. The position involves hands-on Java/Spring Boot code fixes, Kubernetes-based container operations, incident response, and proactive architecture improvements. The engineer drives automation, observability, and security best practices across the SDLC. 🗂️ Requirements: 5+ years experience in SRE or Platform Engineering, Strong proficiency in Java, Strong proficiency in Spring Boot, Experience with Hibernate, Experience with Jenkins, Ability to read, analyze and fix application code, Hands-on experience with Docker, Hands-on experience with Kubernetes, Deep knowledge of Linux systems, Strong understanding of networking, Experience with distributed systems, Experience with monitoring and observability tools, Bachelor’s degree in Computer Science, Systems Engineering or equivalent experience 📃 Skills: Java, Spring, Hibernate, Jenkins, Docker, Kubernetes, Linux, Networking, Prometheus, Grafana, Splunk 🏢 Description: We are looking for an experienced Senior Site Reliability Engineer to join our team and oversee the reliability, resilience, and performance of our core enterprise products. In this role, you will bridge the gap between infrastructure operations and software engineering. You won't just react to alerts - you will proactively analyze system architecture, build automation, and dive deep into the application code (Java/Spring Boot) to fix bugs and eliminate issues at their root. Responsibilities: Architecture & Reliability: Understand the end-to-end product topology from both infrastructure and application perspectives. Identify bottlenecks, scale limitations, and unstable components, driving long-term resolutions before they impact production. Incident Response & RCA: Respond to outages, provide L3 on-call technical support (on rotation), and perform blameless Root Cause Analysis (RCA) to implement permanent fixes. Hands-on Engineering: Address defects, perform code bug fixes directly in production, and recommend architectural improvements during incident analysis. Security & Vulnerability Management: Oversee vulnerability management for applications and containers, manage patching processes, ensure compliance, and monitor certificate expirations and renewals according to global best practices. SRE Advocacy & SDLC: Represent the SRE organization in design reviews, capacity planning, and operational readiness exercises. Partner closely with development teams to embed reliability best practices early in the SDLC. Automation & Mentoring: Build automation tools to reduce manual toil and improve efficiency. Spread SRE culture, create standard documentation, and provide technical mentorship to junior team members. System Health: Oversee the production environment by tracking availability, applying learnings from observability tools, and becoming a Subject Matter Expert (SME) on core issuing products. Min requirements: Experience: 5+ years of experience in Site Reliability Engineering (SRE) or Platform Engineering roles. Software Engineering: Strong proficiency in Java, Spring Boot, Hibernate , and Jenkins. Ability to read, analyze, and fix application code. Containerization: Hands-on expertise with Docker and container orchestration using Kubernetes . Infrastructure: Deep knowledge of Linux systems, networking, and distributed architectures. Observability: Strong understanding of monitoring, logging, and observability tools (e.g., Prometheus, Grafana, Splunk). Education: Bachelor’s degree in Computer Science, Systems Engineering, or equivalent practical experience. Soft Skills: Excellent problem-solving abilities and strong communication skills. Would be a plus: Infrastructure as Code & Cloud: Hands-on experience with tools like Terraform or Ansible, alongside familiarity with major public cloud providers (AWS, GCP, or Azure). Advanced Networking & Service Mesh: Knowledge of service mesh technologies (e.g., Istio, Linkerd) for traffic management, security, and observability in microservices architectures. Industry Experience: Previous background in the FinTech, payments, or banking sectors, with an understanding of high-security compliance standards (e.g., PCI-DSS). We offer: Opportunity to work on bleeding-edge projects Work with a highly motivated and dedicated team Competitive salary Flexible schedule Benefits package - medical insurance, sports Corporate social events Professional development opportunities Well-equipped office About us: Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI , supported by profound expertise and ongoing investment in data , analytics , cloud & DevOps , application modernization and customer experience . Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.

Technology

emagine Polska

IT - Software Engineer RFP-261440-1

Mid

Hybrid

Pune, MH, India

🏢 Summary: Software Engineer role focused on building and maintaining scalable cloud-based applications with modern frontend technologies and DevSecOps practices. The position involves full software development lifecycle responsibilities, frontend architecture, observability, and collaboration with cross-functional teams in a hybrid work environment. Candidates should have strong experience with JavaScript/TypeScript, frontend frameworks, cloud concepts, and software engineering principles. 🗂️ Requirements: Bachelor's degree in Computer Science, Software Engineering, or related field, 5+ years of professional software development experience, JavaScript, TypeScript, React, Angular, Vue, HTML5, CSS3, Browser APIs, Accessibility, Cloud deployment concepts, Azure, Data structures, Algorithms, Software design principles, SQL, NoSQL, MSSQL, MySQL, MongoDB, GitHub, GitLab, DevSecOps, Frontend architecture, Problem-solving, Fluent English 📃 Skills: JavaScript, TypeScript, React, Angular, Vue, HTML5, CSS3, Azure, SQL, NoSQL, MSSQL, MySQL, MongoDB, GitHub, GitLab, Docker, Kubernetes, CI/CD, Scrum, Kanban 🏢 Description: Summary The Software Engineer will play a key role in designing, developing, and maintaining high-quality software solutions that meet the evolving needs of users and the business, while collaborating with cross-functional teams and contributing through the entire software development lifecycle. Main Responsibilities: Design, develop, test, deploy, and maintain robust, scalable, and high-performance software applications using cloud components. Develop reusable UI components and frontend architecture (design systems, state management, routing, bundling). Write clean, efficient, and well-documented code following best practices. Collaborate with product managers, UX/UI designers, and engineers to define, design, and ship new features. Debug and resolve technical issues, ensuring optimal application performance and reliability. Contribute to architectural discussions and decisions, shaping the future of the technical stack. Stay up-to-date with emerging technologies and industry trends to improve development processes and tools. Participate in code reviews, technical documentation, and continuous improvement of engineering standards. Apply DevSecOps practices: dependency management, vulnerability scanning, secrets handling, and secure coding. Establish observability for frontend applications (real-user monitoring, client-side logging, error tracking, performance monitoring). Identify and implement opportunities for automation and process improvements. Key Requirements: Bachelor's degree in Computer Science, Software Engineering, or related field, or equivalent practical experience. 5 years of professional experience in software development. Strong proficiency in JavaScript/TypeScript and modern frontend frameworks (React, Angular, or Vue). Strong understanding of web fundamentals: HTML5, CSS3, browser APIs, security basics, and accessibility. Familiarity with cloud deployment concepts (preferably Azure) and environment configuration. Solid understanding of data structures, algorithms, and software design principles. Experience with SQL and/or NoSQL databases (MSSQL, MySQL, MongoDB, etc.). Experience with version control systems (GitHub/GitLab). Strong problem-solving skills and ability to troubleshoot complex issues. Excellent communication and interpersonal skills. Ability to work independently and manage multiple priorities. Nice to Have: Experience with Docker and Kubernetes. Familiarity with CI/CD pipelines and GitHub actions/workflows. Experience with agile development methodologies (Scrum, Kanban). Proficiency with Agentic IDEs and "Agent in the loop" workflows. Experience in modern CSS architecture and ability to consistently implement design docs. Other Details: Location: Remote (hybrid model) Work Model: Hybrid (3 days in office/week) Work hours: CET Time zone (9 hours) Expected language skill: Fluent English

Technology

Link Group

Senior Site Reliability Engineer

Senior

Hybrid

Warsaw, Poland

170 - 230 PLN

🏢 Summary: The role focuses on ensuring reliability, scalability, and performance of large-scale cloud-based applications by building and maintaining resilient infrastructure. You will manage AWS cloud environments, Kubernetes clusters, and CI/CD pipelines while implementing monitoring, automation, and incident response processes. The position emphasizes Infrastructure-as-Code, observability, and continuous reliability improvements. 🗂️ Requirements: 5+ years experience in SRE, DevOps or similar role, Strong experience with AWS cloud services, Experience with Infrastructure-as-Code tools, Hands-on experience with Kubernetes, Proficiency with Docker, Experience with CI/CD pipelines, Solid knowledge of PostgreSQL or Amazon RDS, Strong SQL knowledge, Knowledge of networking concepts (VPC, DNS, troubleshooting), Strong Linux/Unix administration skills, Experience with observability tools, Experience with automation in infrastructure, Experience with incident management 📃 Skills: AWS, Terraform, Pulumi, Kubernetes, EKS, Docker, GitHub, PostgreSQL, RDS, SQL, VPC, DNS, Linux, Unix, Prometheus, Grafana, Datadog, Dynatrace, CI/CD 🏢 Description: We are looking for an experienced Site Reliability Engineer to ensure the reliability, scalability, and performance of large-scale cloud-based web applications. You will work closely with software development, cloud operations, and platform teams to build and maintain resilient infrastructure and improve system stability. Key Responsibilities: Design and maintain monitoring, alerting, and incident response systems to ensure high availability Collaborate closely with engineering, product, and architecture teams Build and manage cloud infrastructure using Infrastructure-as-Code (e.g., Terraform, Pulumi) on AWS Operate and optimize Kubernetes environments (e.g., EKS) Develop and maintain containerized applications using Docker Improve CI/CD pipelines and drive automation across deployment processes Implement and manage observability tools (logging, metrics, tracing) Participate in incident management, postmortems, and reliability improvements Support capacity planning, disaster recovery, and system scaling Contribute to security, compliance, and operational best practices Develop automation and AI-driven solutions for monitoring and incident prevention Requirements: 5+ years of experience in SRE, DevOps, or similar roles Strong experience with AWS cloud services and Infrastructure-as-Code tools Hands-on experience with Kubernetes and containerized environments Proficiency in Docker and CI/CD pipelines (e.g., GitHub Actions) Solid understanding of databases (e.g., PostgreSQL, Amazon RDS) and SQL Knowledge of networking concepts (VPC, DNS, troubleshooting tools like dig/traceroute) Strong Linux/Unix administration skills Experience with observability tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) Familiarity with automation and AI-based solutions in infrastructure Strong problem-solving and incident management skills