Unlock Full Resume Report

New offer - be the first one to apply!

September 4, 2026

Senior Site Reliability Engineer (Data Platform)

Senior • Remote

19,000 - 22,000 PLN/yr

Warsaw, MZ, Poland

Description

Join a critical global-infrastructure team responsible for a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers data to customers for analytics and reporting. This role focuses on the reliability of large-scale data systems, observability, and complex problems across software, networks, and infrastructure.

What You Will Do

  • Engineer world-class data systems: Focus on the reliability and performance of massive data platforms. Work extensively with large-scale distributed databases to ensure data integrity, availability, and low-latency performance. The environment heavily uses Cassandra and MongoDB; expertise with other modern NoSQL or distributed data systems is also valued.
  • Become the ultimate troubleshooter: Serve as the highest technical escalation point for complex reliability and performance issues, leading investigations across the global stack.
  • Build insightful observability: Design and build the observability fabric for real-time platform health, using Prometheus and Grafana to create actionable insights.
  • Solve problems with code: Build high-quality automation and internal tooling with Python, Go, or Java. Streamline operations, automate diagnostics, including with AI assistance, and enable self-service workflows for other teams.
  • Drive long-term reliability: Partner with Engineering, Product, and Network teams to influence architecture, identify systemic weaknesses, and deliver scalable solutions that prevent future incidents.

Qualifications

  • Deep hands-on experience engineering and operating large-scale distributed data systems, including data-at-scale work, metrics analysis, and troubleshooting complex data-integrity issues.
  • Practical experience with modern NoSQL databases; Cassandra or MongoDB experience is strongly preferred, though similar technologies are considered.
  • At least 5 years of experience as an SRE or Systems/Infrastructure Engineer managing mission-critical distributed systems.
  • Programming experience building automation and tooling in Python, Go, or Java.
  • Strong Linux/Unix administration and low-level system-troubleshooting foundation.
  • Strong root-cause-analysis mindset and ability to deliver robust, production-grade solutions.

Similar jobs you might like