Site Reliability Engineer, WARSZAWA

  • At least 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, or a similar role.
  • Strong experience working with complex, distributed, production-grade systems.
  • Very good knowledge of observability tools, especially Prometheus, Grafana, Loki, Tempo, and OpenTelemetry.
  • Hands-on experience with Kubernetes and Docker.
  • Practical experience with both cloud and on-premises infrastructure.
  • AWS experience is preferred.
  • Ability to automate tasks and workflows using Python, Bash, Go, or similar scripting languages.
  • Good understanding of CI/CD practices, DevOps culture, and agile ways of working.
  • Experience improving deployment pipelines, operational tooling, monitoring, and recovery procedures.
  • Strong ownership mindset, attention to detail, and proactive approach to reliability improvements.
  • Ability to communicate clearly with technical and non-technical stakeholders.
  • Responsible approach to AI-assisted engineering, including validation, security awareness, critical thinking, and practical use of AI tools to improve daily work.

We are looking for a Site Reliability Engineer to help build and strengthen reliability practices across the technology organization. This role will focus on shaping SRE standards, improving operational maturity, and supporting the stability of business-critical trading and production systems.


The position combines hands-on engineering, automation, observability, Kubernetes, cloud, and on-premises environments. You will work closely with DevOps, Cloud, and development teams to improve resilience, scalability, monitoring, and recovery processes across a complex technology landscape.


The company operates in the financial sector and has a strong AI-oriented culture. Artificial intelligence is used in a practical way to support daily work, automate repetitive tasks, improve efficiency, and speed up delivery. As part of the recruitment process, the candidate’s AI mindset will also be assessed, including openness to using modern AI tools, ability to critically evaluate AI-generated outputs, responsible usage, and readiness to identify areas where AI can improve engineering, operations, automation, and incident management.

,[Help define and promote SRE practices, standards, and operating principles across engineering teams., Improve reliability, scalability, and performance of production systems and trading-related platforms., Build and enhance monitoring, logging, tracing, and observability solutions., Work with tools such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry., Review application reliability requirements within Kubernetes-based environments., Support better configuration of services with regard to performance, cost, resilience, and operational stability., Create automation and internal tools to simplify deployments, health checks, recovery processes, and routine operational tasks., Cooperate with development teams to improve fault tolerance, service ownership, and production readiness., Support the implementation of SRE practices such as SLOs, incident reviews, and blameless post-mortems., Participate in an on-call rotation shared across the team.] Requirements: SRE, DevOps, Kubernetes, Cloud, AWS, Python, Bash
Data publikacji: 2026-06-22
APLIKUJ