Senior SRE
What We're Looking For
Must have:
- Strong SRE fundamentals: incident management, production support, capacity planning.
- Experience diagnosing and fixing production issues in high-availability environments.
- Python (preferred) plus knowledge of Java or shell scripting.
- Hands-on experience with Prometheus, Grafana, Elasticsearch.
- Experience in on-premises environments (public cloud exposure is limited).
- Strong English language skills.
- High level of self-motivation and ability to work autonomously.
Nice to have:
- Familiarity with Azure DevOps (ADO).
- Experience building or working with AI agents / prompt engineering.
- Background in banking or another environment of similar complexity and criticality (fintech, telco, large-scale e-commerce).
About the Role
Join a newly formed, international SRE team responsible for the reliability and stability of technology platforms supporting global Markets operations. The team works closely with a central observability platform team, offering a unique opportunity to combine classic SRE work with the development of AI-driven solutions.
This role is ideal for someone who enjoys having a real impact on the stability of critical systems, while also wanting to grow into automation and AI-driven operations.
,[Site Reliability Engineering (primary focus), Monitoring, production incident handling, and root cause analysis for trading and market platforms., Diagnosing production issues and implementing permanent fixes to prevent recurrence., Capacity management and support for the stability of on-premises platforms., Working in a highly critical environment, largely autonomously (without an on-site local team lead)., AI & Automation (development-focused aspect of the role), Building AI agents to automate monitoring and system health checks, e.g.:, start-of-day / start-of-week health checks,, pre-trade, post-trade, and settlement monitoring agents,, order flow monitoring., Working on a central observability platform, with dedicated training support for AI agent development., Using Copilot / LLM-based tools to support troubleshooting and operational decision-making.] Requirements: Python, Observability, Grafana, Prometheus, On-premises, Incident Management, Production Troubleshooting, Java, ADO, Azure DevOps