Senior Site Reliability Engineer AI Hardware & Infrastructure
Who We're Looking For (Your Profile):
- You have a deep background in Site Reliability or Production Engineering, built on a solid Computer Science foundation and proven experience managing large-scale, mission-critical infrastructure.
- You are an exceptional Python programmer. You don't just write scripts; you build scalable, robust operational tools and automation frameworks from the ground up.
- You are a Networking expert. You have a strong, practical understanding of advanced network topologies, high-bandwidth routing and switching, BGP, and the complexities of dual-stack IPv4/IPv6 environments.
- You live and breathe Observability. You have hands-on, expert-level experience with modern monitoring stacks like Prometheus, Grafana, OpenTelemetry, and Loki.
- You are an operational leader. You have extensive experience designing service rollout strategies, defining meaningful alerting thresholds, creating clear technical runbooks, and leading incident response "war rooms."
- You are a natural owner. You thrive on solving ambiguous, complex technical problems and have a proven ability to take a challenge from a vague idea to a production-grade, fully automated solution.
- You are a strong collaborator, able to partner effectively with external data center vendors and coordinate with on-site field technicians to ensure maximum uptime.
We are looking for an elite Site Reliability Engineer to join the core team responsible for our global AI compute infrastructure. Your mission will be to ensure that the physical and virtualized backbone of our AI platform—the bare-metal servers, high-density GPU racks, and the advanced network that connects them—is exceptionally reliable, performant, and scalable.
This is a role for a hands-on engineer who is as comfortable writing Python automation and Infrastructure-as-Code as they are designing BGP routing strategies and collaborating with data center technicians.
,[Become a Master of Fleet Automation: You will write sophisticated tooling and automation in Python to manage the entire lifecycle of our server fleet. Your code will handle everything from initial provisioning and configuration to ongoing maintenance and decommissioning, eliminating manual effort across thousands of machines., Architect Seamless Operational Workflows: You will integrate our core operational systems, connecting platforms like JIRA, Siebel, and PagerDuty through robust APIs. Your goal is to create automated workflows that dramatically reduce the time it takes to resolve hardware and network incidents., Build the Future of Observability: You will design and implement a world-class observability stack tailored for bare-metal and virtualized hardware. This includes building custom telemetry pipelines, creating insightful Grafana dashboards with data from Prometheus, and pioneering the use of AI-driven anomaly detection to predict failures before they happen., Lead in Times of Crisis: As a senior member of the team, you will be a leader during critical incidents. You'll participate in a 24/7 on-call rotation, spearhead the response to high-severity outages, and drive comprehensive, blameless post-mortems that result in concrete architectural improvements., Leverage AI to Build Better Systems: We believe in using our own tools. You will actively use advanced AI utilities and LLM-assisted development to enhance your own technical execution, from generating complex automation scripts to evaluating system performance.] Requirements: Python, Networking, BGP, Prometheus, Grafana Additionally: Private healthcare, Sport subscription, Foreign languages classes, Life Insurance, Cafeteria system.Podobne oferty
Senior Mechatronics Engineer
SoftServe
WROCŁAW
2026-09-11
Senior GenAI Engineer
EY GDS
KATOWICE WROCŁAW
2026-09-11
Senior DevOps Engineer
DCG
GDAŃSK GDYNIA WARSAW ŁODŹ
2026-09-10
Senior Backend Engineer Python&Java with AI/ML skills
GFT Poland
KRAKÓW
2026-09-10
Senior Lead DevOps Engineer Hungary
Deloitte
BUDAPEST
2026-09-09
Senior SRE ACDC Platform Infrastructure Automation
Link Group
REMOTE
2026-09-09
DevOps Engineer Mid or Senior Packer Core &AI
Ericsson
KRAKÓW
2026-09-09
Staff Engineer
Air Space Intelligence
GDAŃSK
2026-09-09
Azure Red Hat OpenShift OpenShift Administrator
ASTEK Polska
REMOTE
2026-09-09
Senior Software Engineer Senior UNIX Automation Engineer
Sopra Steria Poland
KATOWICE
2026-09-09