Senior Site Reliability Engineer Linux & Virtualization Platform

Link Group REMOTE 2026-09-03

Who We're Looking For (Your Profile):

  • You possess an expert-level understanding of Linux internals and system administration. You are comfortable navigating the complexities of the Linux kernel, drivers, and the underlying hardware they run on.
  • You have deep, practical experience with virtualization technologies, particularly KVM/QEMU.
  • You are a skilled diagnostician, able to systematically debug complex issues and determine if the root cause is in the physical test environment, hardware firmware, BIOS/UEFI, a kernel driver, or a user-space application.
  • You are a proficient software developer, with strong skills in Python and shell scripting, capable of independently building automation frameworks and operational tools.
  • You have significant experience in a Senior SRE or Development role, preferably working on large-scale, distributed systems where reliability is paramount.
  • You are experienced with modern infrastructure-as-code and configuration management tools like SaltStack, Ansible, Chef, or Puppet.
  • You have excellent communication skills and a collaborative mindset, with a proven ability to work effectively across different engineering teams.

We are looking for a deeply technical Senior Site Reliability Engineer to become a core member of our Virtualized Host Platform (VHP) team. Your mission will be to ensure the absolute reliability, performance, and scalability of the foundational compute platform that powers our global services.

This is not a typical cloud SRE role. This is for engineers who are passionate about Linux internals, virtualization, and the complex interplay between software and hardware. You will be the ultimate owner of the host environment, using your expertise to automate, debug, and build one of the most resilient platforms on the planet.

,[Become a Master of the Host: You will develop an expert-level understanding of our entire virtualized host stack, from the physical hardware and BIOS/UEFI, through the Linux kernel and KVM/QEMU virtualization layer, to the software and services running on top., Write Code to Solve Infrastructure Problems: You will leverage your advanced skills in Python and shell scripting to build robust automation and tooling. Your code will solve complex operational challenges, automate large-scale infrastructure management (using tools like SaltStack or Ansible), and eliminate manual work., Build Proactive Observability: You will design and implement sophisticated monitoring and observability systems that go beyond simple metrics. Your goal is to identify and resolve potential issues in the kernel, drivers, or hardware before they can impact our customers., Lead the Toughest Investigations: You will be the ultimate escalation point for the most complex, system-level issues. You will collaborate with support, operations, and engineering teams, using your deep diagnostic skills to pinpoint the root cause of problems, whether they lie in the hardware, firmware, or software., Drive Reliability and Stability: You will participate in a 24/7 on-call rotation, guiding the team through service-impacting incidents and ensuring rapid restoration. You are not just fixing problems; you are architecting a more resilient system for the future.] Requirements: Linux, Virtualization, KVM, Python, Shell, SRE, SaltStack, Ansible, Progress Chef, Communication skills Additionally: Private healthcare, Sport subscription, Foreign languages classes, Life Insurance, Cafeteria system.