Infrastructure Engineer GPU & Kubernetes, KRAKÓW BIAŁYSTOK

Who we are looking for:


Infrastructure engineer who's genuinely comfortable at the intersection of GPUs and Kubernetes, someone who's provisioned GPU nodes, debugged scheduling weirdness, and knows why nvidia-smi looking fine doesn't mean your workload is actually using the GPU efficiently.

You'll be working on the infrastructure that underpins testing environments and internal AI/ML tooling.

Must have:

  • 3+ years working with Kubernetes in production, not just spinning up minikube for a demo
  • Hands-on experience with GPU infrastructure: NVIDIA driver/CUDA stack, GPU scheduling in K8s (device plugins, MIG/time-slicing), and diagnosing GPU-related performance issues
  • Solid Linux systems fundamentals: networking, storage, containers, the usual
  • Experience with infra-as-code (Terraform, Helm, Ansible, or similar)
  • Comfortable with a scripting language (Python or Bash) for automation and tooling
  • Familiarity with CI/CD practices and GitOps workflows

Nice to have:

  • Experience with distributed training frameworks (PyTorch DDP, NCCL) or inference serving (Triton, vLLM, KServe)
  • Exposure to telco cloud or network functions virtualization (NFV/CNF) environments
  • Experience with bare-metal GPU provisioning, not just cloud-managed GPU instances
  • Background in a QA/testing or platform reliability context
  • Familiarity with tools like Prometheus/Grafana, ArgoCD, or Rancher

We offer:

  • Professional team
  • International projects
  • Agile methodology
  • Flexible hours
  • Possibility to work hybrid from our office in Krakow or Bialystok
  • Friendly working environment
  • Career development opportunities, skills and proficiency growth
  • Private healthcare
  • Employment on the basis of different types of contracts (B2B/ UoP/Umowa zlecenie)

EPOL IT operates within the EPOL HOLDING capital group.
For over 15 years EPOL HOLDING group has completed over 100 projects that support business and give an advantage over competitors in the field of telecommunications, industry and health care.
Our key domains are Telecommunications, Internet of things, Automation of business processes, Healthcare, Portal solutions and Artificial intelligence.

,[Design, build, and maintain Kubernetes clusters running GPU-accelerated workloads (training, inference, and validation test suites), Manage GPU resource scheduling and sharing (MIG, time-slicing, device plugins) to keep utilization high without workloads stepping on each other, Own the infrastructure lifecycle: provisioning, upgrades, monitoring, cost/capacity planning — for both cloud and on-prem/bare-metal GPU nodes, Build and maintain CI/CD pipelines for infra changes and containerized workloads, Troubleshoot performance issues across the stack, from driver/CUDA versioning to network fabric to pod scheduling, Work closely with our QA and product teams to support Touchstone AI Factory's validation environments, Set up observability (metrics, logging, tracing) for GPU workloads and cluster health, Contribute to infra-as-code practices] Requirements: Kubernetes, GPU, Linux, Bash, Python, Helm, Ansible, Terraform, CI/CD, Docker, Grafana, Prometeus, PyTorch
Data publikacji: 2026-07-27
APLIKUJ