Kafka Site Reliability Engineer

VERITA HR POLSKA Sp. z o.o. Kraków 2026-10-09
  • Ability to communicate effectively on a technical level with engineers and stakeholders.
  • Strong problem-solving and analytical skills.
  • In-depth understanding of Kafka architecture.
  • Hands-on experience with Kafka cluster maintenance, including implementing changes and recommended fixes to clusters and topics to keep production stable.
  • Experience working in an environment built on infrastructure as code and automation-first principles.

Key Technologies

  • Messaging: Apache Kafka, Confluent Kafka
  • DevOps tools: GitHub, Jira, Confluence, Jenkins
  • Monitoring and observability: Prometheus, Grafana, Splunk
  • Container platform: Kubernetes
• Prestigious position at one of the world’s largest banks.
• Stable, long-term projects.
• Competitive salary with a B2B contract.
• Hybrid work (6 days per month from the office in Cracow).
• Private healthcare and multisport card.
• Personal growth and development opportunities with the possibility to rotate between projects.


This role requires working in conjunction with a globally located team. There will be occasions where activity may need to align with another geographic location.

,[Provide deep technical expertise across the Kafka ecosystem, including brokers, ZooKeeper/KRaft, Kafka Connect and Schema Registry., Deploy and operate Kafka components on container platforms, with hands-on Kubernetes experience (e.g., scaling, upgrades, configuration and troubleshooting)., Administer and operate the Kafka platform end-to-end, including provisioning of Topics and Role-bindings., Configure, deploy and maintain Kafka connectors, including MQ, Splunk, MongoDB, GCP Pub sub, as well as Connect components such as workers, tasks, converters and SMTs (transforms)., Implement monitoring and alerting to ensure platform observability and resilience., Automate operational and maintenance activities using scripting and automation tools., Perform benchmarking, performance analysis and tuning to ensure scalable and efficient data streaming., Lead root cause analysis for production incidents, document findings and drive preventative improvements to enhance reliability.] Requirements: Kafka, Infrastructure as Code, Apache Kafka, DevOps, GitHub, Jira, Confluence, Jenkins, Prometheus, Grafana, Splunk, Kubernetes Additionally: Private healthcare, Sport subscription, Free coffee, Bike parking, Shower, Free beverages, No dress code.