Production Support & Operations Engineer

DCG WARSZAWA GDAŃSK 2026-08-21
  • Minimum 5 years of experience in IT operations, application support (2nd/3rd line), or a similar production-facing role
  • Proven experience managing incidents end-to-end, from alerting and troubleshooting through RCA and prevention
  • Minimum 2 years of experience working within ITIL processes including incident, problem, and change management
  • Experience working in Agile delivery environments alongside development teams
  • Excellent English communication skills (C1)
  • Strong knowledge of Jenkins for building, maintaining, and troubleshooting deployment pipelines
  • Hands-on experience with CI/CD pipelines, automation, and continuous improvement initiatives
  • Excellent troubleshooting and problem-solving skills for complex production issues
  • Proficiency with Splunk and Sysdig for log analysis and alerting
  • Strong understanding of Prometheus and Grafana for monitoring, dashboards, and alert tuning
  • Practical experience operating services running on Kubernetes, including pod health checks, log analysis, and service restarts
  • Expertise in Ansible for controlled configuration changes in operational environments
  • Strong knowledge of Docker and Docker Compose
  • Basic scripting skills in Bash and Python for operational automation and data reconciliation

Nice to have:


  • Experience with IBM Datastage operations
  • Awareness of or willingness to learn Pega and Airflow
  • Experience with Oracle and DB2, including querying, execution plan interpretation, and data incident analysis
  • Understanding of ETL application behavior and REST API communication
  • Experience supporting distributed systems and Kafka-based message flows
  • Java or development background supporting understanding of solutions and integrations

Offer:


  • Private medical care
  • Co-financing for the sports card
  • Constant support of dedicated consultant
  • Employee referral program
,[Own RCA activities for production incidents, including diagnosis, resolution, and preventive actions, Monitor production services, identify anomalies, and proactively address issues before they become incidents, Design, implement, maintain, and improve CI/CD pipelines using GitHub Actions and related tooling, Automate build, test, security scanning, and deployment processes, Oversee day-to-day stability and alignment of Pre-Production and Production environments, Maintain operational documentation, runbooks, known issues, and resolution procedures, Collaborate with development and platform teams to troubleshoot issues and clarify operational requirements, Identify recurring operational pain points and propose automation or tooling improvements, Enhance observability through dashboards, alerts, and log queries, Support service continuity initiatives and participate in disaster recovery exercises] Requirements: Splunk, Prometheus, Kubernetes, Docker, Grafana