Senior/Lead DevOps Engineer
- 5+ years of experience in DevOps/Cloud Engineering, with strong, hands-on experience deploying and operating AI/ML models and LLM-based applications in production
- Strong scripting/automation skills (Python, Bash, Terraform)
- Experience with IAM, network security, and compliance frameworks in regulated industries (banking/finance preferred)
- Familiarity with Kubernetes, including deploying, operating, and troubleshooting production clusters, is a nice-to-have
- Exposure to NVIDIA GPU infrastructure/tooling (e.g., NVIDIA NIM, Run: AI) and on-premises infrastructure is a nice-to-have but not required
- Cloud provider certifications (Architect, Engineer, Network Engineer level) are a plus
- Demonstrated LLMOps experience, including model versioning, inference optimization, production monitoring and evaluation, and ongoing lifecycle management of deployed models (required)
About the role
You'll join a cloud engineering team responsible for designing and deploying AI/ML infrastructure and deployment pipelines for high-profile, security-conscious clients.
Our client delivers secure, large-scale cloud and AI infrastructure for the enterprise and banking sectors, helping them modernize legacy systems and safely adopt AI/ML technologies.
This is a US-based role (New York City, Jersey City, Lake Mary, FL, or Pittsburgh) with an anticipated base salary range of $113,120–$141,400 annually.
SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.
,[Design, deploy, and maintain secure cloud infrastructure (landing zones, IAM, network security, encryption/key management) across one or more major cloud providers, Deploy, serve, and optimize AI/ML models and LLM-based applications in production, including inference pipelines, model-serving infrastructure, and scaling of AI workloads, Build and operate container platforms (Kubernetes preferred) supporting AI/ML and application workloads, Build and maintain CI/CD pipelines and infrastructure-as-code (Terraform, Jenkins/GitLab), Implement cost optimization strategies across production and non-production environments, Ensure compliance and security best practices for handling sensitive data in regulated industries, Mentor engineers and act as technical point of contact for client engagements] Requirements: DevOps, Cloud, AI, Python, Bash, IAM, Network Security, Kubernetes