Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Lead AI Engineer
Ready to turn bold ideas into real-world impact?
At Genpact, we don’t just adapt to change, we lead it. AI and digital innovation are transforming the way businesses work, and we’re at the forefront of it. Genpact’s AI Gigafactory, our industry-first accelerator, exemplifies how we scale advanced technology solutions to help global enterprises work smarter, grow faster, and transform at scale. Whether tackling complex challenges through large-scale models or agentic AI, our breakthrough solutions tackle companies’ most complex challenges.
If you thrive in a fast-moving, innovation-driven environment, love building and deploying cutting-edge AI solutions, and want to push the boundaries of what’s possible, this is your moment.
Genpact (NYSE: G) is an agentic and advanced technology solutions company. We leverage process intelligence and artificial intelligence to deliver measurable outcomes. With a strong partner ecosystem and decades of client trust, we provide innovative solutions that transform how businesses run. Powered by a team with an active learning mindset and client centricity at its core, we deliver lasting value for the world’s leading enterprises.
Get to know us at genpact.com and on LinkedIn, YouTube, X, and Facebook.
Job Description
Role Overview
We are looking for a skilled, experienced DevOps Engineer at the Manager level to join our Advanced AI Practice within the AI Application Engineering team. This is a hands-on role with progressive team leading opportunities for someone who can own the AI infrastructure, automation, and reliability backbone that takes Generative AI and Agentic AI solutions from prototype to production at scale. You will design and run the CI/CD, cloud, container, and MLOps/LLMOps platforms that our applied AI teams build on, while also mentoring engineers and driving DevOps best practices across the practice.
Key Responsibilities
Platform ownership: Own end-to-end DevOps/MLOps/LLMOps infrastructure for AI application engineering initiatives across on-prem, hybrid, and multi-cloud (AWS/Azure/GCP) environments.
CI/CD & GitOps: Design, build, and govern CI/CD pipelines for application code as well as AI models and agentic workflows; drive GitOps-based delivery using tools such as ArgoCD or Flux.
AI/LLM infrastructure: Enable scalable, cost-efficient model-serving and agentic-application infrastructure — GPU orchestration, autoscaling for LLM inference, model registries, and vector database infrastructure.
Containerization & orchestration: Manage Kubernetes-based platforms (EKS/AKS/GKE), Helm charts, and service-mesh/API-gateway layers for high-availability deployments.
Infrastructure as Code & cloud governance: Drive IaC standards (Terraform/Bicep/CloudFormation), policy-as-code, and FinOps practices to govern cloud and GPU spend.
Observability & SRE: Build monitoring, logging, and tracing for microservices and LLM/agent pipelines (latency, token usage, cost, drift); own incident management and on-call practices.
Security & compliance: Embed DevSecOps practices — image/dependency scanning, secrets management, SBOMs — and align infrastructure with responsible-AI and data-governance requirements.
Team leadership: Manage, mentor, and grow a team of DevOps engineers; define engineering standards, career paths, and hiring plans.
Stakeholder partnership: Work closely with Data Science, Applied AI R&D, product, and client teams to translate solution requirements into infrastructure design and SLAs.
Technology evaluation: Assess and pilot new tools and platforms — agent-orchestration infra, GPU cloud providers, MCP servers/gateways — and provide build-vs-buy recommendations.
Process maturity: Champion an automation-first culture; own disaster recovery, business continuity, and DevOps/MLOps maturity roadmaps for the practice.
Must-Have Skills & Experience (Mandatory)
8+ years of DevOps/Site Reliability/Infrastructure engineering experience, including at least 2 years in a technical leadership or team management capacity.
Strong scripting and automation skills in Python, Bash, and Ansible.
Hands-on expertise with CI/CD tools such as Jenkins, GitLab CI/CD, GitHub Actions, or Azure DevOps.
Deep experience with containerization and orchestration: Docker, Kubernetes (EKS/AKS/GKE), and Helm.
Strong command of Infrastructure as Code: Terraform (mandatory), with working knowledge of CloudFormation/Bicep/Pulumi.
Hands-on expertise in at least two hyperscalers (AWS, Azure, or GCP) with working exposure to the other two.
Solid Git-based version control and GitOps workflow experience (ArgoCD, Flux, or equivalent).
Proven experience supporting production ML/AI workloads, including model-deployment pipelines and model-serving frameworks (e.g., KServe, Seldon, NVIDIA Triton, vLLM, TensorRT-LLM).
Working knowledge of observability stacks: Prometheus, Grafana, ELK/EFK, OpenTelemetry.
Security fundamentals: DevSecOps practices, vulnerability scanning (Trivy, Snyk), and secrets management (HashiCorp Vault, AWS/Azure Secrets Manager).
Strong Linux administration, networking fundamentals, and experience designing HA/DR architectures.
Good-to-Have Skills
LLMOps/GenAI infrastructure: experience serving LLMs at scale, GPU cluster management (Ray, KubeRay, NVIDIA Triton), and token/cost monitoring.
Agentic AI infrastructure: supporting deployment of agent-orchestration frameworks (LangGraph, CrewAI, AutoGen) and Model Context Protocol (MCP) servers/gateways.
Vector database infrastructure: deployment and scaling of Pinecone, Weaviate, Milvus, or FAISS.
FinOps: hands-on experience with cloud and GPU cost-optimization tooling and practices.
Service mesh and API gateway technologies: Istio, Envoy, Kong.
Policy as Code: OPA/Gatekeeper for governance and compliance automation.
Prior experience embedded within a Data Science / Applied AI organization, supporting research-to-production pipelines.
Qualifications
Education: Bachelor's or Master's degree in Computer Science, Data Science, Statistics or a related field.
Relevant certifications: AWS/Azure/GCP DevOps or Solutions Architect, CKA/CKAD, HashiCorp Terraform Associate.
Bonus: Publications, patents or open-source contributions in AI/ML.
Qualifications
Certifications
Required Skills
AI/ML Ops
Language
English (Required)
Language Proficiency -
Proficient - C2
Additional Job Location -
Job Type
Regular
Master Skill List -
Advanced Analytics / AI / ML
Remote Type -
Hybrid
Work Shift -
Day Job (India)
Why join Genpact?
• Lead AI-powered transformation – Drive innovation and solve real-world business challenges that matter
• Make an impact – Help global enterprises solve business challenges that matter
• Accelerate your career – Gain hands-on experience, mentorship, and world-class learning opportunities to stay ahead
• Work with the best – Join 140,000+ bold thinkers and pr
More openings worth a look
Recently tracked roles with full details and direct application links.