Live opening · Posted 2 days ago

Site Reliability AI Engineer

Atos · Bangalore | Chennai | Mumbai
Instahyre 6-10 yrs
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyAtos
LocationBangalore | Chennai | Mumbai
Experience6-10 yrs
SourceInstahyre
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
19 min from Instahyre publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,945 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
Design, build, and ship LLM-powered and agentic product features that enhance the team's efforts and outcomes.
Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.
Work on integrating the existing AI tools, and should know major AI frameworks and libraries.
Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritisation, risk decisions and planning.
Architect and continuously optimise the observability of the platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.
Engineer advanced alerting and automation capabilities with Kibana alerting and anomaly detections, and integrating response workflows (routing, runbooks, remediation scripts) to standardise on-call execution and accelerate restoration of services.
Lead incident response for customer-impacting issues across teams, coordination, communications, service restoration, and blameless RCA, then corrective actions that prevent recurrence and reduce operational risk.
Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.
Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimising infrastructure and observability spend.
The core requirements for the job include the following:
Primary Skills:
Observability - ELK (Elastic/Kibana), Prometheus, Grafana, PromQL.
Software and automation - Java, Vertx, Python/Shell/Bash, Rest-SOAP API, Docker containerization, Kubernetes, Kafka.
Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services, and event-driven architecture.
Cross-team coordination, incident triage and resolution, leadership and stakeholder management.
Secondary Skills:
Lang-chain, Langraph, RAG, MCP.
Experience with working on LLM's and integrating with the existing applications.
Python - FastAPI.
Cache - Redis.

Experience
6-10 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App