Live opening · Posted 11 days ago

SDE - III DevOps

AiDASH · Bangalore
Instahyre 6-10 yrs
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
CompanyAiDASH
LocationBangalore
Experience6-10 yrs
SourceInstahyre
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
1 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,051 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are looking for an SDE-III DevOps who takes end-to-end ownership of the infrastructure that runs our satellite-imagery and ML-inference platform. This is a senior, hands-on, individual-contributor role you will set the bar across the DevOps function, influence technical direction, and lead by example, without managing a team. What sets this role apart at AiDASH is what runs on the infrastructure: a globally deployed platform that ingests [petabytes of satellite imagery], runs [millions of inference requests per day] across CPU and GPU fleets, and serves utilities, transportation, and construction customers on tight SLOs. You will build the systems that underpin that and you will do it with AI as a first-class tool, not an afterthought. You will work closely with developers, QA, security, and product teams to design systems that are reliable, secure, and easy to operate while extending an already mature DevSecOps program.
Responsibilities:
Own scalable, secure infrastructure across AWS, Azure, or GCP using AI coding assistants to accelerate IaC authoring, policy validation, and cost reviews.
Architect and maintain CI/CD pipelines that support rapid, safe deployments, with AI-assisted failure triage, smart test selection, and automated release notes.
Lead container orchestration on Kubernetes for production satellite-data and ML-inference workloads, including GPU scheduling, autoscaling, and model-serving infrastructure.
Own observability standards (SLIs, SLOs, error budgets, alerting) and be accountable for keeping platform availability at our committed SLO targets.
Implement automation using Terraform, Ansible, or equivalent, with AI pair-programming as a normal part of the workflow, not a side experiment.
Define and harden security, secrets management, and access-control practices in partnership with the DevSecOps function.
Establish disaster-recovery strategies and backups for critical systems and prove them with regular game days.
Build or extend internal tooling that uses LLMs to make engineers faster at log triage, runbook drafting, alert summarization, and ChatOps for routine ops tasks.
Participate in a follow-the-sun on-call rotation with the global DevOps team.
Raise the bar on engineering standards, including how the team adopts and governs AI tooling in infrastructure workflows.
Requirements:
6+ years in DevOps, infrastructure engineering, or SRE, with proven ownership of production systems at meaningful scale.
Deep experience with at least one major cloud (AWS, Azure, or GCP) and a working grasp of what the others do differently.
Strong infrastructure-as-code (Terraform or equivalent) and CI/CD pipeline experience (Jenkins, GitLab CI, GitHub Actions, or similar).
Production Kubernetes beyond what I have used with Helm. You can debug a stuck pod, design an autoscaling strategy, and reason about cost.
Strong scripting in Python, Bash, or equivalent.
Hands-on with at least one observability stack (Prometheus, Grafana, ELK, or equivalent) and able to define meaningful SLOs.
Shipped real work using AI coding assistants, IaC, debugging, incident triage, and internal tooling and can speak to where they helped and where they got in your way.
Built or extended at least one internal tool using LLM APIs (or are clearly keen to). A hacky prototype counts.
Comfortable in a fast-moving, engineering-driven environment and good at influencing without authority.
Nice to have:
Production experience with ML infrastructure model serving (Triton, KServe, TorchServe, or similar), GPU workload management, feature stores, or data-pipeline orchestration (Airflow, Argo, or equivalent).
Familiarity with compliance frameworks relevant to critical infrastructure (SOC 2 ISO 27001 NERC, or similar).
Certifications such as AWS DevOps Engineer, CKA / CKAD, or equivalent.
Experience with serverless platforms (Lambda, Cloud Functions, etc. ).
Exposure to managing and tuning production databases (PostgreSQL, MySQL, or NoSQL).

Experience
6-10 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App