Live opening · Posted 27 days ago

Infrastructure Developer

Adobe · Noida
Instahyre 6-9 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyAdobe
LocationNoida
Experience6-9 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
93 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,210 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are looking for an experienced Infrastructure Developer (6-9 years) to help design, build, and scale the platform that powers our most demanding ML training workloads. This is a hands-on engineering role where you will write production-grade code, drive meaningful technical initiatives, and contribute to the reliability of an infrastructure that thousands of GPU hours depend on every day. You bring strong Kubernetes skills, solid networking fundamentals, a developer's mindset, and the ability to own projects end-to-end with limited supervision. You have operated systems at a significant scale and are ready to step up into broader technical leadership.
Responsibilities:
Build for scale: Design and improve Kubernetes-native infrastructure that runs distributed GPU training jobs reliably and efficiently. You will own significant components and drive their evolution.
Lead focused initiatives: Own meaningful projects end-to-end, write design docs, gather input from stakeholders, and deliver under realistic timelines, often collaborating with engineers across time zones.
Codify infrastructure: Define and ship cloud infrastructure through IaC (Terraform/Pulumi). Apply the same rigour, testing, and review discipline to infra changes as to application code.
Strengthen observability: Contribute to and extend deep observability stacks, metrics, distributed tracing, log aggregation, SLO/SLI frameworks that surface problems before they become incidents.
Write production code: Build automation, internal tooling, operators, and platform services in Go, Python, or Rust. This is not a YAML-only role.
Own reliability: Participate in incident response, post-mortems, and reliability reviews. Drive systemic fixes, not just workarounds. Be a strong contributor to the on-call culture.
Solve hard networking problems: Debug and resolve complex cluster networking issues, CNI, BGP, service mesh, DNS at scale, east-west traffic, and throughput tuning.
Mentor and grow: Raise the technical bar through code reviews, design feedback, and knowledge sharing with peers and more junior engineers.
The core requirements for the job include the following:
Kubernetes and GPU Infrastructure:
6-9 years in SRE, platform engineering, or infrastructure roles.
Strong working knowledge of Kubernetes internals: scheduler, kubelet, CRDs, operators, admission controllers.
Hands-on experience running GPU/accelerator training workloads in production.
Familiarity with multi-cluster management and workload placement strategies.
Helm, Kustomize, GitOps (Flux/ArgoCD) practical experience and good judgment on when to use them.
Cloud and Infrastructure as Code:
Solid hands-on AWS experience (VPC, EKS, EC2 S3 IAM; TGW a plus).
Production experience with Terraform or Pulumi, modular and tested.
CI/CD for infrastructure: drift detection, plan gating, and rollback strategies.
Working understanding of cost optimisation, reserved capacity, and spot instance management.
Observability:
Prometheus, Grafana, AlertManager production experience, not just lab setups.
Exposure to distributed tracing: OpenTelemetry, Jaeger, or Tempo.
Log aggregation: Loki, Elasticsearch/OpenSearch.
Comfort with SLO/SLI design, error budgets, and multi-tier alerting.
Networking Fundamentals:
Strong TCP/IP, DNS, TLS, HTTP/2 gRPC fundamentals.
Practical experience with CNI plugins: Cilium, Calico, or Flannel and their trade-offs.
Familiarity with service mesh (Istio/Linkerd), ingress controllers, and API gateways.
Ability to debug under load: packet captures, eBPF traces, kernel counters.
Coding and System Design:
Production-quality code in Go, Python, or Rust you ship, not just a script.
Solid grasp of distributed systems fundamentals: consistency, availability, failure modes.
Experience writing Kubernetes operators or working with controller-runtime patterns.
Engaged code reviewer, thoughtful, constructive, and consistent.
Clear technical writer: design docs, ADRs, runbooks that others can actually use.
Collaboration and Ownership:
Has delivered meaningful, cross-functional projects from design to production.
Being comfortable with ambiguity can break down a problem and make progress without a perfect spec.
Experience working async across distributed teams and time zones.
A strong communicator can explain infra trade-offs clearly to peers and partner teams.
Self-driven identifies problems, proposes solutions, and follows through to outcomes.
Bonus Points:
Azure / GCP hands-on experience.
Familiarity with ML training pipeline internals.
eBPF-based observability or networking.
Chaos engineering or game day participation.
Open-source infrastructure contributions.
Security, compliance, or audit exposure.

Experience
6-9 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App