Live opening · Posted 6 days ago

Member of Technical Staff | Observability & Reliability

Avra · São Paulo
Ashby Yes FullTime
You are 6 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 6 days ago
CompanyAvra
LocationSão Paulo
Job typeFullTime
Work modeYes
SourceAshby
Listed6 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
15 min from Ashby publishing this role to us finding it
12 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
73,819 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT THE ROLE
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.
In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.
WHAT YOU'LL DO
- Evolve our observability stack for logs, metrics, traces, and alerting.
- Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.
- Bring telemetry into customer clusters within a model where agents only make outbound connections.
- Detect drift between the desired state and what's actually running in each environment.
- Monitor the health of our deployment and runtime agents.
- Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.
- Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.
- Reduce telemetry cost: less redundant data, more useful signal.
HOW WE MEASURE SUCCESS
- 99.9% serving availability, with incidents trending down.
- MTTR, including on-premise incidents.
- Near-zero drift between desired and actual state.
- All agents active and reporting, across every dataplane.
WHAT WE'RE LOOKING FOR
- Deep experience with OpenTelemetry and observability backends.
- Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.
- Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).
- Experience operating software in environments you don't fully control.
- Production-quality code and reviews, and a willingness to operate what you build.
NICE TO HAVE
- Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).
- GCP or GKE, AWS or EKS.
- ML multi-node/multi-cluster workloads in production.
- Financial services or regulated environments.

Employment type
FullTime

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App