Live opening · Posted 5 days ago

Member of Technical Staff | Observability & Reliability

Jobgether · United States (Remote)
Linkedin Yes
You are 5 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 5 days ago
CompanyJobgether
LocationUnited States (Remote)
Work modeYes
SourceLinkedin
Listed5 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
9 min from Linkedin publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
73,753 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | Observability & Reliability based in Brazil.
This is a high-impact platform engineering role focused on making distributed systems observable, reliable, and operationally resilient.
You’ll own observability across cloud and customer-hosted environments, ensuring teams can understand system health wherever workloads run.
The role spans logs, metrics, traces, alerting, SLOs, incident response, and reliability engineering.
You’ll work closely with Kubernetes-based infrastructure and software operating across environments you may not fully control.
Your work will directly support high availability, faster incident resolution, and consistent deployment health.
You’ll also help reduce telemetry costs by improving the quality and efficiency of the signals collected.
This is an autonomous, hands-on position where you’ll build, operate, and continuously improve critical platform capabilities.
Accountabilities
Evolve and maintain the observability platform covering logs, metrics, traces, alerting, and system health across cloud and customer-hosted dataplanes.
Ensure every environment reports critical operational information, including active releases, health status, heartbeats, logs, metrics, and usage to the central control plane.
Implement telemetry collection within customer Kubernetes environments using outbound-only connectivity models.
Detect and investigate differences between desired infrastructure or deployment state and what is actually running in each environment.
Monitor the health and availability of deployment and runtime agents, including ephemeral workloads such as Ray clusters supporting batch inference.
Define and maintain Service Level Objectives (SLOs), establish actionable alerting, and contribute to error-budget practices.
Lead or participate in incident response and postmortems, identifying improvements that reduce recurring failures and mean time to recovery (MTTR).
Coordinate incident resolution across internal teams and customers when fixes involve customer-managed environments.
Optimize telemetry pipelines to reduce redundant data, control infrastructure costs, and improve the signal-to-noise ratio of operational information.
Write production-quality code, review technical changes, and take operational ownership of the systems you build.
Requirements
Deep professional experience with OpenTelemetry and modern observability platforms or backends.
Hands-on experience defining SLOs, working with error budgets, designing actionable alerts, and managing production incidents.
Strong experience with Kubernetes and infrastructure-as-code tools such as Terraform and Helm.
Experience operating software across distributed or customer-hosted environments where infrastructure and connectivity may not be fully under your control.
Strong software engineering fundamentals, with experience producing maintainable, production-ready code and conducting effective code reviews.
Willingness to participate in operational ownership, troubleshooting, incident response, and continuous reliability improvements.
Strong analytical and problem-solving abilities, with an ability to investigate complex distributed-system behavior.
Experience communicating clearly across engineering teams and, when required, working directly with external customers or stakeholders.
Experience with GCP/GKE or AWS/EKS is an advantage.
Familiarity with multi-node or multi-cluster ML workloads in production is a plus.
Experience deploying software to customer-hosted Kubernetes environments, including Helm-based deployments and outbound-only connectivity, is valuable.
Experience in financial services or other regulated environments is an additional advantage.
Benefits
Fully remote role based in Brazil.
Full-time position within an engineering-focused environment.
Opportunity to own critical observability and reliability systems with direct impact on production availability.
Work across cloud and customer-hosted Kubernetes environments, providing broad exposure to distributed infrastructure.
Opportunity to work with modern observability, Kubernetes, infrastructure-as-code, and ML infrastructure technologies.
High degree of technical ownership, with responsibility for both building and operating the systems you develop.
Exposure to complex reliability challenges involving real-time services, batch workloads, customer environments, and distributed systems.
Opportunity to contribute to incident management, platform architecture, and long-term reliability practices.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App