Live opening · Posted 5 days ago

Member of Technical Staff | Inference Platform

Jobgether · Brazil (Remote)
Linkedin Yes
You are 5 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 5 days ago
CompanyJobgether
LocationBrazil (Remote)
Work modeYes
SourceLinkedin
Listed5 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
13 min from Linkedin publishing this role to us finding it
10 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
74,236 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | Inference Platform based in Brazil.
This role focuses on the infrastructure that powers reliable, scalable machine learning inference across cloud and customer environments.
You’ll own the systems that execute models for both large-scale batch workloads and real-time APIs.
Your work will directly influence availability, latency, throughput, GPU utilization, and the cost of every prediction.
You’ll build and evolve Kubernetes-based inference infrastructure while solving challenging problems in scheduling, scaling, recovery, security, and isolation.
The role combines hands-on systems engineering with a strong product mindset, treating compute efficiency and operational reliability as core product features.
You’ll work across model serving, batch execution, data processing, GPU optimization, and customer-hosted infrastructure.
It’s an opportunity to take broad ownership of production ML infrastructure where technical decisions have measurable customer and business impact.
Accountabilities
Evolve and operate a Kubernetes-based online and batch inference runtime supporting production machine learning workloads.
Run large-scale batch inference through ephemeral jobs, implementing multi-dimensional admission control across CPU, memory, and GPU resources.
Build and extend Kubernetes controllers and custom resources to support reliable model execution and scheduling.
Optimize model inference engines and feature-processing pipelines using efficient, vectorized, and columnar operations.
Develop efficient mechanisms for serving graphs and data from Lance-based storage.
Own the execution of training, post-training, and fine-tuning jobs across both cloud infrastructure and customer-hosted Kubernetes environments.
Design and improve autoscaling strategies, GPU serving, inference performance, and infrastructure cost efficiency.
Implement comprehensive telemetry and monitoring for model execution, enabling performance, reliability, and cost optimization.
Solve complex infrastructure challenges including deterministic job sizing, checkpointing, recovery of batch workloads, difficult input files, and automated profiling of newly accepted models.
Design robust approaches to per-customer encryption, workload isolation, and secure execution across shared and customer environments.
Improve serving availability, online inference latency, batch throughput, GPU utilization, and cost per prediction or training job.
Ensure training and batch workloads complete reliably and on schedule without requiring manual intervention or repeated retries.
Write production-quality code, participate in rigorous code reviews, and take operational ownership of the systems you build.
Requirements
Professional experience operating model-serving infrastructure or large-scale batch compute workloads on Kubernetes.
Strong experience building Kubernetes controllers, operators, or comparable Kubernetes-native infrastructure.
Strong software engineering skills with production-quality Python and experience developing reliable, maintainable systems.
Demonstrated ability to profile and optimize data-intensive Python pipelines and identify performance bottlenecks.
Understanding of distributed systems, workload scheduling, resource allocation, and production infrastructure.
Experience with performance optimization across compute, memory, storage, and GPU resources.
Strong cost-awareness and the ability to treat infrastructure efficiency as an important product requirement.
Experience operating production systems and willingness to take responsibility for the reliability and performance of systems you build.
Strong engineering judgment, problem-solving skills, and ability to work effectively on open-ended infrastructure challenges.
Familiarity with ML inference workloads and the operational requirements of serving models at scale.
Experience with Ray, Ray Serve, or KubeRay in production is a strong advantage.
Knowledge of Kueue or other batch scheduling and admission-control technologies is beneficial.
Experience with GPU serving and performance optimization is a plus.
Familiarity with Arrow, Parquet, Lance, or other columnar data formats is advantageous.
Experience shipping and operating software on customer-hosted Kubernetes environments is valuable.
Experience with GCP or AWS and platforms such as GKE or EKS is a plus.
Experience working in financial services or other regulated environments is beneficial.
Benefits
Full-time, fully remote position based in Brazil.
Opportunity to own critical production infrastructure powering both real-time and large-scale batch machine learning workloads.
Broad technical scope across Kubernetes, ML inference, distributed systems, GPU infrastructure, data processing, and cloud platforms.
Hands-on opportunity to build and evolve Kubernetes controllers, scheduling systems, autoscaling infrastructure, and model-serving platforms.
Direct impact on measurable engineering and business outcomes, including availability, latency, throughput, GPU utilization, and cost per prediction.
Opportunity to solve complex infrastructure challenges involving reliability, recovery, security, encryption, isolation, and multi-tenant execution.
Exposure to cloud and customer-hosted environments, including production Kubernetes deployments.
Engineering culture centered on ownership, production quality, operational responsibility, and measurable outcomes.
Opportunity to work on infrastructure where compute efficiency is treated as a core product capability.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App