Live opening · Posted 3 days ago

Research Engineer - ML Infrastructure

Chai Discovery · San Francisco office
Ashby No FullTime
You are 3 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 3 days ago
CompanyChai Discovery
LocationSan Francisco office
Job typeFullTime
Work modeNo
SourceAshby
Listed3 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
1004 min from Ashby publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
67,898 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT CHAI DISCOVERY
Chai builds the design suite for molecules. We train frontier models that learn the underlying foundations of biochemical structure and interaction, so scientists can move faster and pursue targets that other methods cannot reach.
AI is reinventing life sciences the same way it reinvented software engineering, and Chai is at the forefront of this shift. Leading pharmaceutical companies like Eli Lilly https://endpoints.news/eli-lilly-chai-discovery-sign-ai-software-deal/, Pfizer https://www.forbes.com/sites/amyfeldman/2026/06/04/why-pfizer-and-eli-lilly-are-betting-on-this-13-billion-ai-drug-discovery-startup/, and Novartis https://www.chaidiscovery.com/news/novartis-partnership are adopting our platform to power their drug discovery programs.
We value diverse perspectives and are ready to find greatness in unexpected places.
About the role
Research Engineers on ML Infra make our models train and run performantly, reliably, and at scale by owning the distributed systems that sit underneath every model our researchers ship. As a ML Infra Research Engineer, you will:
- Architect, debug, and optimize the distributed ML training stack across the model, layer, and kernel levels — eliminating runtime and reliability bottlenecks on large GPU clusters.
- Profile end-to-end training runs to find bottlenecks across compute, communication, and storage, and build tooling to monitor throughput, utilization, and uptime across clusters.
- Optimize ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels.
- Work closely with Research Scientists to ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs.
- Own reliability of the training stack: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.
About you
- 4+ years of industry experience working within AI/ML infrastructure teams.
- Proficiency in Python and PyTorch or JAX.
- Strong software systems design skills, with comfort operating across the stack from model code down to kernels.
- Experience with orchestrating GPU clusters and large-scale model training.
- Experience with optimizing ML workloads: parallelism, quantization, CUDA/Triton kernels.
WE OFFER
The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters. We protect & promote a culture of high velocity and ownership. We compensate our team accordingly.

Employment type
FullTime

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App