Live opening · Posted 11 days ago

Agent Engineer - Evals and Harness

Lexsi Labs · Mumbai Metropolitan Region (Remote)
Linkedin No
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
CompanyLexsi Labs
LocationMumbai Metropolitan Region (Remote)
Work modeNo
SourceLinkedin
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
13 min from Linkedin publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
69,052 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission is to build safety aware autonomous systems in the extreme near term. Our research work spans areas like AI alignment methodologies, interpretability-led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences including ICLR, ICML, WWW, IJCNN, MICCAI and EurIPS. Our labs are located in India (Mumbai and remote), Paris, and London.
We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior.
The Role
Our current sprint on building autonomous systems for complex problems, across software engineering, data science, and AI research, involves building the harness, execution substrate and evaluation system, and each is designed to run inside a customer's environment rather than ours.
This role sits underneath all three agents. The coding agent, the data science agent and the AI engineering agent look different from the outside, but they are the same system underneath: a loop that plans, acts, observes, recovers, and knows when to stop. You will build that shared layer, and the evaluation system that tells us whether any change to it made things better. Both halves matter equally. A harness we cannot measure is a harness we cannot improve.
What you'll work on:
Harness
The core agent loop and its execution model, covering orchestration, concurrency, sub-agent coordination, cancellation, timeouts, checkpointing and resume.
Context and state management for long-running tasks, including retention, compaction, and how an agent reasons over what it has already done.
The shared tool protocol, its schemas, versioning, and the internal libraries every agent team builds against.
The execution substrate. Sandboxed environments that are reproducible, snapshot-able and resource-bounded, and that behave identically for a repository task, a training run and a data analysis.
Failure semantics. Retry policy, partial failure, idempotency, and the distinction between a recoverable error and a task that should stop and hand back.
The trace schema every agent emits, which serves at once as our debugging surface, our training signal, and the audit record our customers keep.
Evals
Task suites for each agent type, built from real work rather than synthetic benchmarks, and the infrastructure to run them at scale in parallel.
Verifiers and scoring. Deterministic checks where the task allows it, model-graded rubrics where it does not, and calibration of the graders themselves.
Regression gates that run on every harness change, with cost and latency accounted alongside quality.
The statistics to say a result is real. Seeds, variance, pass@k, confidence, and knowing how many runs are needed before anyone claims an improvement.
Turning observed production failures into permanent test cases.
Tooling for how we work. Trace inspection, replay, and diffing one run against another.
You will work across all three agent teams and closely with our research team on evaluation design, post-training and interpretability of agent behavior.
What We Are Looking For
This is a systems and infrastructure role. Most of the difficulty here is concurrency, state and measurement, not prompting.
Strong software engineering fundamentals and advanced Python. You have shipped and operated production services, not only written them.
Concurrency and distributed systems. Async execution, worker pools and queues, retries and idempotency, backpressure, cancellation, and reasoning clearly about what happens when a step fails halfway through.
Containers and sandboxing. Docker and OCI internals, resource isolation, reproducible environments, and an understanding of why a job that passes locally fails in a sandbox.
Test and CI infrastructure. You have built test harnesses, run large suites in parallel, and dealt with flakiness as an engineering problem rather than an annoyance.
Measurement literacy. Comfort with variance, sampling and significance, and healthy skepticism toward benchmark results including your own.
Observability instincts. Distributed tracing, structured logging, and building the inspection tooling that makes non-deterministic systems debuggable.
You debug systematically. You find out what actually happened rather than adjusting things until the symptom disappears.
You are comfortable when problems are loosely specified and ownership is assumed rather than assigned.
Experience with evaluation or benchmarking work for AI systems is a strong plus.
Experience with agentic systems, LLM inference and serving, or developer tooling is a plus.
Open source contributions we can read are a plus.
We are hiring several engineers for this team at a range of experience levels, including engineers early in their careers who have strong fundamentals and want to work on agents from the infrastructure side.
We move quickly and expect candidates to do the same. We value substance over polish and execution over rhetoric.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App