Live opening · Posted 3 days ago

Full Stack AI Engineer (LLM Agents)

Metric Tree Labs · India (Remote)
Linkedin Yes
You are 3 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 3 days ago
CompanyMetric Tree Labs
LocationIndia (Remote)
Work modeYes
SourceLinkedin
Listed3 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
7 min from Linkedin publishing this role to us finding it
24 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
35,957 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Location: Remote (India), working with a San Francisco-based founding team
Engagement: Full-time, long-term, dedicated
Experience: 4–5 years in software engineering, including 1+ year building LLM applications and agents
Openings: 1 (founding engineering hire)
Time zone: Flexible hours, with a daily live overlap with the founders (details below)
About the Client
Our client is an early-stage AI startup in San Francisco, founded by alumni of Harvard and MIT. The company is in stealth. It is building an agent-based product that takes complex, high-stakes paperwork off the hands of US families: working out what applies to them, researching official sources, handling correspondence and following up over weeks, with every step checked and backed by evidence.
Parts of the product are already built and running in test. You will be the third person on the product and the first engineering hire, working directly with the two co-founders to take it into production with real users.
Why This Role Is Different: Product Ownership
This is not a ticket-taking role. As the first hire alongside the founders, you will be expected to own the product, not just the code:
• Take whole parts of the system from design to production and be accountable for how they perform.
• Think about the user. The people using this product are going through a hard time, so the quality of every screen, message and automated action matters.
• Weigh trade-offs, push back on ideas, and help decide what to build next.
• Work on your own for hours at a time, and leave clear updates on what shipped, what is blocked and what needs a decision.
• Help set the engineering standards and practices as the team grows.
What You'll Work On
This is a broad role across the whole product:
• Agents and the agent harness: build and test agents that call tools in bounded loops, with budgets, stop conditions, failure handling and checks on every output.
• Model layer: call models from several vendors, route between them and fall back when one fails.
• Long-running workflows: agent runs that wait days or weeks for a reply and survive restarts.
• Full-stack product: ship features end to end, from the agent to the screen a family sees on phone or desktop.
• Evaluation and observability: build eval suites that show whether each agent is correct and when it is safe for an agent to act without a person approving it; trace every run.
• Data and security: store sensitive personal records securely, with access control on every row.
• Infrastructure: help move services onto AWS and Azure, and build performance-critical services in Go, Rust or Python where TypeScript is not the right tool.
In your first month: ship your first change in week one, build and test agents, and deliver features end to end.
By six months: own one or more parts of the platform from design to production, and help take the product live with real users.
Core Requirements
Agents and AI systems
• At least 1 year of hands-on work building LLM applications and agents, including at least one agent you designed and shipped yourself.
• A working understanding of agent architecture: tool/function calling, planning vs. execution, single vs. multi-agent, context management, human-in-the-loop approvals and guardrails.
• Knowing how agent frameworks work (e.g. LangGraph, OpenAI Agents SDK, Claude Agent SDK, Vercel AI SDK), and when it is better to write the loop yourself.
• Structured output: typed, schema-validated data from models, and handling the cases where the model fails.
• Evaluation: building test sets and graders that measure model correctness over time.
• Designing around interfaces, so a model vendor, cloud or database can be swapped without a rewrite.
Engineering
• 3–5 years of professional software engineering.
• Strong TypeScript across the stack: Node.js on the server and React in the browser.
• Production experience in at least one of Go, Rust or Python.
• Deploying and running services on AWS or Azure.
• SQL and relational data modelling, preferably PostgreSQL.
• Background jobs, queues or workflow engines (e.g. Temporal, Inngest, Celery, Step Functions): work that runs outside a web request, retries and survives failure.
• Secure handling of sensitive personal data: access control, encryption, and keeping secrets out of code and logs.
• Automated testing as a habit, including tests of behaviour that depends on a model.
• Daily use of AI coding tools, and the judgement to review and fix what they produce.
Nice to Have
• Experience on both AWS and Azure; infrastructure as code (Terraform, Pulumi).
• Go or Rust in addition to Python.
• Row-level security in Postgres.
• Research agents that search the web and cite sources; browser automation or voice agents; MCP servers you have built.
• Document processing: OCR and pulling data from scans and PDFs.
• Experience at an early-stage startup or as a founding engineer.
• A public repo, technical write-up or demo of an agent you built. This will make your application stand out.
What We Look For Beyond Code
• An ownership mindset: comfortable with ambiguity and able to own outcomes, not just tasks.
• "Evidence before claims": when you say something works, you show the test, the log or the screenshot.
• Excellent written and spoken English, and comfortable on video (daily stand-up with cameras on, plus short recorded video updates).
• Able to work on your own and communicate clearly across time zones.
• A reliable internet connection, a quiet workspace, and your own laptop (16 GB RAM minimum, 32 GB preferred, disk encryption on).
Working Hours
The founders are in San Francisco. In the first few weeks, the team will try a few working patterns and keep the one that works best. Candidates should be open to any of these (India time):
• 2:30 p.m. – 11:30 p.m.
• 7:00 a.m. – 4:00 p.m.
Engagement Details
• Full-time, long-term engagement through Metric Tree Labs, working only with this client as part of its core team.
• Interview process: an intro call with the co-founders, then a 90-minute technical session where you walk through an agent you built and work on a design problem together. Decisions are made quickly.
• Target start: October 2026. Joining date is agreed on selection.
• Competitive compensation, reviewed after six months.
To apply: send your CV along with a link to an agent you've built (repo, write-up or demo)

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App