Live opening · Posted 8 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are a legal-intake advertising operation. We built an internal platform that watches every ad we run, transcribes it, breaks it into scenes, maps viewer retention onto those scenes, tags the creative on a structured genome (hook pattern, spokesperson, claim type, emotional arc, offer framing), and measures all of it against one number: cost per signed case.
On top of that sits what we call the Brain. It diagnoses why a creative is failing down to the exact transcript line where viewers drop off, reviews a brief scene by scene, proposes the next iteration, and records a dated prediction that a nightly job later settles against what actually happened.
It is real, it is running on live spend, and we do not yet know how good it is.
That is the job.
The mandate
Make the Brain something a creative strategist trusts enough to work from every day. Not to replace them: to make one strategist as productive as five, by answering the questions they currently answer on instinct.
Why is this angle working in one market and dying in another
Why does this hook hold and that one lose most of the audience in three seconds
What should the next iteration of this brief change, and what must it not touch
Is this brief actually failing, or is the sample too thin to say
What you would actually do
Prove it. Build out the evaluation layer until we can state the Brain's accuracy with a straight face: diagnosis quality against expert judgment, verdict accuracy against realized cost per signed case, calibration of its stated confidence. We already have a prediction ledger that settles bets nightly and throws out any bet a human contaminated. Turn it into a number leadership can read.
Open the genome. Most creative attributes are extracted but not yet published to the performance rollups, because each one is gated behind an inter-rater agreement check. Get the subjective dimensions through that gate, which means making extraction consistent enough to survive it.
Put real statistics under it. Signed cases are rare events. At genome depth our slices have single-digit counts, and today we simply drop thin rows. Replace that with partial pooling so a thin slice borrows strength from its parent instead of disappearing, and so the Brain can say "probably" without lying.
Decompose the diagnosis. Today it is one large prompt. Make it a system: retrieval over past briefs and outcomes, grounding verification per claim, and graceful behavior when the evidence is genuinely thin.
Make it causal. Design and run real creative holdouts so we can separate "this changed and things improved" from "this caused things to improve."
You might be coming from
We care far more about what you have measured than where you worked. These are the backgrounds that map most directly:
Mobile game or app user acquisition. You have run thousands of creative variants, tagged them on a real taxonomy, attributed installs and downstream value back to individual creatives, and predicted long-term value from a thin early signal. This is the closest match on the list
Creative intelligence or ad-creative analytics products. You have already built a version of this, just as a SaaS instead of in-house
Experimentation platforms. You build sequential testing, variance reduction and shrinkage estimators, and you are constitutionally unable to call a win on five conversions
Incrementality and marketing measurement. Geo holdouts, media mix modeling, Bayesian hierarchical models over sparse channel data
LLM evaluation or AI quality infrastructure. Golden sets, LLM-as-judge and its failure modes, calibration, prompt regression testing, shipped as a product
Rare-event risk modeling. Fraud, credit or underwriting. You make confident decisions on imbalanced data where being wrong costs money that day
Recommender systems or ranking. Sparse feedback, cold-start, exploration versus exploitation, and off-policy evaluation from logged data
Growth engineering inside a heavy-spending brand or performance agency. You were the whole data-and-tooling function, you built a scrappy version of this, and you know which parts a creative team actually opens
Lead generation, insurance or legal marketplaces. Same funnel, same rare and valuable end event, same long lag between spend and the outcome that matters
Video understanding. Production pipelines over transcription, scene segmentation and frame analysis
If you see yourself in two or more of these, please apply even if you do not clear every line below.
Required
Language and framework
TypeScript at a high level. The platform is TypeScript end to end. This is not a Python role, and porting the Brain to Python is not on the table
Next.js App Router (we are on 16) and React 19: server components, server actions, streaming responses
Node or Bun runtime fluency. We run on Bun
Data
PostgreSQL and strong analytical SQL: window functions, CTEs, aggregation across a deep dimensional model. The Brain is a data modeling problem wearing an AI costume
pgvector, including HNSW indexes and cosine similarity. Our embeddings, transcript lines, creative segments and insight clusters all live in Postgres
An ORM in the Drizzle or Prisma family. We use Drizzle with hand-managed migrations
Comfort designing schemas, not just querying them
AI and LLM
Production LLM application experience, not notebook experiments: constrained JSON output, strict schema validation, tool calling, agent loops over read-only tools, token and cost control, retries and graceful degradation
RAG done properly: chunking, embedding choice, hybrid retrieval, and above all grounding verification. Every figure and quote the Brain produces must trace to a citation
LLM evaluation. This is the single most important requirement in this post
Multi-model routing. We run through OpenRouter across Gemini, GPT, Claude and Grok depending on the task, and we switch when one gets better or cheaper. Be model-agnostic by habit
Zod or equivalent runtime validation on everything crossing a model boundary
Statistics
Rare-event and sparse-data analysis: hierarchical or Bayesian modeling, empirical Bayes, shrinkage, partial pooling. Concretely, you know why a 40 percent conversion rate on 5 observations means nothing, and you know what to do about it besides deleting the row
Experiment design: holdouts, power, significance, and the discipline to say a result is not real yet
Infrastructure
Docker, and a container-deployed app in production. We run on Railway behind a health check
Background and scheduled jobs: long-running nightly pipelines, streaming responses, and the edge-timeout problems that come with them
Supabase auth or equivalent session and role-based access work
Strong plus
Meta Marketing API, Google Ads API, TikTok Marketing API
Meta Conversions API and offline conversion uploads
Tailwind, shadcn, TanStack Query and Table. You will occasionally build the surface that shows your own output
You have worked alongside creative teams and understand exactly why they distrust dashboards
Not required
A marketing background. We will teach you the domain
A PhD. We care what you have shipped and whether you measured it
How we work
Small team. Direct access to the person spending the money. Everything is judged against cost per signed case. There is a rule in this codebase we do not break: the Brain proposes, a person decides, and it never states a number or a quote it cannot cite.
You will also read roughly 18,000 lines of someone else's careful, well-documented work before you write your own. Bring that kind of temperament.
To apply
Send a short note with:
One AI or measurement system you shipped where you can tell us its accuracy and exactly how you measured it
One time your own evaluation caught your system being confidently wrong, and what you changed
A code sample or repository
Skip the cover letter. We read every note and reply to all of them.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.