Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
We are seeking an Applied AI Engineer, Creative Intelligence to help advance an internal AI-powered platform that analyzes and improves advertising creative performance.
Our platform processes every advertisement we run by transcribing the content, segmenting it into individual scenes, mapping viewer-retention data, and classifying creative attributes such as hook structure, spokesperson type, claims, emotional arc, and offer framing. These insights are then connected to our primary performance metric: cost per signed case.
At the center of the platform is an AI system known internally as the Brain. It evaluates creative performance, identifies potential causes of audience drop-off, reviews creative briefs, recommends future iterations, and records predictions that are later compared with actual campaign results.
This role will be responsible for improving the accuracy, reliability, and statistical foundation of the Brain so that creative strategists can confidently use its insights in their daily work.
Key Responsibilities
Design and expand the Brain’s evaluation framework, including diagnosis quality, prediction accuracy, confidence calibration, and performance against realized cost per signed case.
Transform prediction data into clear, reliable metrics that leadership and creative teams can use to assess system performance.
Improve the consistency of AI-generated creative classifications and establish inter-rater reliability across subjective attributes.
Develop statistical approaches for sparse and rare-event data, including hierarchical modeling, empirical Bayes, shrinkage, and partial pooling.
Build methods that allow low-volume data segments to generate appropriately qualified insights rather than being excluded entirely.
Decompose the current diagnostic workflow into a reliable system incorporating retrieval, claim-level grounding, structured outputs, and citation verification.
Develop and maintain evaluation datasets, prompt regression tests, model-comparison frameworks, and quality-monitoring processes.
Design controlled creative experiments and holdouts to distinguish correlation from causal impact.
Collaborate with creative strategists to ensure that recommendations are practical, understandable, and useful in real workflows.
Contribute to the continued development of a production TypeScript application and its supporting data infrastructure.
Required Qualifications
Software Engineering
Advanced professional experience with TypeScript in production environments.
Strong experience with Next.js App Router and React, including server components, server actions, and streaming responses.
Fluency with Node.js or Bun runtimes.
Experience designing, maintaining, and improving production-grade applications.
The platform is built entirely in TypeScript and currently runs on Bun. This is not a Python-based role.
Data and Database Systems
Advanced knowledge of PostgreSQL and analytical SQL.
Experience with window functions, CTEs, dimensional data models, and complex aggregations.
Strong database schema-design skills.
Practical experience with pgvector, vector embeddings, HNSW indexes, and cosine similarity.
Experience with an ORM such as Drizzle or Prisma.
We currently use Drizzle with hand-managed database migrations.
AI and LLM Applications
Demonstrated experience shipping production LLM applications beyond prototypes or notebook experiments.
Experience with structured model outputs, constrained JSON generation, schema validation, tool calling, and agent workflows.
Strong knowledge of retrieval-augmented generation, including chunking strategies, embedding selection, hybrid retrieval, and grounding verification.
Direct experience developing or maintaining LLM evaluation systems.
Familiarity with golden datasets, LLM-as-judge methods, confidence calibration, prompt regression testing, and their limitations.
Experience managing model retries, fallbacks, token usage, latency, cost, and graceful degradation.
Familiarity with multi-model routing and model-agnostic architectures.
Experience using Zod or an equivalent framework to validate data crossing model boundaries.
Our system uses OpenRouter to route tasks across models from providers including Gemini, GPT, Claude, and Grok.
Statistics and Experimentation
Experience analyzing rare events and sparse datasets.
Understanding of hierarchical or Bayesian modeling, empirical Bayes, shrinkage, or partial pooling.
Strong knowledge of experiment design, holdouts, statistical power, significance, and sample-size limitations.
Ability to communicate uncertainty and avoid drawing conclusions from insufficient evidence.
Infrastructure
Experience with Docker and container-based production deployments.
Experience developing background processes, scheduled jobs, and long-running data pipelines.
Understanding of streaming responses, execution limits, timeouts, and production reliability.
Familiarity with Supabase Auth or comparable authentication, session-management, and role-based access systems.
Our application is deployed on Railway and supported by scheduled nightly pipelines.
Preferred Experience
Candidates may be especially well suited to this role if they have worked in one or more of the following areas:
Mobile application or gaming user acquisition
Creative intelligence or advertising analytics platforms
Experimentation or A/B testing platforms
Marketing measurement or incrementality
LLM evaluation or AI quality infrastructure
Fraud, credit, underwriting, or other rare-event risk models
Recommendation systems or ranking
Growth engineering for a performance-focused brand or agency
Lead generation, insurance, or legal marketplaces
Video transcription, scene segmentation, or content-understanding systems
Experience with the following is also beneficial:
Meta Marketing API
Google Ads API
TikTok Marketing API
Meta Conversions API and offline conversion uploads
Tailwind CSS
shadcn/ui
TanStack Query
TanStack Table
Working directly with creative or performance-marketing teams
A marketing background and advanced academic degree are not required. We prioritize demonstrated experience building, deploying, and rigorously evaluating real systems.
How We Work
We are a small, highly collaborative team working closely with the stakeholders responsible for advertising investment and performance.
Our creative intelligence system is designed to support human judgment—not replace it. The system may propose recommendations, but final decisions remain with the team. Every figure, quotation, and performance claim generated by the platform must be traceable to supporting evidence.
The successful candidate should be comfortable reviewing and understanding an established, carefully documented codebase before introducing significant changes.
How to Apply
Please include the following with your application:
A brief description of an AI or measurement system you personally helped ship, including its measured accuracy and the methodology used to evaluate it.
An example of a time your evaluation process identified that a system was confidently incorrect, including what you changed as a result.
A code sample or repository that demonstrates relevant technical work.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.