Live opening · Posted 17 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Auckland, New Zealand | Come for 3–6 Months. Stay for the Adventure | Relocation Support
Tired of Silicon Valley? Come build the future from New Zealand.
What if your next big AI adventure came with better weather, a better lifestyle and a lot more space to think?
We're heading into summer in Auckland - so why not spend the next 3–6 months here while the weather is at its best, working on some genuinely ambitious AI problems? Think long summer evenings, beaches, mountains, great food and an incredible lifestyle - without stepping away from the cutting edge of technology.
We're looking for an exceptional Agent Factory Lead to join us in Auckland and help build the systems that will power how work gets done at MacroActive.
If you're ready to swap the Bay Area grind for Auckland sunshine - without stepping away from the cutting edge of AI, this could be your next move.
And yes, we're serious about the work. We're building an AI factory where agents write, test, review and ship code, analyse data, build campaigns and transform how our business operates.
You won't be watching the future happen. You'll be building it.
Most companies are still asking whether AI agents can write code. We're past that question.
The next one is harder: what happens when agents do most of the producing?
Agents that turn a feature request into tested, reviewed, shippable code. Agents that test that code. Agents that produce the analysis and insight the business runs on. Agents that build campaigns and find operational breakthroughs.
That workforce is the FACTORY. We're looking for the person who will help design it, build it, run it while improving it every week.
You won't be writing our features. You'll be building, maintaining and optimising the agents that do.
About MacroActive
MacroActive is a white-label platform for fitness and wellness creators in 30+ countries. Our creators have earned more than $220M on our platform. We take a revenue share, so we only win when our creators win.
Why this is Interesting?
Data most AI teams don't have. Hundreds of thousands of real coach and client conversations, tens of thousands of correct-form exercise videos, and more than a decade of attribution data linking specific pieces of social content to actual signups, revenue and user behavior. That means you will build evals against ground truth (money earned), not vibes.
Cost is a feature. A sloppy retry loop or an unbatched pipeline is a line item. You'll own the unit economics of everything the factory produces.
Humans have to trust it. The people directing and reviewing the factory's output need to know what it did, why, and when it wasn't sure. Explainability, sensible escalation and knowing when NOT to act matter as much as capability.
You set the standard. There's no inherited playbook. The architecture, the conventions and the bar are yours to define.
What You'll Own??
The Production Line
The agents that do the work:
Code: agents that take a scoped request to working code in a sandbox, with repo context, conventions and guardrails baked in.
Quality: agents that write and run tests, review changes, and catch regressions before a human has to.
Insight: agents that turn our data into analysis leadership can act on, with answers that are verifiably correct, not just plausible.
Campaigns and operations: agents that draft, test and improve marketing and operational work against real outcomes.
The Factory Floor
What every line runs on:
The agent spec. A standard way to define any agent before it's built: how success is measured (in the world, not inside the agent), what it can sense, what it can act on, and what kind of environment it operates in. Then choose the simplest architecture that behaves rationally there.
The toolbelt. Typed tool contracts and MCP servers, with allow-lists, scoped permissions, idempotency keys, and human approval gates on anything that ships to production, moves money or messages a customer.
Orchestration. Deterministic, graph-based workflows for multi-agent chains, on durable execution. No roulette-wheel planning.
Retrieval that holds up. Production RAG: chunking strategy, hybrid search, reranking, and the judgment to know when a graph beats a vector store.
The eval harness. Unit tests for tools, integration tests for chains, simulation for whole systems. Golden sets, LLM-as-judge calibrated against human labels, regression gates in CI.
The runway to production. Per-step tracing and replay, token and cost meters per agent, SLOs on latency, correctness and safety. Shadow, then canary, then guarded rollout, then scale.
Resilience and cost controls. Rate limits, budgets, step caps, backoff with jitter, circuit breakers, caching, batching, and graceful fallback to a human.
Governance engineers don't hate. A registry and manifest standard so leadership can see what every agent does, what it costs and who owns it.
Continuos Improvement
Measure every line's throughput, quality and cost. Find the constraint. Fix it. Repeat. The factory should be measurably better every month.
The part most job ads leave out
Every new way of working meets resistance. Not because people are difficult, but because change costs something before it pays anything back. The factory will be no different.
So we need someone with real agency. Someone who has introduced new systems into established teams, and made them stick. You don't wait for permission or a perfect spec. You make the call, ship it with limited blast radius, and let the results do the talking to the executive owner.
And when someone pushes back (someone always does), you can defend the decision on its merits. With evidence, not seniority. Calmly, in a room full of sceptics. And if the evidence proves you wrong, you change your mind just as fast.
If you'd rather be handed a backlog and wait for consensus, that's completely fine. This just isn't that role.
You're probably who we're looking for if you...
Have shipped at least one agentic system real users depend on, and can explain how you KNEW it was working (and what broke, how and what was improved over time).
Have worked in the post-agentic SDLC, where agents write and review most of the code and humans own intent, specs, architecture and acceptance. You know what that changes about planning, code review, testing and release, and what it breaks.
Have built or heavily customised coding agents used on real codebases: repo context, sandboxed execution, test generation, automated review, and a clear view on where they still fail.
Have introduced a new system or way of working into an established team, and can tell us how you got it adopted when not everyone wanted it.
Are fluent in tool use and function calling with frontier model APIs, and treat prompting and context engineering as engineering: zero-shot, few-shot, chain-of-thought, structured outputs, and when each is worth the tokens.
Have built RAG in production and have real opinions about embeddings, vector databases, reranking and how you evaluate retrieval.
Treat evals as part of the product. If you can't measure it, you haven't built it.
Write solid production software: TypeScript and/or Python, relational databases, queues, APIs, CI/CD. Agents with distributed systems that had a stochastic component.
Reason about quality, latency and cost trade-offs with numbers, not adjectives.
Have built and run systems at real scale, where tens of thousands of people use your applications every day, and sometimes all at once when a live broadcast starts. You know what a thundering herd looks like, and you design for it before it arrives.
Strong signals
Multi-agent orchestration on graph frameworks and durable workflow engines.
Analytics agents (text-to-SQL or similar) where correctness was verified, not assumed.
Automated prompt and context optimisation against evals (DSPy, GEPA or similar), and the judgement to know when it's worth the compute.
Reward design, reinforcement learning or bandit approaches applied to agent behaviour.
Knowledge graphs or GraphRAG, and the judgement to know when they aren't worth t
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.