Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
DISCO recently brought its bidding and ad-serving stack in-house. The model that decides which offer a shopper sees, and what we bid to show it, is now ours: our features, our training data, our evaluation loops, and our production incidents. That was the hard part. The next part is making it substantially better, faster, and cheaper to run. We're hiring a senior machine learning engineer who is also a data scientist and an AI engineer. You own the models and the signal pipelines that feed them. You own the eval harnesses, the eval loops, and the agent graphs that decide what ships and what gets killed. And you use every AI tool available to move faster than a team this size should be able to move.
The core responsibilities for the job include the following:
Bidding and Ranking Models:
Own DISCO's in-house bidding and offer selection models end-to-end: training, evaluation, deployment, monitoring, and the incident when something goes wrong at 11 pm.
Improve predicted conversion rate modeling and the real-time bid calculation that sits on top of it, where bids are computed at request time against the live pCvR score.
Build DISCO-specific mechanisms the generic approaches don't handle well, including cold start for new advertisers and creative and ad rotation logic that balances exploration against near-term revenue.
Hold the line on serving performance. Model quality that doesn't survive production latency requirements isn't model quality.
Evaluation Loops, Harnesses, and Graphs:
Build the offline eval harness that rejects bad ideas cheaply before they touch traffic. This is the core loop: dataset, metric, harness, eval, decision.
Design the eval loops that tell us quickly and honestly whether a model change is a real lift or a well-dressed one.
Build the agent graphs and orchestration that carry a model idea from experiment to eval to A/B with minimal human glue.
Own the A/B system that evaluates multiple bidding models concurrently against defined metrics, including RPL, CTR, CPA, and ROAS.
Define and enforce the keep/kill process: what evidence a model needs to graduate, what triggers a rollback, and who decides.
Feature and Signal Pipelines:
Build and operate the pipelines that feed the model: counter and aggregate features, text embeddings, cart and basket-level signals, and identity and transaction attributes from across the network.
Own the feature store and the contract between offline training and online serving, including the parity checks that catch training and serving skew before it costs revenue.
Build the feature graphs and DAGs so new signals compose cleanly instead of becoming one-off plumbing.
Continuously test new signals against baseline. Ship what lifts. Write up and kill what doesn't, on a documented timeline, without sunk-cost attachment.
Network Intelligence for GTM:
Build data-driven models that surface where the network is strong and where it's weak, by publisher, category, advertiser, and shopper segment.
Turn those findings into something Sales, Customer Success, and the Publisher team can act on: which advertisers to pitch for which inventory, where yield is being left on the table, and which accounts are at risk before the renewal conversation starts.
Partner with Finance on the unit economics questions that only network-level data can answer.
AI Enablement and Agent Tooling:
Extend DISCO's tooling so a non-ML user can build a model, run it in a production A/B, read the results against the baseline in a visual interface, and iterate without ML support.
Own the guardrails on that: what a non-ML user can and cannot push to traffic, what gets reviewed, and how we prevent an enthusiastic experiment from becoming a revenue incident.
Use AI tools, including Claude, Cursor, and agents, as a genuine multiplier in your own work, and build the harnesses and loops so the whole team operates the same way.
You lead on how this team uses AI. You set the standard for where AI output is trusted and where it gets verified before it ships.
Requirements:
5+ years building and shipping ML models into production, with meaningful time in ads, marketplaces, recommendations, search ranking, or another domain where model quality maps directly to revenue.
You have owned a model in production, not handed one off. You have debugged training and serving skew, watched a metric move the wrong way after a deploy, and been the person who decided whether to roll back.
You think in loops and harnesses, not just models. When you build something, you build the way to evaluate it.
You reach for AI tooling first and hand-write only what needs it. You know where AI output is trustworthy and where it gets verified before it ships.
Strong Python and SQL. Comfortable in production code, feature pipelines, and the infrastructure around model serving, not only in notebooks.
Practical experience with auction and bid mechanics, pCvR or pCTR modeling, or closely adjacent problems. You understand why a model that looks better offline can lose money online.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.