Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The candidate will have responsibilities across the following functions:
Inference Optimisation:
Drive TTFT below 400ms for multi-step agent pipelines.
Streaming optimisation: first token to user while sub-agents are still running.
KV cache strategy, prompt compression, and dynamic context window management.
Multi-provider routing: model selection by latency, cost, and task type across OpenAI, Anthropic, Gemini, and open-weight models.
Agent Architecture:
Design and implement Plan-Execute-Synthesise pipelines that run sub-agents in parallel DAGs, not sequential chains.
Build reliable orchestration on top of Temporal: retries, timeouts, partial failure recovery, and idempotency.
Structured output enforcement: JSON schema validation, retry loops on malformed LLM output, and graceful degradation.
Tool call design: schema design that LLMs actually follow reliably across providers.
Evaluation and Harness:
Own the eval framework end to end: ground truth datasets, automated scoring pipelines, and regression detection on every PR.
LLM-as-judge pipelines for qualitative output assessment.
Latency regression testing - p50/p95/p99 tracked across every deployment.
Adversarial test case design: ambiguous queries, missing data, conflicting sources, malformed tool responses.
Infrastructure:
Model serving and cold start optimisation.
Async worker architecture for parallel sub-agent execution.
Observability: trace every token, every tool call, every synthesis step.
Experience
6-10 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.