Live opening · Posted 8 hours ago

Applied AI Engineer

IFS · London, England, United Kingdom
Smartrecruiters Hybrid Full-time
You are 8 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 8 hours ago
CompanyIFS
LocationLondon, England, United Kingdom
Job typeFull-time
Work modeHybrid
SourceSmartrecruiters
Listed8 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
9 min from Smartrecruiters publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,186 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About the AI Software Factory
The AI Software Factory is a new software-engineering function inside R&D. We build the systems that let AI agents do real engineering work on IFS software: design, implementation, testing, review.
We are after an order-of-magnitude improvement in delivery speed. That is a hypothesis we intend to prove or disprove in production, not a slogan. Faster only counts if the software still works, stays secure, and can be supported afterwards.
Our output is software. Orchestration, evaluation, codebase analysis, verification, and the interfaces that connect all of that to the way engineers actually work. Product teams are our users and our pilot partners. The team is new, so the first pilots and much of the technical architecture are still open questions. You would be joining to turn the proposition into something that runs, rather than to inherit a finished system.
One of our core deliverables is a reusable framework for parallel software engineering at IFS. It has to define how a feature gets broken into work several agents can do at once, how dependencies and shared state limit that, what context, tools and guardrails each agent receives, where an engineer reviews or decides, and how separately produced changes come back together as one releasable result. It must cover the full lifecycle: plan, design, build, test, review, document, operate. Parallel code generation on its own is not the goal.
Why this role?
Pointing one agent at one ticket is increasingly common. Building a repeatable framework where many agents plan, build, test and review parts of one outcome at the same time and still produce coherent software is not, and you would be helping define how it works.
You get a meaningful slice of that capability. Clear responsibility for working components and the evidence behind one pilot, with senior architecture support behind you, instead of a queue of disconnected tickets. What you build is production engineering infrastructure: evaluation, regression and codebase-analysis systems used to decide whether an agentic workflow is safe enough to expand. Those results inform whether a pilot proceeds, where human controls stay necessary, and which practices get adopted more widely across IFS.
Our interview process uses realistic work samples from the Factory's problem space, such as agent-generated changes and evaluation evidence, so both sides can look at actual work instead of talking around it.
The problem this role exists to solve
An agent opens a pull request and CI goes green. That proves less than it looks. The agent may have written the tests itself, missed an indirect dependency, or made a change that works alone and collides with what another agent is doing three files away.
So verification splits in two:
Is this agent's output correct? Eval suites that mean something, deterministic checks, human validation where judgement cannot be avoided, and regression tests that keep known failures fixed when a model or prompt changes.
Do many agents' outputs compose? A codebase is a graph of dependencies. Running agents at the same time means understanding blast radius, spotting overlapping work, and verifying the merged result rather than approving a collection of individually green pull requests.
You help build that framework and prove it on real pilots. Principles become working software: work decomposition and dependency models, agent orchestration and isolation, coordination and reintegration, and the checks that show the composed result holds. You own substantial components and the evidence they produce, with architecture direction from a senior engineer. You contribute to the Factory-wide architecture without being expected to define it alone.
What you'll do
You help define and implement the framework for parallel software engineering across planning, design, build, test, review, documentation and operation. For each phase that means making it explicit what agents execute, what engineers review, and what stays human-owned.
Much of the work is mechanism. Turning a feature or engineering objective into a dependency-aware work graph: which tasks can run concurrently, which have to be sequenced, what context each agent needs, and where results have to synchronise. Then the controls that make running them at once safe, which means isolation of concurrent changes, dependency and blast-radius analysis, detection of overlapping work, failure and retry handling, and verification that a batch holds together as a whole.
You build and maintain an eval suite for a real agentic engineering workflow, covering planning, implementation, testing and review. You wire a pilot codebase and its CI pipeline into the Factory's verification harness: ground truth, smoke tests, full-suite execution, reporting someone can actually read. Every agent failure you observe becomes a durable regression test, so it stays fixed across model, prompt and tooling changes.
Some of it is judgement rather than code. Reviewing held-out agent outputs through a structured human-validation process, then working out when automated judging agrees with independent human reviewers closely enough to be worth trusting. Investigating runs that failed or came out strange, and making a defensible call: what failed, why it matters, which check backs the conclusion.
You also work with pilot teams on real delivery. Set the delegate/review/own boundary, watch where parallel work succeeds or falls apart, and turn what you learn into reusable capability instead of one-off pilot fixes.
Your centre of gravity is building and proving the parallel software-engineering framework through one pilot. The reusable capability is the product. The pilot is where its assumptions meet a real codebase, a real delivery workflow and a real engineering team.
In your first six months: take one pilot workflow from an engineering objective through dependency-aware decomposition, parallel agent execution, human review, reintegration and end-to-end evaluation against its real codebase and CI. Deliver reusable framework components, a documented set of known failure modes, and evidence showing where parallelism is safe, where work has to sequence, and why.
What you'll bring
You enjoy building with AI and want to help define a new way of producing software rather than bolt an agent onto the existing one. You see work that can be decomposed, dependencies that constrain it, and the failure paths that show up when several streams run at once. You want real ownership, and you also value architecture guidance and rigorous review. Most importantly, when someone asks “Can these run in parallel?” and “How do you know the result composes?”, your instinct is to show the model and the check.
We are not screening for one programming language or framework. The Factory works across different IFS product codebases, so adaptability and engineering fundamentals matter more than stack matching.
A track record of shipping and maintaining production-quality software, with the judgement to own a well-scoped system from proposed approach through implementation, review and operation. This is not an entry-level role, though we care about demonstrated judgement more than a particular number of years.
Practical experience building with AI agents or tool-using LLM systems, beyond a single chat-completion call. Professional work, open source, research, formal study or substantial independent projects all count.
You can enter an unfamiliar codebase, learn its stack, and reason about work as a dependency graph: what proceeds independently, what shares state, what has to sequence, and what a change can reach.
Strong testing instincts, and the writing to go with them. You ask what evidence separates “this appears to work” from “we have checked the behaviour that matters”, and you can explain what was checked, what failed, and how confident anyone should be in the result.
Nice to have
None of this is required. All of it would help.
Evaluation or testing methods beyond conventional unit tests: experimental design, inter-rater agreement, A/B testing, statistical analysis.
Graph-shaped systems and concurrent work: static analysis, call graphs, monorepo build systems, DAGs, queues, fan-out/fan-in, isolation, partial failure, multiple workers modifying shared state. Agent systems, build systems, CI and data pipelines all pose the same underlying problems.
Familiarity with MCP or another tool-calling or agent-orchestration framework.
Work in regulated, security-sensitive or high-consequence engineering environments, where a working result also has to be demonstrably controlled.
These are genuinely optional. If you have built practical agentic systems, care about proving their behaviour and can grow into the less familiar parts, we would like to hear from you.

Employment type
Full-time

Work arrangement
Hybrid

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App