Live opening · Posted 6 days ago

Staff Machine Learning Engineer

ServiceNow · Santa Clara, California, United States
Smartrecruiters No Full-time
You are 6 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 6 days ago
CompanyServiceNow
LocationSanta Clara, California, United States
Job typeFull-time
Work modeNo
SourceSmartrecruiters
Listed6 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
4 min from Smartrecruiters publishing this role to us finding it
10 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,939 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Team Overview
We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms. This isn't a thin LLM wrapper. We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our from-scratch, industry-leading agent execution runtime). Our systems are backed by a first-party MCP surface into admin/product/rules/transaction systems and interoperate with other AI agents over the A2A protocol. Below that sits a conventional Java/Spring Boot microservices fleet and a React/TypeScript + Lit frontend that the agents ultimately drive.
Role Overview
We're looking for someone who already operates at a Senior-Staff bar in the agentic/LLM domain but is building out breadth across the rest of the stack. You'll be one of the most senior technical voices on how agentic systems get designed here — state management, tool boundaries, streaming protocols, prompt/context architecture, and multi-agent coordination — while staying credible end-to-end: able to read a Spring Boot service, unblock a frontend integration, or reason about a classical ML model pipeline when the problem calls for it.
What you get in this role:
Multi-agent orchestration — LangGraph/LangChain agents over frontier LLMs for transaction editing, conversational configuration, and multi-product quote planning with plan/approve/refine loops and parallel task execution
The Harness — we're crystallizing our own agent execution runtime into an industry-leading, state-of-the-art harness. Full-duplex sessions where a user can interrupt, redirect, or answer a clarifying question mid-execution while other work keeps streaming, built on a from-scratch async runtime rather than a bolted-on wrapper around someone else's agent loop. This is as much a performance and UX problem as a backend one — low-latency streaming, backpressure, live progress, partial results, graceful cancellation — and it's the part of the stack we're most invested in owning outright. You'd be a primary owner of where this goes next.
MCP as a secondary interface — we maintain a first-party MCP server and clients into our admin/product/rules/transaction systems, but as the Harness matures it becomes the primary way our own agents interact with the platform, with MCP kept as the secondary, standards-based surface for external interop. You'd help decide what stays MCP-first and what moves onto the Harness.
A2A protocol — agent-to-agent task delegation and streaming, surfaced through an external gateway so other systems (including core ServiceNow) can drive our agents directly
Forward Deployed Engineering — expect real time embedded with customer- and product-facing teams against live deployments. Adapting the Harness and our agents to actual customer workflows under real constraints, not just building platform capability in the abstract
RAG / context engineering — tenant-uploaded document ingestion, categorization, and aggregation into agent context. Prefix-cacheable prompt design for cost/latency
Classical ML, when the problem isn't a good fit for an LLM — we have a separate PyTorch/scikit-learn training and serving pipeline (field-value prediction) that a whole-stack ML engineer should be able to read, extend, or evaluate against LLM-based alternatives
Full-stack fluency — enough comfort in Spring Boot/Java services and the React/TypeScript + Lit frontend to unblock an integration end-to-end without waiting on a handoff
To be successful in this role you have:
8+ years building production software, with several years specifically shipping LLM-powered / agentic systems (not just API wrapper calls — real tool-use loops, state management, multi-turn orchestration)
Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
Deep, hands-on expertise with LangGraph and/or LangChain (or the judgment to know when to skip them and hand-roll something better)
Strong understanding of MCP — ideally having built an MCP server, not just consumed one
Production async Python (FastAPI, asyncio) — comfortable with WebSockets, streaming, and the failure modes of long-lived stateful connections
Track record of making real architecture decisions on agent systems — tool boundaries, context/state design, cost/latency tradeoffs, when a bounded agent beats a fully autonomous one
Enough range outside Python to read/modify a Spring Boot service and a React component without hand-holding — this is explicitly a whole-stack role, not "Python specialist with a frontend allergy"
Comfort operating with ambiguity and setting technical direction, not just executing a spec — this is a Staff-level bar on judgment
A demonstrated habit of pulling the latest from the industry — new agent frameworks, protocol standards, model capabilities — into production quickly
Willingness and ability to build real fluency in the Core ServiceNow platform. Our systems increasingly need to interoperate with it directly, and this role is expected to help drive that, not just react to it
Willingness to work directly with customers/deployments as part of Forward Deployed Engineering efforts — this isn't a purely internal-platform role
Desired Qualifications
Experience with classical ML (PyTorch/scikit-learn) in addition to LLM-based systems
Experience with A2A or other agent-to-agent interop protocols
Experience with RAG / knowledge-graph systems (embeddings, vector or graph-based retrieval)
Multi-tenant SaaS experience, especially around per-tenant isolation of stateful connections/resources
CPQ, quoting, or transaction/pricing domain experience
Prior Forward Deployed Engineering experience, or time spent embedded with customers shipping bespoke solutions exposure
For positions in this location, we offer a base pay of $176,100 - $308,200, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Employment type
Full-time

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App