Live opening · Posted 4 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Prancer | San Diego, CA / Remote (US) | Full-time
About Prancer
Prancer builds continuous frontier attack validation. Our platform finds attack paths across an enterprise, proves which ones are real, and retests until they are closed. At its core is SwarmHack™, a deterministic, multi-agent autonomous pentesting engine with more than 100 specialized agent capabilities. It covers web and API, Active Directory and Entra, cloud and Kubernetes, Security Service Edge, robotics and IoT, and AI agents and MCP/A2A trust boundaries.
Prancer is a member of Anthropic's Cyber Verification Program. Our multi-model router assigns work to Anthropic Mythos, OpenAI cyber models, private open-weight models, or pure deterministic execution, depending on the task and the customer's policy. We hold three USPTO patents on our Penetration Testing as Code (PAC) framework. We serve enterprise customers including Accenture and NTT Data, and we deliver as SaaS, private cloud, or air-gapped appliance.
Our principle: model intelligence proposes, SwarmHack executes, evidence decides. A finding counts as Exploited only when we have captured proof from the target.
The role
We're hiring AI engineers who have broken into systems and understand how models reason. You'll work on the layer where frontier and open-weight models meet a live exploitation engine. You'll teach models to plan attacks, route them safely to specialist agents, and make sure nothing gets reported without evidence behind it.
What you'll do
Design and improve the multi-model router that decides which model, or which deterministic path, handles each attack task based on efficacy, cost, and deployment policy.
Build agent planning and re-planning logic so SwarmHack adapts as new evidence appears. That includes captured credentials seeding authenticated re-crawls and discovered networks feeding lateral pivots.
Integrate and evaluate frontier and open-weight models (Mythos, GLM, Qwen, and others) for exploit discovery, attack chaining, and remediation guidance.
Build evaluation harnesses against vulnerable labs and our 200-host AWS range to measure exploit success, false positives, and evidence quality over time.
Extend the AI-agent attack battery that tests LLM applications, RAG pipelines, tool use, and MCP/A2A integrations for prompt injection, tool abuse, data exfiltration, and trust-boundary failures.
Strengthen the safety envelope: signed scope, kill switch, non-destructive payloads, and full provenance for every action a model proposes.
Support private and sovereign deployments where models, GPUs, and evidence stay inside the customer's environment.
What you bring
4+ years in software engineering, with meaningful hands-on work building LLM-based systems, agents, or ML pipelines in production.
A real offensive security background, such as penetration testing, red teaming, exploit development, bug bounty, or security research. You know what a working exploit chain looks like.
Strong Python. Go, Rust, or C is a plus.
Experience with agent frameworks, tool calling, structured outputs, and prompt and evaluation design.
Working knowledge of at least a few of these: web/API exploitation, Active Directory and Entra attacks, cloud IAM and Kubernetes, network protocols, or container escape.
Comfort running open-weight models locally or on GPUs (vLLM, SGLang, or similar).
The discipline to treat "the model said so" as a hypothesis, not a finding.
Nice to have
OSCP, OSEP, OSWE, CRTO, or equivalent practical certifications.
Published CVEs, conference talks, or open-source offensive tooling.
Experience with AI red teaming or LLM security (OWASP LLM Top 10, MITRE ATLAS).
Background in reinforcement learning, planning systems, or model fine-tuning.
Familiarity with OCSF, MITRE ATT&CK mapping, or compliance frameworks such as PCI DSS 4.0, NIST CSF 2.0, DORA, and NIS2.
Why Prancer
You'll work directly on frontier-model offensive security as a member of Anthropic's Cyber Verification Program.
You'll join a small, senior team where your work ships to enterprise customers quickly.
Our engine is built on years of development and patented technology, not a thin wrapper around a model.
You'll have real ownership, working directly with the founder, who brings 25 years in cybersecurity.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.