Live opening · Posted 9 days ago

Principal Security Researcher, AI Model Evaluation (PhD, Contract)

Cobalt · San Francisco Bay Area (Remote)
Linkedin No
You are 9 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 9 days ago
CompanyCobalt
LocationSan Francisco Bay Area (Remote)
Work modeNo
SourceLinkedin
Listed9 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
7 min from Linkedin publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
70,707 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About the role
Cobalt builds expert data and evaluation infrastructure for AI developers. We are recruiting senior security researchers for a contract project that evaluates how well frontier AI models perform real technical work in a command line environment. Your job is to design original, realistic security tasks that a leading AI model fails to solve. Accepted tasks are used to measure and improve the capabilities of frontier models.
What you will do
Design self-contained command line tasks drawn from real security work, such as vulnerability analysis, exploit development, reverse engineering, cryptographic implementation review, incident forensics, and system hardening.
Build the task environment, write a reference solution, and write automated tests that verify whether a solution is correct.
Test your task against a frontier AI model and refine it until the model fails for substantive reasons rather than because of ambiguity or trick wording.
Respond to reviewer feedback until the task is accepted.
Who we are looking for
A PhD in computer science, security, or a closely related field.
Industry or academic experience in security research or security engineering.
At least one publication, either academic (for example a peer-reviewed paper) or professional (for example a conference talk, a published vulnerability disclosure, or a widely used open-source tool).
Fluency in the Linux command line, shell scripting, Docker, and Python.
The ability to write precise task specifications that another expert could follow without asking questions.
Why Cobalt AI:
Advance frontier AI where it counts. Apply your research expertise to the data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while deepening your understanding of how frontier models are trained and assessed.
Work with a top-tier network. Collaborate with researchers from leading institutions and labs on high-impact, flexible work.
Set your own schedule. Flexible 10 to 40 hour weeks that fit around your research position and your life.
Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App