Live opening · Posted 4 days ago

Research Engineer, Privacy and Anonymization

Clera · San Francisco
Ashby No FullTime
You are 4 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 4 days ago
CompanyClera
LocationSan Francisco
Job typeFullTime
Work modeNo
SourceAshby
Listed4 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
7 min from Ashby publishing this role to us finding it
6 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,116 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT THE ROLE
This Research Engineer role sits at the intersection of privacy engineering and AI data infrastructure. You will own the full pipeline for detecting and removing sensitive information from raw, real-world data before it flows into processing, training, evaluation, and synthetic data workflows. The work is foundational: protecting privacy without destroying the structure and signal that make data valuable for training frontier AI agents.
WHAT YOU'LL DO
- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, and design appropriate transformations based on data type and downstream use case.
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
- Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
- Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
WHAT WE'RE LOOKING FOR
- 2+ years of experience building reliable production data or ML systems in Python.
- Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content.
- Experience building end-to-end data processing pipelines without a fully prescribed roadmap.
- Strong experimental instincts with the ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
- Solid understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation, and when each is appropriate.
- Experience designing systems that are robust to schema drift, unusual formats, and edge cases.
- Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is a strong plus.
- Experience with low-latency or high-throughput ML inference and data processing systems is a plus.
- Prior work with sensitive data in healthcare, finance, security, or related domains is a plus.
LOCATION
This role is on-site in San Francisco, California, USA. Visa sponsorship is available.

Employment type
FullTime

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App