Live opening · Posted 13 hours ago

Data/ML Engineer

Princeton University · Princeton, NJ (Remote)
Linkedin Yes
JobBeeper subscribers received an alert for this role.

At a glance

The key details from the original listing.

Posted 13 hours ago
CompanyPrinceton University
LocationPrinceton, NJ (Remote)
Work modeYes
SkillsPython, Azure, Terraform
SourceLinkedin
ListedPosted 13 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
5 min from Linkedin publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
64,123 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

The Accelerator seeks a Data/ML Engineer to strengthen our data team and advance the engineering, enrichment, and provisioning of the data we collect.
The Accelerator at Princeton includes a portfolio of multiple planned independent and intersecting tools, built on a shared data and compute platform serving computational social scientists at research institutions across North America, Europe, and Africa. The Data/ML Engineer will work within our team to help drive data engineering and machine learning initiatives and collaborations. They will play a crucial role in building and operating the pipelines that transform large-scale social media and web behavior data into research-ready data products, and in developing the machine learning and enrichment capabilities that extend their value. They will work on problems that have no precedent and little source material, requiring novel solutions. They will also be responsible for working with the other teams within the Accelerator and our external partners to help foster collaboration and create an impactful environment for our users.
This is a 6-month term role with potential for extension. A remote work arrangement within the United States may be considered for candidates with the appropriate background and experience.
Strategy
Responsibilities
Work closely with the Accelerator leadership team to align data engineering and machine learning initiatives with overarching goals and long-term vision.
Identify and prioritize development projects that benefit from data engineering and machine learning methodologies and innovations.
Data Engineering
Design, build, and operate data pipelines across the Accelerator's medallion architecture, with end-to-end ownership of transformation layers that serve researchers.
Ensure the accuracy, integrity, and quality of data to be made available through the Accelerator, including data quality validation at pipeline boundaries and enforcement of versioned schema contracts.
Diagnose and optimize distributed data processing workloads at production scale.
Develop deployment automation, CI/CD, and release processes for data products, including versioned data releases and researcher-facing change documentation.
Machine Learning And Data Science
Design, develop, and operate ML and NLP enrichment pipelines over large-scale text and behavioral data, including language identification, translation, and topic and content classification.
Own the full lifecycle of enrichment models: selection, evaluation against labeled data, batch inference architecture, cost efficiency, and reprocessing and versioning strategy.
Develop ML-ready feature layers and data products to support advanced research use cases.
Evaluate and apply large language model workflows and other emerging AI methods where they demonstrably improve outcomes, with attention to their validity for downstream scientific analysis.
Apply statistical analysis and modeling to characterize datasets, estimate coverage, and support research design.
Platform Operations And Cost Engineering
Contribute to cost attribution, visibility, and governance across institutional workspaces, including cluster policies, budget controls, and storage lifecycle management.
Design data and ML workloads to operate within the platform's cost governance framework.
Develop automation for workspace and project provisioning as institutions and research projects onboard.
Operate within Unity Catalog governance, multi-tenant isolation, and research data security requirements.
Research And Collaboration
Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
Modern Software Engineering Foundations: agile (Scrum), DevOps, CI/CD, code review, and pair programming, with working knowledge of cloud compute platforms to support collaborative, scalable, and efficient development.
Author and maintain researcher-facing documentation and provide direct technical support to research users of the platform.
Collaborate with research teams to define data products, sampling frames, and enrichment requirements, and apply state-of-the-art techniques to ongoing scientific challenges.
Stay current with the latest advancements in data engineering, machine learning, and relevant fields to continuously innovate.
Build strong relationships with external partners, driving collaborations that enhance the Accelerator's scientific impact.
Qualifications
Essential Qualifications:
3+ years of relevant experience as a data engineer, machine learning engineer, or data scientist, which may include graduate research and internship experience, with a record of building production systems that operate reliably at scale. Experience working in a remote, agile environment.
Bachelor's degree or equivalent in a relevant field.
Strong proficiency in Python and SQL, and hands-on experience with distributed data processing (e.g., Apache Spark) on large data volumes.
Experience building, evaluating, and operating machine learning or NLP pipelines, including batch inference.
Working knowledge of cloud data platforms.
Strong communication and interpersonal skills to effectively collaborate with researchers in the field, other engineers at various levels of experience, and administrative and leadership team members.
Preferred Qualifications
Experience with Azure and Databricks, including Unity Catalog.
Experience with infrastructure-as-code (e.g., Terraform), containers, and CI/CD tooling.
Experience with large-scale social media, web behavior, or text-as-data research.
Familiarity with large language model annotation workflows and their evaluation.
Publications in reputable scientific journals or conferences is desirable.
Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.
The University considers factors such as (but not limited to) scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.
If the salary range on the posted position shows an hourly rate, this is the baseline; the actual hourly rate may be higher, depending on the position and factors listed above.
The University also offers a comprehensive benefit program to eligible employees. Please see this link for more information.
Standard Weekly Hours
36.25
Eligible for Overtime
No
Benefits Eligible
Yes
Probationary Period
180 days
Essential Services Personnel (see Policy For Detail)
No
Physical Capacity Exam Required
No
Valid Driver’s License Required
No
Experience Level
Mid-Senior Level
Salary Range
$120,000 to $135,000

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App