Live opening · Posted 3 days ago

Remote Engineering Expert

Turing · Türkiye (Remote)
Linkedin Yes
You are 3 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 3 days ago
CompanyTuring
LocationTürkiye (Remote)
Work modeYes
SourceLinkedin
Listed3 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
12 min from Linkedin publishing this role to us finding it
7 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
45,720 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About Turing:
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.
Role Overview:
We are seeking experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines—including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance.
Job Requirements:
Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience.
Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, Gmsh) combined with strong Python scripting skills.
AI Evaluation & Failure Diagnostics: Hands-on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool-use failures.
Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup.
Job Responsibilities:
Model-Breaking Problem Design: Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
Environment & Simulation Integration: Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes (e.g., misinterpreting simulator feedback, premature design convergence, physically impossible geometries).
Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity or missing information.
Cross-Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the model evaluation pipeline.
Domains:
Electrical Engineering
Mechanical Engineering
Aerospace Engineering
Education & Experience:
Bachelor's degree or equivalent practical experience in any field.
Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required.
Offer Details:
Commitments Required: 40 hours per week with 4 hours of overlap with PST.
Engagement type: Contractor
Engagement Length: upto 24 weeks
Evaluation Process:
Shortlisted candidates will be sent a Job Interest Form.
Finalized talents will go through delivery review & proceeded further accordingly.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App