Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
ABOUT THE ROLE
As an AI Evaluation Engineer, you'll own the quality bar for both our AI-powered features and the broader product experience they sit inside. You'll design and run evaluation frameworks that catch regressions in model behavior (accuracy, hallucination, safety, compliance-sensitive edge cases) as well as classic product/QA issues — so that what we ship to customers handling real securities and compliance data is trustworthy every time. This is a hands-on, blended role: part evaluation engineering for LLM/AI systems, part product quality ownership.
WHAT YOU'LL DO
- Design, build, and maintain evaluation frameworks and benchmark suites for AI/LLM-powered features (e.g., document extraction, compliance checks, automated workflows), covering accuracy, consistency, hallucination rate, and safety.
- Define golden datasets, rubrics, and scoring methodologies (human-in-the-loop and automated/LLM-as-judge) to measure model and product quality objectively.
- Build automated eval pipelines that run in CI/CD, flag regressions before release, and produce clear, trackable quality metrics over time.
- Extend evaluation coverage beyond the model layer into full product/QA testing — functional, regression, and end-to-end testing of AI-powered features and the surrounding product.
- Partner with product managers and engineers to translate ambiguous quality bars ("is this good enough to ship?") into measurable, repeatable evaluation criteria.
- Investigate failures and edge cases, perform root-cause analysis across the model/product boundary, and drive fixes with engineering.
- Maintain traceability and reporting on eval/QA results for compliance-sensitive workflows, given the regulated nature of the data we handle.
WHAT WE'RE LOOKING FOR
- 5 plus years of experience in QA/test engineering, ML evaluation, or a related quality-focused engineering role.
- Hands-on experience testing or evaluating AI/LLM-powered features — building eval sets, scoring rubrics, or benchmark harnesses (or strong adjacent automation/QA experience with a demonstrated interest in AI evaluation).
- Solid automation/testing fundamentals: scripting (Python and/or TypeScript/Java), API testing, SQL for data validation, and CI/CD integration.
- Comfort working with LLMs and AI tooling directly — prompt engineering, RAG pipelines, or AI-assisted development tools — either as a builder or a rigorous evaluator of them.
- Strong analytical mindset: comfortable defining metrics for fuzzy, subjective quality questions and defending them with data.
- Clear written communication — you'll be documenting failure modes and quality bars for both engineers and non-technical stakeholders.
- Bonus: experience in fintech, compliance, or another regulated domain where correctness and auditability matter.
NICE TO HAVE
- Experience with vector databases / semantic search (FAISS, ChromaDB, or similar) and RAG evaluation.
- Familiarity with eval tooling/frameworks (e.g., Ragas, DeepEval, promptfoo, custom LLM-as-judge pipelines).
- Background in Agile/Scrum environments and defect-management tooling (Jira or similar).
WHAT WE OFFER
- Competitive compensation based on experience
- Equity participation
- Comprehensive health insurance (self + family)
- Paid leave and wellness benefits
To apply: send your résumé to careers@vinylequity.com.
Employment type
FullTime
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.