Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are seeking a proactive and detail-oriented SDET 2 to establish and scale the automated testing, evaluation, and quality infrastructure for our multi-agent procurement platform. In an agentic system, quality assurance extends beyond traditional API and UI testing to include non-deterministic AI evaluation, dynamic workflow verification, and reliability in multi-agent coordination.
You will design and build test harnesses to ensure that autonomous agents accurately plan, call external enterprise tools, adhere to safety guardrails, and execute high-stakes procurement workflows with deterministic consistency.
Responsibilities:
Agentic and LLM Evaluation Infrastructure: Build and maintain automated evaluation (evals) frameworks (using tools like DeepEval, Ragas, Langfuse, or custom harnesses) to measure agent reasoning accuracy, tool-calling precision, hallucination rates, and prompt regressions.
Workflow and State Machine Testing: Design automated test suites for complex, asynchronous, long-running workflows validating state persistence, checkpointing, retries, and Human-in-the-Loop (HITL) pause/resume logic across workflow engines (e. g., Temporal, Celery/Redis).
End-to-End Automation: Develop and scale robust test automation frameworks across modern Python backend microservices, async APIs (FastAPI), web interfaces (Playwright/Cypress), and enterprise integration touchpoints (ERP mock layers).
Synthetic Data and Scenario Generation: Generate synthetic enterprise procurement datasets, edge cases, and adversarial prompt injections to stress-test agent guardrails, schema enforcement (Pydantic), and fallback systems.
Performance and Load Testing: Conduct latency benchmarking, rate-limit testing, and load testing across multi-agent graphs to identify bottlenecks in LLM token consumption and parallel tool execution.
CI/CD Integration: Embed automated test gates and regression evals seamlessly into CI/CD pipelines (e. g., GitHub Actions, GitLab CI) to prevent agent drift and API regressions before deployment.
Requirements:
Experience: 4 to 7 years of software testing and automation experience, with strong exposure to backend distributed systems and modern AI/LLM-enabled applications.
Strong Python Proficiency: Deep hands-on coding expertise in Python (PyTest, AsyncIO, Requests/HTTPX) for building custom test frameworks and automation utilities.
API and Integration Testing: Advanced knowledge of REST/gRPC API testing, schema validation, contract testing, and mocking asynchronous third-party enterprise services.
LLM Quality Fluency: Practical understanding of LLM evaluation metrics (groundedness, context precision, answer relevancy, safety guardrails) and familiarity with agentic architectures (ReAct, multi-agent state graphs).
UI Automation: Experience with modern browser automation frameworks (Playwright, Cypress, or Selenium).
Preferred Qualifications:
Prior experience testing durable execution platforms (e. g., Temporal) or distributed state engines.
Exposure to enterprise SaaS, fintech, or procurement and supply-chain domains.
Skills
Automation Testing, Cypress, Playwright, Python, SDET, Selenium, Software Development Engineer in Test
Experience
4-8 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.