Live opening · Posted 12 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
Architect agentic solutions end to end: task decomposition, tool/function design, memory and state models, orchestration topology (supervisor, hierarchical, sequential, parallel), and human-in-the-loop checkpoints. Design and tune production RAG pipelines — chunking, hybrid retrieval, reranking, query rewriting, metadata filtering, grounding and citation — and own the retrieval quality metrics, not just the pipeline diagram. Build the evaluation layer alongside the system: golden datasets, trace-level assertions, LLM-as-judge rubrics, regression suites that gate deployment. Select and justify the AWS service composition for each workload, with an explicit view on cost, latency, data residency, and failure modes. Define guardrails and the security posture for agents that take real actions: least-privilege tool permissions, prompt injection defenses, PII handling, audit trails. Set engineering standards for the practice — reference architectures, reusable agent and tool patterns, observability conventions — and mentor engineers into them. Partner with client stakeholders on solution shaping, technical discovery, effort estimation, and proof-of-value scoping.
Responsibilities
Architect agentic solutions end to end: task decomposition, tool/function design, memory and state models, orchestration topology (supervisor, hierarchical, sequential, parallel), and human-in-the-loop checkpoints.
Design and tune production RAG pipelines — chunking, hybrid retrieval, reranking, query rewriting, metadata filtering, grounding and citation — and own the retrieval quality metrics, not just the pipeline diagram.
Build the evaluation layer alongside the system: golden datasets, trace-level assertions, LLM-as-judge rubrics, regression suites that gate deployment.
Select and justify the AWS service composition for each workload, with an explicit view on cost, latency, data residency, and failure modes.
Define guardrails and the security posture for agents that take real actions: least-privilege tool permissions, prompt injection defenses, PII handling, audit trails.
Set engineering standards for the practice — reference architectures, reusable agent and tool patterns, observability conventions — and mentor engineers into them.
Partner with client stakeholders on solution shaping, technical discovery, effort estimation, and proof-of-value scoping.
Qualifications
2-3 year experience in Agentic system design
2-3 year experience in Agent frameworks
3 year experience in RAG engineering
AWS AI stack – 3 years of experience
5 years of experience in AWS core services
7+ years of experience in Python engineering
2+ years of experience Evaluation and LLMOps
2+ years of experience AI security and responsible AI
Required Skills
Agentic system design: Planning and execution loops, ReAct and plan-execute patterns, tool/function calling design, short- and long-term memory, state persistence and recovery, multi-agent orchestration and handoff design, MCP and A2A for tool and agent interoperability. Practical judgment on when an agent is the wrong answer and a deterministic workflow is the right one.
Agent frameworks: Deep hands-on experience with at least two of: Strands Agents SDK, LangGraph, CrewAI, AutoGen, LlamaIndex. Comfortable dropping to a custom orchestration loop where a framework gets in the way.
RAG engineering: Production experience with document processing and chunking strategy, embedding model selection, vector and hybrid (BM25 + dense) retrieval, reranking, query decomposition, GraphRAG and agentic RAG patterns, context-window budgeting, and grounding/citation enforcement. Able to diagnose whether a bad answer came from retrieval, ranking, or generation.
AWS AI stack – 3 years of experience:
Amazon Bedrock — Converse API, model selection and routing, Knowledge Bases, Guardrails, Flows, prompt caching, batch vs. real-time inference
Amazon Bedrock AgentCore — Runtime, Memory, Gateway, Identity, Observability, Code Interpreter, Browser
Amazon SageMaker AI for custom model hosting, training, and endpoint operations
Vector and search: Amazon OpenSearch Serverless, S3 Vectors, Aurora PostgreSQL with pgvector, Amazon Kendra.
AWS core services – 5 years of experience: Lambda, Step Functions, ECS/EKS/Fargate, API Gateway, EventBridge, SQS, DynamoDB, S3, CloudWatch, IAM, KMS, Secrets Manager, VPC and PrivateLink. Able to design a VPC-isolated, private-endpoint deployment for a regulated client without help.
Python engineering – 7+ years of experience: Production-grade Python — async, typing, Pydantic, FastAPI, structured testing. Clean, reviewable code; not notebook-only.
Evaluation and LLMOps – 2+ years of experience: Offline and online evaluation design, RAGAS or equivalent retrieval metrics, trace-based observability (OpenTelemetry GenAI conventions, CloudWatch, Langfuse/LangSmith or similar), prompt and model versioning, token and cost governance, drift and regression detection.
AI security and responsible AI – 2+ years of experience: Prompt injection and tool-abuse threat modeling, scoped tool permissions and action approval design, data classification and PII handling, content filtering, auditability. Working familiarity with an AI governance framework such as NIST AI RMF.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.