Live opening · Posted 18 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for an SDE II Gen AI Engineer who can build production-grade AI systems with a strong focus on agentic AI, multi-agent workflows, LLM applications, evaluations, clustering, open-weight model exploration, Python, FastAPI, and PyTorch. The candidate should be hands-on, strong in engineering, and comfortable building AI agents that can reason, use tools, call services, maintain context, and deliver reliable outputs in production.
Responsibilities:
Build AI agents that can understand user intent, plan tasks, call tools/APIs, use memory, and generate structured responses.
Design and develop multi-agent workflows where specialized agents handle tasks such as retrieval, reasoning, personalization, safety, evaluation, summarization, and orchestration.
Build reusable agent skills/tools that can be plugged into an agent orchestration layer.
Work on LLM-based applications for user query understanding, content generation, summarization, classification, recommendation, personalization, and decision support.
Explore and benchmark open-weight model architectures, including decoder-only LLMs, encoder models, vision-language models, diffusion models, embedding models, and multimodal architectures.
Build backend AI services using Python and FastAPI for agent execution, model inference, prompt generation, tool calling, evaluations, and orchestration.
Work with PyTorch for model loading, inference, experimentation, fine-tuning, embeddings, and optimization.
Build clustering and semantic grouping pipelines for content deduplication, topic discovery, user-interest grouping, article/event grouping, personalization, and retrieval improvement.
Use embeddings, vector search, similarity scoring, and clustering techniques to improve retrieval, recommendations, personalization, and agent context selection.
Build automated AI evaluation pipelines for LLM, GenAI, and agent outputs.
Define and track evaluation metrics such as relevance, factuality, fluency, coherence, safety, consistency, tool-call accuracy, reasoning quality, and intent match.
Implement LLM-as-judge, embedding-based evaluations, golden datasets, regression checks, and CI/CD quality gates.
Design prompt templates, prompt versioning, structured outputs, fallback strategies, model routing, and guardrail logic.
Optimize agent workflows for latency, cost, quality, safety, reliability, and scalability.
Integrate AI agents with backend services, APIs, queues, databases, vector stores, storage systems, and observability tools.
Work on agent memory, session state, context management, tool registry, planning, retry logic, and error handling.
Collaborate with backend, ML, product, and platform teams to ship scalable AI-powered product experiences.
Requirements:
Strong programming experience in Python.
Hands-on experience with FastAPI.
Working knowledge of PyTorch.
Good understanding of LLMs, prompts, embeddings, vector search, and evaluations.
Understanding of agentic AI, tool calling, RAG, memory, planning, and multi-agent workflows.
Familiarity with open-weight model architectures and experimentation workflows.
Basic understanding of clustering, similarity search, semantic grouping, and recommendation systems.
Ability to build clean APIs and production-grade AI services.
Good debugging, problem-solving, and system-thinking skills.
Good to Have:
Experience with agent frameworks such as LangGraph, CrewAI, AutoGen, Semantic Kernel, or custom agent orchestration systems.
Experience with model serving, quantization, LoRA/fine-tuning, batching, caching, or inference optimization.
Exposure to multimodal AI, image generation, video generation, vision-language models, or diffusion-based systems.
Familiarity with Kubernetes, Docker, GCP/AWS, Pub/Sub/Kafka, CI/CD, and observability.
Experience with evaluation tools or custom eval frameworks for LLMs and agent systems.
Understanding of AI safety, content moderation, guardrails, and responsible AI practices.
Experience
2-5 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.