Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About The Role
The LLM / GenAI Engineer will build production AI systems that combine foundation models, retrieval, structured data, and application services. The role covers the full lifecycle of generative AI features, including RAG pipelines, agentic workflows, model adaptation, evaluation, observability, and deployment.
Working with applied scientists, backend engineers, and product teams, this role will turn ambiguous business requirements into reliable systems with measurable quality, latency, cost, and safety targets. The position is remote and based in Phoenix, AZ, with meaningful ownership over the technical direction of GenAI capabilities.
Key Responsibilities
Design and implement production-grade RAG systems using Python, LangChain, LlamaIndex, or custom orchestration frameworks
Build ingestion, chunking, embedding, retrieval, reranking, and citation pipelines using vector stores such as Pinecone, Weaviate, OpenSearch, or pgvector
Develop agentic workflows that integrate LLMs with internal APIs, tools, structured outputs, and deterministic business logic
Create LLM evaluation systems covering retrieval quality, groundedness, factuality, safety, latency, and cost using offline benchmarks and online monitoring
Fine-tune or adapt open-source and hosted models using supervised fine-tuning, LoRA, QLoRA, prompt optimization, and carefully curated datasets
Deploy and operate AI services on AWS, GCP, or Azure using Docker, Kubernetes, CI/CD, and observability tooling for tracing, logging, and performance monitoring
Partner with engineering and product teams to define technical requirements, conduct architecture reviews, and deliver tested, documented systems into production
What We Are Looking For
3–8 years of experience in software engineering, machine learning engineering, or applied AI, including at least 1 year building or operating LLM-powered applications in production
Strong Python skills with experience developing scalable services, asynchronous workflows, REST APIs, and automated tests
Hands-on expertise with RAG architectures, embedding models, semantic search, vector databases, and retrieval evaluation
Experience using LLM APIs or open-source models such as GPT, Claude, Llama, or Mistral, including structured generation, tool calling, and prompt versioning
Proficiency with at least one cloud platform and production engineering tools such as Docker, Kubernetes, Terraform, GitHub Actions, or equivalent
Working knowledge of LLM adaptation and evaluation techniques, including fine-tuning, LoRA, tokenization, benchmark design, and LLM-as-judge methods
Bachelor’s or master’s degree in computer science, machine learning, data science, or a related technical field; Bonus: experience with model serving frameworks such as vLLM or Triton, GPU optimization, multimodal models, safety guardrails, or ML observability platforms
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.