Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About The Role
The LLM / GenAI Engineer will design, build, and operate production AI systems spanning retrieval-augmented generation, tool-using agents, model adaptation, and automated evaluation. The role focuses on turning foundation models into reliable products with measurable gains in quality, latency, cost, and user experience.
Working alongside applied scientists, platform engineers, and product teams, this role will own critical components of the GenAI stack—from data and prompt pipelines to inference services, observability, and continuous improvement. The position is remote and aligned with the Atlanta, GA engineering organization.
Key Responsibilities
Design and implement production RAG systems using Python, LangChain, LlamaIndex, or custom orchestration frameworks
Build ingestion, chunking, embedding, retrieval, reranking, and grounding pipelines using vector stores such as Pinecone, Weaviate, Elasticsearch, or pgvector
Develop agentic workflows with tool calling, structured outputs, memory, human-in-the-loop controls, and robust failure handling
Create LLM evaluation frameworks covering offline benchmarks, LLM-as-judge workflows, golden datasets, hallucination detection, and regression testing
Fine-tune and adapt open-source models using supervised fine-tuning, LoRA, QLoRA, prompt optimization, or preference-based methods where appropriate
Deploy and optimize inference services on AWS, GCP, or Azure using Docker, Kubernetes, and APIs such as vLLM or Hugging Face TGI
Instrument production systems for latency, token usage, quality, safety, and cost; establish monitoring, alerting, rollback, and incident response practices
What We Are Looking For
3–8 years of software engineering, machine learning engineering, or applied research experience, including at least 1 year delivering LLM or GenAI systems to production
Strong Python skills with experience building maintainable services, asynchronous workflows, REST APIs, and automated tests
Hands-on expertise with LLM application patterns including RAG, embeddings, vector search, prompt engineering, function calling, and structured generation
Experience with PyTorch, Hugging Face Transformers, model fine-tuning, and inference optimization for open-source or hosted foundation models
Proficiency with cloud and production infrastructure, including AWS, GCP, or Azure; Docker, Kubernetes, CI/CD, and observability platforms
Bachelor’s or master’s degree in computer science, machine learning, data science, electrical engineering, or a related technical field
Bonus: Experience with Ray, vLLM, Triton, distributed training, multimodal models, privacy and safety controls, or enterprise-scale AI platform development
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.