Live opening · Posted 14 hours ago

AI Engineer Lead

GoAudits · India (Remote)
Linkedin Yes
You are 14 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 14 hours ago
CompanyGoAudits
LocationIndia (Remote)
Work modeYes
SourceLinkedin
Listed14 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
10 min from Linkedin publishing this role to us finding it
21 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
64,522 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About GoAudits
GoAudits is a fast-growing, profitable, founder-led company trusted by 2,000+ brands and 30,000+ daily users across 97 countries. We help businesses run smarter audits and inspections, and our customers stay because the product genuinely works for them. Now we're building our own AI capability in-house.
AI Lead
GoAudits | Full-Time
About the Role
We are hiring an AI Lead to design, train, adapt, and productionize in-house large language models using open-source stacks. This is a hands-on technical role, not a slide-deck or "vibe coding" position. You will own model architecture decisions, training and fine-tuning pipelines, evaluation rigour, and the engineering needed to run private models reliably.
You should be able to reason about transformers at the level of attention, tokenisation, context windows, alignment, quantisation, and data quality, and then implement that reasoning in Python with reproducible code, not copy-paste notebooks or thin API wrappers.
What You Will Do
Build in-house LLMs: Lead the design and build of private LLMs from open-source bases (Llama, Qwen, Mistral, Gemma, DeepSeek, and similar).
Own the full loop: Data curation, continued pretraining, SFT, preference alignment (DPO / ORPO / KTO / RLHF-style methods), evaluation, and release.
Retrieval where it serves the model: Build RAG systems (embeddings, chunking, rerankers, grounded generation) with measurable quality.
Training and inference infrastructure: Multi-GPU and multi-node jobs, DeepSpeed / FSDP / Megatron-style parallelism, and vLLM / TGI / TensorRT-LLM serving.
Evaluation beyond "it looks good": Define suites covering capability, safety, hallucination, domain accuracy, latency, and cost.
Engineering standards: Readable Python, tests for data and eval pipelines, experiment tracking, model cards, and reproducible training configs.
Mentorship: Review PRs for correctness of training code, not just style, and reject shallow wrapper-only contributions.
Product partnership: Work with product and domain teams to decide when to train, when to fine-tune, and when a smaller specialised model is the right answer.
Required Skills
Deep model knowledge: Transformers, attention variants, positional encodings, tokenizer design, scaling laws, overfitting vs. undertraining, catastrophic forgetting.
LLM training in practice: Continued pretraining, LoRA / QLoRA / full-parameter SFT, mixture-of-experts familiarity, packing, loss masking, curriculum and data mixing.
Python as a systems language: Clean packages, typing, logging, config management (Hydra / pydantic), multiprocessing, and debugging CUDA / NCCL failures.
Core stack: PyTorch (required), Hugging Face Transformers / TRL / PEFT / Datasets, tokenizers, Accelerate. TensorFlow only as secondary.
Serving and efficiency: Quantization (GPTQ, AWQ, GGUF, bitsandbytes), speculative decoding, KV-cache, batching, and cost/latency trade-offs.
Data for models: Filtering, dedup, synthetic data with quality controls, contamination checks, preference and pairwise datasets.
Evaluation: Automatic benchmarks plus human eval design. You do not ship on vibes or a single chatbot demo.
Open-source fluency: You have trained or substantially adapted open models, not only called closed APIs.
Strong fundamentals: Data structures, algorithms, and systems instincts. You can profile a slow data loader or a bad attention kernel as readily as you can discuss architecture.
What We Do Not Want
Vibe coders who only prompt ChatGPT/Claude, paste unverified snippets, and cannot explain why a training run diverged.
API-only "AI engineers" whose entire portfolio is LangChain wrappers around a third-party LLM.
Leads who manage process but cannot open a training script, read a loss curve, or debug a tokeniser mismatch.
Resume-only familiarity with RAG/agents, with no understanding of embeddings, retrieval failure modes, or eval.
Nice to Have
Published work, strong open-source contributions, or shipped internal models with documented eval.
Experience with multimodal or code models.
Kubernetes, Slurm, or cloud GPU fleets (AWS / GCP / Azure) for training clusters.
Security and privacy constraints for on-prem or VPC-only model serving.
Prior people leadership of 2 to 8 ML engineers without losing hands-on depth.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App