Live opening · Posted 1 day ago

[Remote_VietNam] Senior AI Engineer| English must have

Lighting Labs Lichtdesign · Vietnam (Remote)
Linkedin Yes
You are 1 day behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 1 day ago
CompanyLighting Labs Lichtdesign
LocationVietnam (Remote)
Work modeYes
SkillsPyTorch
SourceLinkedin
Listed1 day ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
19 min from Linkedin publishing this role to us finding it
14 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,379 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About the roleWe build domain-specific agents and models for enterprise clients. This role trains, fine-tunes, and optimises open-weight LLMs and multimodal models for retrieval, reasoning, and workflow automation. You adapt current models to client problems, optimise for quality and cost, and deploy what you build. The role is hands-on and sits between research and production engineering.
ResponsibilitiesModel Development & Fine-tuningResearch, train, and fine-tune open-source LLMs (LLaMA, Mistral, Qwen, Gemma, etc.) for domain-specific tasks.
Train and adapt BERT and other Transformer encoder models (RoBERTa, DeBERTa, MPNet) for classification, retrieval, and embedding workloads.
Implement supervised fine-tuning (SFT), instruction tuning, and preference-based alignment (RLHF, DPO, ORPO).
Run distributed training jobs on Ray / KubeRay clusters with DeepSpeed or Hugging Face Accelerate.
Develop efficient data pipelines for model training: data cleaning, tokenization, chunking, and labeling.
Optimize models for RAG pipelines, grounding responses in canonical data and metadata.
Evaluation, Serving & DeploymentEvaluate models with LangSmith, custom benchmarks, and human-in-the-loop feedback loops.
Deploy optimized models to production environments (cloud, on-prem, or air-gapped setups) using vLLM, SGLang, or TGI.
Route and govern model traffic with liteLLM for multi-model, multi-provider serving patterns.
Collaborate with platform engineers on inference infrastructure: tensor parallelism, continuous batching, KV-cache tuning.
Maintain experiment tracking and ensure reproducibility across training runs.
QualificationsMust-Have Technical Expertise5+ years in applied ML/AI research or engineering, with at least 2 years focused on LLM or Transformer model work.
Strong background in PyTorch, Hugging Face Transformers, and tokenizers. Hands-on with both decoder (Llama-family) and encoder (BERT-family) architectures.
Proven ability to adapt open-source models to real-world, production-grade tasks.
Experience with distributed training using Ray / Ray Train, DeepSpeed, or Hugging Face Accelerate.
Working knowledge of an inference-serving stack (vLLM, SGLang, TGI) and of liteLLM for multi-provider routing.
Deep understanding of training efficiency tradeoffs: memory, throughput, and cost optimization.
Proficiency with AI-assisted development tools (Cursor, Claude Code, GitHub Copilot, or similar).
Research & ExecutionAbility to balance rapid prototyping with rigorous benchmarking and reproducibility.
Strong analytical skills for experiment design and result interpretation.
Excellent documentation and communication of research findings.
Preferred/BonusParameter-efficient fine-tuning techniques: LoRA, QLoRA, adapters.
Quantization techniques: 4-bit/8-bit inference, AWQ/GPTQ, GGUF/GGML optimizations.
Experience training or distilling domain-specific embedding models on top of BERT / MPNet / E5 backbones.
Multimodal training experience (vision + text for document understanding).
Experience running Ray jobs on KubeRay with gang scheduling and GPU-aware autoscaling.
Familiarity with continuous batching, PagedAttention, and tensor-parallel inference in vLLM.
Experiment tracking with Weights & Biases, MLflow, or similar tools.
Serving stacks: vLLM, TGI, TensorRT-LLM, Ray Serve, SGLang.
Strong Vietnamese and English communication skills.
BenefitsCompetitive salary and performance incentives
Applied model work that ships to clients
Access to compute resources for model training and experimentation
Training budget and conference attendance
Flexible work arrangements
A team that expects both rigour and shipped results

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App