Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
We are looking for an AI Research & Solution Engineer to build capabilities around Open AI Models and Sovereign AI, helping enterprises adopt, optimize, and deploy AI within their own private infrastructure.
The role involves researching emerging open models such as Llama, GLM, Qwen, DeepSeek, Mistral, and Gemma, benchmarking them across enterprise use cases, fine-tuning and optimizing them on GPU infrastructure, and building production-ready Sovereign AI solutions and products.
This is a hands-on role combining AI research, model evaluation, model optimization, enterprise solution development, and private AI deployment.
Key Responsibilities
Research and evaluate emerging Open / Open-Weight AI models and understand their capabilities, limitations, architecture, and infrastructure requirements.
Design and execute model benchmarks and evaluations across enterprise use cases and datasets.
Compare open models with closed models based on accuracy, reasoning, latency, throughput, cost, and other relevant metrics.
Architect and build enterprise AI use cases using open models, including RAG, AI agents, copilots, document intelligence, and domain-specific applications.
Optimize open models for specific customer, industry, and business requirements.
Deploy and optimize models on NVIDIA GPU infrastructure, focusing on GPU utilization, inference latency, throughput, concurrency, and cost.
Experience with technologies such as PyTorch, Hugging Face, vLLM, TensorRT-LLM, CUDA, quantization, LoRA/QLoRA and related AI tools is a must.
Design and implement Sovereign AI solutions for enterprise customers using open models and dedicated/private AI infrastructure.
Translate customer business, data privacy, security, and AI requirements into appropriate Sovereign AI architectures.
Deploy AI models and applications within customer data centers, private clouds, or dedicated GPU environments.
Work across the AI stack including GPU infrastructure, containers, Kubernetes, model serving, RAG, vector databases, AI agents, APIs, and enterprise applications.
Integrate AI models with enterprise data sources and business systems while maintaining appropriate data privacy and security controls.
Work closely with customer IT, infrastructure, security, data, and application teams during AI deployments.
Build reusable capabilities for Private AI / Sovereign AI / Open Model Cloud platforms.
Convert successful research and experiments into production-ready AI solutions and products.
Qualifications & Skills
3–6 years of relevant hands-on experience in AI/ML, with significant practical experience in LLMs and Open Models.
Bachelor's or Master's degree in Computer Science, AI/ML, Data Science, or a related field. Equivalent practical experience will also be considered.
Strong hands-on experience with open models such as Llama, Qwen, Mistral, DeepSeek, GLM, Gemma, or similar models.
Experience running, benchmarking, fine-tuning, optimizing, or deploying LLMs on GPUs.
Strong understanding of LLMs, Transformers, RAG, fine-tuning, model evaluation, and inference. Also Python programming skills and experience with PyTorch and/or JAX.
Experience with Hugging Face Transformers and LLM inference frameworks such as vLLM or TensorRT-LLM.
Understanding of NVIDIA GPUs, GPU memory, quantization, model parallelism, and inference optimization.
Working knowledge of Linux, Docker, Kubernetes, and private/on-premise AI deployments.
Understanding of enterprise security, data privacy, access control, and AI workload isolation.
Strong experimental and analytical mindset with the ability to design benchmarks, analyze results, and derive meaningful insights.
Preferred Experience
Experience in LLM benchmarking, fine-tuning, post-training, or inference optimization.
Experience building RAG systems, AI agents, enterprise copilots, or domain-specific AI applications.
Experience working directly with enterprise IT, infrastructure, security, or application teams.
Contributions to open-source AI projects, research papers, technical blogs, or AI communities are a plus.
Candidates with experience building, deploying, benchmarking, fine-tuning, or optimizing real-world applications using Open Models will be strongly preferred.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.