Live opening · Posted 20 hours ago

Applied Machine Learning Engineer (Model Layer)

BruntWork · India (Remote)
Linkedin No
You are 20 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 20 hours ago
CompanyBruntWork
LocationIndia (Remote)
Work modeNo
SourceLinkedin
Listed20 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
45 min from Linkedin publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
17,430 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Our client is building a simple, private AI chat platform designed to give users easy access to powerful open-source and leading AI models without requiring any technical setup. The product focuses on delivering high-quality AI responses while keeping the experience fast, affordable, and privacy-conscious.
The front-end and website are managed by a separate application team. This role is focused specifically on the AI model layer—the engine that powers the chat experience.
The team is small and early-stage, so this is an opportunity for a hands-on engineer to take significant ownership of a critical part of the product and work directly with founders and engineering stakeholders.
About the Role
As the Senior LLM / AI Model Routing Engineer, you will own the model layer from integration through production optimization. You will build the systems that determine which AI model should handle each request, evaluate model performance, and continuously improve the balance between quality, speed, reliability, and cost. This is a production engineering role, not a research-only position. The ideal candidate has experience taking LLM technology beyond prototypes and building reliable systems used by real users.
Key Responsibilities
LLM Integration & Model Management
Integrate open-source and third-party LLMs through APIs and inference providers.
Maintain the product's model lineup and evaluate new models as they become available.
Compare models based on quality, reliability, speed, cost, and suitability for different use cases.
Determine when models should be added, replaced, or retired.
Implement reliable model and provider fallbacks.
Model Routing
Design and maintain intelligent model routing logic.
Route requests to the most appropriate model based on factors such as task type, quality, latency, cost, and model behavior.
Continuously refine routing decisions using production data and user feedback.
Explore advanced approaches such as model cascades, ensembles, and multi-model systems when appropriate.
Evaluation & Optimization
Build practical systems for evaluating LLM output quality and routing decisions.
Establish and monitor key performance metrics, including quality, latency, reliability, and cost per request.
Analyze production data to identify opportunities for improvement.
Optimize prompts, model parameters, configurations, and fallback strategies.
Balance response quality with performance and operating costs.
Performance & Reliability
Monitor model performance, latency, usage, and inference costs.
Identify and resolve issues with models, providers, routing, and integrations.
Improve response speed and reliability as usage grows.
Implement appropriate logging, monitoring, error handling, and fallback mechanisms.
Maintain clear documentation for the model architecture and operational processes.
Collaboration
Work closely with the application engineering team to integrate the model layer into the product.
Collaborate on the connection between the AI, application, and policy layers.
Communicate technical concepts and trade-offs clearly to founders and non-ML stakeholders.
Provide data-driven recommendations on model strategy and product performance.
Requirements
Proven experience building LLM-powered applications in production.
Strong Python development skills.
Hands-on experience working with LLM APIs and inference platforms.
Experience with providers such as OpenAI, Anthropic, Together, Groq, or similar.
Strong understanding of LLM capabilities, prompt design, model evaluation, and production behavior.
Ability to make practical trade-offs between quality, latency, reliability, and cost.
Strong analytical and problem-solving abilities.
Comfortable working independently in a small, fast-moving startup environment.
Strong written and verbal English communication skills.
Preferred Qualifications
Experience with LLM model routing, multi-model architectures, ensembles, or AI gateways.
Familiarity with open-weight models such as Llama, Qwen, or Gemma.
Experience with LLM evaluation frameworks or automated evaluation pipelines.
Experience optimizing inference cost, latency, throughput, or reliability.
Experience with model observability and production monitoring.
Experience building privacy-focused or privacy-conscious AI products.
Previous experience working in an early-stage or lean engineering team.
The Ideal Candidate
We are looking for a builder who enjoys ownership and execution.
You don't need to be an academic AI researcher. You need to understand how modern LLMs work, know how to evaluate them in the real world, and be able to turn that knowledge into reliable production systems.
You are:
Pragmatic: You make smart trade-offs instead of over-engineering.
Data-driven: You measure performance and use evidence to make decisions.
Curious: You stay current with new models, providers, and AI techniques.
Proactive: You identify problems and opportunities without waiting for instructions.
Ownership-oriented: You take responsibility for the performance and reliability of your systems.
Collaborative: You can work effectively with engineers, founders, and non-technical stakeholders.
What Success Looks Like
Requests are consistently routed to the best model for the task.
Users receive high-quality, accurate, and relevant responses.
Latency remains fast and consistent as usage increases.
Cost per request is continuously optimized.
Model and provider failures are handled reliably.
New models are evaluated and integrated quickly when they add value.
Routing and model decisions are based on measurable production data.
The model layer is reliable, well-documented, and continuously improving.
Why Join?
Own a critical AI layer: Take end-to-end ownership of the engine behind the product.
Work with leading LLMs: Continuously evaluate and work with new open-weight and provider models.
Make a direct impact: Your decisions will directly influence product quality, performance, and cost.
Work with a lean team: Collaborate closely with founders and have meaningful technical influence.
Remote & global: Work from anywhere with a flexible schedule.
Build for real users: Move beyond experimentation and build production AI systems.
Growth opportunity: Expand your technical ownership as the product and team scale.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App