Live opening · Posted 1 day ago

AI engineering intern - custom models & multi-model memory

Recollia · India (Remote)
Linkedin Yes
You are 1 day behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 1 day ago
CompanyRecollia
LocationIndia (Remote)
Work modeYes
SourceLinkedin
Listed1 day ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
27 min from Linkedin publishing this role to us finding it
15 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
62,674 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

This is a founding engineer offer with a stipend of Rs 20,000 with a possibility for an incentive of Rs 5,000 every bi-weekly depending on the performance. Internship will be converted into founding engineer role which comes with equity.
Recollia builds AI tools for medical students. Our responses run on Claude and OpenAI today. We're replacing most of that with our own fine-tuned models, and building RAV, a unified-memory layer that works across Claude, ChatGPT, open-weight models, and our internal tools. You'll work on both, inside the architecture, not on the edges of it.
WHAT YOU'LL BUILD
• Fine-tuning pipelines: data prep, SFT and LoRA/QLoRA on open-weight models, evals that prove the model got better
• Inference serving, measured on cost per response against the API baseline
• RAV's memory layer and a model gateway that routes across providers with fallbacks
REQUIRED
Model training
• Transformer internals well enough to reason about them: attention, KV cache, tokenization, chat templates, context-length trade-offs
• Fine-tuning beyond the tutorial: SFT, LoRA/QLoRA, preference tuning (DPO or similar), and when each is the wrong tool
• Training mechanics: mixed precision, gradient accumulation and checkpointing, LR schedules, multi-GPU (FSDP/DeepSpeed), and the memory math to know what fits on a given card
• Dataset work: cleaning, dedup, formatting, synthetic data generation, and recognizing when the data is the problem rather than the model
• Evaluation: held-out eval sets, LLM-as-judge and its failure modes, regression suites; a loss curve alone doesn't convince you
Inference and serving
• Serving stacks (vLLM, SGLang, TGI, llama.cpp): continuous batching, KV-cache management, what actually drives throughput
• Quantization (AWQ, GPTQ, GGUF, FP8) and measuring the quality you lose
• Profiling: TTFT, tokens/sec, p95 latency, cost per response against an API baseline
• GPU cloud deploys (RunPod, Modal, Lambda, AWS) with Docker; autoscaling and cold starts
Memory and retrieval
• Embedding models: selection, evaluation, domain fine-tuning
• Retrieval pipelines: chunking, hybrid search (BM25 + dense), rerankers, metadata filtering, query rewriting; measured with recall@k, not vibes
• Memory design: episodic vs. semantic memory, compaction, write policies, staleness and conflict resolution, entity or graph memory
• Vector stores (pgvector, Qdrant, LanceDB) and Postgres well enough to design schemas and indexes
Multi-model orchestration
• Provider-agnostic gateway across Anthropic, OpenAI, and OpenAI-compatible open-weight endpoints: routing by cost, quality, and latency; fallbacks, retries, rate limits, streaming
• Structured outputs and tool calling across providers, including where they differ
• Agent loops, MCP servers and clients, connector ingestion via OAuth APIs (Gmail, Calendar, Drive, GitHub, Linear, Notion)
• Observability: tracing (Langfuse, OpenTelemetry), token accounting, cost dashboards
• Prompt injection and PII risk in connector-fed systems, and how to contain it
Engineering foundations
• Python at library level, not script level: async, typing, packaging, profiling; you've read transformers or vLLM source when the docs ran out
• FastAPI or equivalent, Postgres, Redis, Docker, Linux, GPU driver and CUDA troubleshooting
• Git workflow, CI, tests that catch regressions
• Enough math to read a loss curve and know what to change: optimization, LR schedules, overfitting on small data
• Reads papers and repos, reproduces results, keeps experiment logs (W&B or equivalent)
No degree required. Skills over credentials.
WHAT HUNGER LOOKS LIKE TO US
• You ship without being told. Your GitHub shows it.
• You'd rather run the experiment than debate it.
• You ask for more scope, not more instructions.
• You say when something is broken, including our decisions.
HOW TO APPLY
Send email to careers@recollia.ai with one thing you built: repo, demo, or write-up. In three sentences, tell us what it does, what broke, and what you'd do differently.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App