Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This is a founding engineer offer with a stipend of Rs 20,000 with a possibility for an incentive of Rs 5,000 every bi-weekly depending on the performance. Internship will be converted into founding engineer role which comes with equity.
Recollia builds AI tools for medical students. Our responses run on Claude and OpenAI today. We're replacing most of that with our own fine-tuned models, and building RAV, a unified-memory layer that works across Claude, ChatGPT, open-weight models, and our internal tools. You'll work on both, inside the architecture, not on the edges of it.
WHAT YOU'LL BUILD
• Fine-tuning pipelines: data prep, SFT and LoRA/QLoRA on open-weight models, evals that prove the model got better
• Inference serving, measured on cost per response against the API baseline
• RAV's memory layer and a model gateway that routes across providers with fallbacks
REQUIRED
Model training
• Transformer internals well enough to reason about them: attention, KV cache, tokenization, chat templates, context-length trade-offs
• Fine-tuning beyond the tutorial: SFT, LoRA/QLoRA, preference tuning (DPO or similar), and when each is the wrong tool
• Training mechanics: mixed precision, gradient accumulation and checkpointing, LR schedules, multi-GPU (FSDP/DeepSpeed), and the memory math to know what fits on a given card
• Dataset work: cleaning, dedup, formatting, synthetic data generation, and recognizing when the data is the problem rather than the model
• Evaluation: held-out eval sets, LLM-as-judge and its failure modes, regression suites; a loss curve alone doesn't convince you
Inference and serving
• Serving stacks (vLLM, SGLang, TGI, llama.cpp): continuous batching, KV-cache management, what actually drives throughput
• Quantization (AWQ, GPTQ, GGUF, FP8) and measuring the quality you lose
• Profiling: TTFT, tokens/sec, p95 latency, cost per response against an API baseline
• GPU cloud deploys (RunPod, Modal, Lambda, AWS) with Docker; autoscaling and cold starts
Memory and retrieval
• Embedding models: selection, evaluation, domain fine-tuning
• Retrieval pipelines: chunking, hybrid search (BM25 + dense), rerankers, metadata filtering, query rewriting; measured with recall@k, not vibes
• Memory design: episodic vs. semantic memory, compaction, write policies, staleness and conflict resolution, entity or graph memory
• Vector stores (pgvector, Qdrant, LanceDB) and Postgres well enough to design schemas and indexes
Multi-model orchestration
• Provider-agnostic gateway across Anthropic, OpenAI, and OpenAI-compatible open-weight endpoints: routing by cost, quality, and latency; fallbacks, retries, rate limits, streaming
• Structured outputs and tool calling across providers, including where they differ
• Agent loops, MCP servers and clients, connector ingestion via OAuth APIs (Gmail, Calendar, Drive, GitHub, Linear, Notion)
• Observability: tracing (Langfuse, OpenTelemetry), token accounting, cost dashboards
• Prompt injection and PII risk in connector-fed systems, and how to contain it
Engineering foundations
• Python at library level, not script level: async, typing, packaging, profiling; you've read transformers or vLLM source when the docs ran out
• FastAPI or equivalent, Postgres, Redis, Docker, Linux, GPU driver and CUDA troubleshooting
• Git workflow, CI, tests that catch regressions
• Enough math to read a loss curve and know what to change: optimization, LR schedules, overfitting on small data
• Reads papers and repos, reproduces results, keeps experiment logs (W&B or equivalent)
No degree required. Skills over credentials.
WHAT HUNGER LOOKS LIKE TO US
• You ship without being told. Your GitHub shows it.
• You'd rather run the experiment than debate it.
• You ask for more scope, not more instructions.
• You say when something is broken, including our decisions.
HOW TO APPLY
Send email to careers@recollia.ai with one thing you built: repo, demo, or write-up. In three sentences, tell us what it does, what broke, and what you'd do differently.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.