Live opening · Posted 11 days ago

Data Scientist

1Buy.AI · Delhi
Instahyre 7-11 yrs
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
Company1Buy.AI
LocationDelhi
Experience7-11 yrs
SourceInstahyre
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
1 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
17,046 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Requirements:
You ship code daily. Your GitHub is active. Python is your primary language. You debug model traces, read stack traces, and review your own PRs; you do not wait for an engineering team to implement your ideas.
6-12 years of hands-on experience in data science, ML engineering, or software engineering with a clear arc toward production AI/ML systems.
At least 3 years owning production ML systems, not research prototypes. Models that received real traffic, failed in real ways, and were debugged, retrained, and monitored over time.
Strong classical ML proficiency: XGBoost, LightGBM, and CatBoost for tabular tasks, time-series forecasting, clustering and unsupervised methods, and SHAP-based interpretability. You know when gradient boosting beats a neural network and have the benchmark to prove it.
Feature engineering depth: encoding strategies, temporal and lag features, interaction terms, domain-specific transforms, and leakage-free cross-validation design. This is still more impactful than model choice on most business problems.
Practical GenAI engineering: you have built RAG pipelines in production. You understand chunking strategy trade-offs, embedding model selection, retrieval quality metrics, and reranking layer design.
Vector database fluency: hands-on experience with Qdrant, Weaviate, ChromaDB, or pgvector index types (HNSW, IVF), metadata filtering, and hybrid dense-sparse search.
LLM orchestration experience: LangChain or LlamaIndex at production scale, chains, retrieval pipelines, and version-controlled prompt management.
Prompt engineering as a discipline: version-controlled, systematically evaluated prompts. Chain-of-thought, few-shot, and XML-tagged system prompts. Not ad-hoc iteration in a notebook.
MLOps hygiene: MLflow or W& B for experiment tracking, Docker for containerisation, AWS basics (SageMaker, S3 Lambda), and model monitoring in production.
HuggingFace ecosystem: model hub, transformers library, datasets, and sentence transformers. You can benchmark and select an embedding model for a specific retrieval task, not just call from_pretrained.
Preferred:
Hands-on experience with MCP (Model Context Protocol), building or consuming MCP-compatible tool endpoints for AI agent integration.
Fine-tuning experience: LoRA, QLoRA, PEFT on open-source LLMs (Llama, Mistral, Qwen) for domain adaptation or task-specific performance improvement.
LLM evaluation frameworks: RAGAS, LLM-as-judge patterns, or custom eval harness design for RAG pipeline quality measurement.
Experience contributing to agentic workflows, building tool schemas, function signatures, or supporting multi-agent state machine design in LangGraph or equivalents.
Exposure to supply chain, procurement, logistics, or industrial B2B data, pricing models, demand forecasting, or lifecycle risk in physical goods contexts.
Familiarity with Ollama or vLLM for local inference and cost benchmarking against cloud API alternatives.
Good to Have:
Advanced degree (M. S. or Ph. D. ) in computer science, statistics, mathematics, or a related quantitative field, or equivalent depth demonstrated through shipped products or open-source contributions.
Experience with multimodal AI, parsing PDFs, images (datasheets, BOMs), or combined structured/unstructured inputs for information extraction pipelines.
Open-source contributions to AI/ML libraries, tools, published evaluations, or fine-tuned model releases.
Familiarity with data governance, lineage tooling, or responsible AI frameworks at an engineering level (implemented, not just read about).
Experience with model quantisation (GGUF, AWQ, GPTQ) or edge inference for latency-sensitive production environments.

Experience
7-11 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App