Live opening · Posted 2 days ago

Senior Data Scientist (HMIS, Analytics)

Apeiro · Bangalore | Noida
Instahyre 5-9 yrs
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyApeiro
LocationBangalore | Noida
Experience5-9 yrs
SourceInstahyre
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
25 min from Instahyre publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,099 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are looking for a Senior Data Scientist who can take a loosely defined problem why are stockouts rising in this region or can we predict claim denials before submission and drive it end to end: framing the question, sourcing and validating the data, building the model, and translating the result into something a product team or health authority can act on. This is a role for someone comfortable owning ambiguity, not waiting for a fully specified ticket. You will help build the NLP, generative AI, and RAG capabilities that let users query complex health databases in plain language safely, accurately, and grounded in trusted data.
Responsibilities:
Own data science problems end-to-end from ambiguous business or clinical questions to scoped problems, models, and measurable outcomes.
Develop predictive and statistical models across health domains: supply chain demand forecasting, claims and denial analytics, and population health indicators.
Design experiments and validation approaches defining success metrics, baselines, and evaluation criteria before building.
Work with large-scale, multi-country datasets, handling the messiness, gaps, and inconsistencies of real-world health data.
Build and improve natural-language-to-data capabilities: LLM- and RAG-based systems that let users query health databases in plain language, including text-to-SQL, retrieval over schema and documentation, and grounding model outputs in trusted data.
Apply NLP and generative AI techniques to unstructured health data clinical notes, documents, and free-text fields where structured extraction adds value.
Partner with data engineers to productionise model feature pipelines, retraining, monitoring, and drift detection rather than handing off notebooks.
Collaborate with product and domain teams to translate clinical and operational questions into well-scoped data problems with clear success criteria.
Communicate findings and their limitations honestly to technical and non-technical stakeholders, including government counterparts.
Set analytical standards for the team's methodology, rigour, reproducibility, and documentation and mentor junior analysts and data scientists.
Contribute to architecture and tooling decisions for the analytics and ML stack.
Requirements:
5+ years of experience in data science, with a track record of shipping models and analyses that reached real users or decisions.
Demonstrated experience on large-scale projects working with high-volume data and systems where scale and performance matter.
Proven ability to handle ambiguity: you can take an open-ended problem, structure it, and drive it to a defensible answer with limited direction.
Strong Python for data science: Pandas, NumPy, scikit-learn, and at least one deep learning or advanced modelling framework where relevant.
Strong SQL and hands-on experience with large or analytical datasets; comfort with columnar databases such as ClickHouse, BigQuery, Snowflake, or DuckDB is a plus.
Solid grounding in statistics and machine learning fundamentals: you can choose the right method and explain why, and you know the difference between correlation and a usable signal.
Hands-on experience with Natural Language Processing (NLP) text processing, embeddings, and language model fundamentals.
Practical experience with generative AI and large language models (LLMs) prompting, evaluation, and building applications on top of foundation models via APIs or open-weight models.
Experience designing and building Retrieval-Augmented Generation (RAG) systems, chunking, embeddings, vector search, and retrieval quality and an understanding of how to ground LLM outputs to reduce hallucination.
Familiarity with natural-language-to-SQL / text-to-data approaches and their failure modes, especially the accuracy and safety concerns of letting an LLM query a real database.
Experience taking models to production and maintaining them, monitoring, retraining, and model quality over time, not just training and handing off.
Clear communication: you can defend a methodology to a technical peer and explain the same result plainly to a non-technical decision-maker.
Strong data quality and epistemic discipline: you surface uncertainty and caveats rather than overstating confidence.
Nice to Have:
Experience with health data FHIR, claims, supply chain, or logistics is a strong differentiator but not a hard requirement.
Experience with time-series forecasting, causal inference, or optimisation in an operational context.
Familiarity with LLM tooling and frameworks like LangChain, LlamaIndex, or equivalents and vector databases such as pgvector
Experience evaluating LLM and RAG systems, systematically building eval sets and measuring answer accuracy, faithfulness, and retrieval quality rather than relying on spot checks.
Familiarity with MLOps tooling experiment tracking, model registries, and pipeline orchestration (Dagster, Prefect, or Airflow).
Experience with dbt or equivalent transformation tooling and lakehouse or medallion architecture patterns.
Comfort working in a B2G context where data sovereignty, auditability, and reproducibility are first-class concerns.
Prior experience mentoring or setting technical direction for a small data team.

Experience
5-9 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App