Live opening · Posted 1 day ago

Senior Data Engineer

Apeiro · Bangalore | Noida
Instahyre 6-9 yrs
You are 1 day behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 1 day ago
CompanyApeiro
LocationBangalore | Noida
Experience6-9 yrs
SourceInstahyre
Listed1 day ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
25 min from Instahyre publishing this role to us finding it
12 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,956 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
Build and maintain production services and APIs that serve analytics and ML model outputs at scale across multi-country GCC deployments.
Design and implement ETL/ELT pipelines feeding analytical and ML workloads including supply chain demand forecasting, claims adjudication signals, and population health indicators.
Develop predictive and statistical models for health supply chain commodity flows, eClaims payer analytics, and population health use cases.
Collaborate with product and domain teams to translate clinical and operational questions into well-scoped data problems with measurable success criteria.
Own model monitoring, retraining pipelines, and data quality frameworks, not just initial deployment.
Contribute to and extend the data lakehouse architecture built on Apache Iceberg, ClickHouse, and Parquet with Dagster orchestrating the pipeline.
Participate in architecture decisions around the analytics data stack, including ClickHouse DR, lakeFS as Iceberg REST catalogue, and dbt-clickhouse transformations.
Write clean, testable, production-grade code, not research notebooks passed to an engineering team.
Requirements:
6-9 years of total experience with a demonstrable split between data science and backend or data engineering, not purely one or the other.
Strong Python (non-negotiable): Pandas, NumPy, scikit-learn, SQLAlchemy, and FastAPI or equivalent async frameworks.
Solid SQL and hands-on experience with columnar or analytical databases; ClickHouse strongly preferred; BigQuery, Snowflake, or DuckDB accepted.
Production model deployment experience: you have shipped models to live environments and maintained them, not just trained and handed off.
Familiarity with Kafka or event-driven architectures is a meaningful plus.
Experience with dbt or equivalent transformation tooling and an understanding of medallion or lakehouse architecture patterns.
Experience in health data - FHIR, claims, supply chain, or logistics is a strong differentiator but not a hard requirement.
Comfort working in a B2G context where data sovereignty, auditability, and schema compliance are first-class concerns
Nice to Have:
Exposure to GS1 or EPCIS data formats for pharmaceutical or medical device traceability.
Experience with Dagster, Prefect, or Airflow for pipeline orchestration.
Familiarity with lakeFS, Apache Iceberg, or OpenTable format ecosystems.

Experience
6-9 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App