Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Build and maintain production services and APIs that serve analytics and ML model outputs at scale across multi-country GCC deployments.
Design and implement ETL/ELT pipelines feeding analytical and ML workloads including supply chain demand forecasting, claims adjudication signals, and population health indicators.
Develop predictive and statistical models for health supply chain commodity flows, eClaims payer analytics, and population health use cases.
Collaborate with product and domain teams to translate clinical and operational questions into well-scoped data problems with measurable success criteria.
Own model monitoring, retraining pipelines, and data quality frameworks, not just initial deployment.
Contribute to and extend the data lakehouse architecture built on Apache Iceberg, ClickHouse, and Parquet with Dagster orchestrating the pipeline.
Participate in architecture decisions around the analytics data stack, including ClickHouse DR, lakeFS as Iceberg REST catalogue, and dbt-clickhouse transformations.
Write clean, testable, production-grade code, not research notebooks passed to an engineering team.
Requirements:
6-9 years of total experience with a demonstrable split between data science and backend or data engineering, not purely one or the other.
Strong Python (non-negotiable): Pandas, NumPy, scikit-learn, SQLAlchemy, and FastAPI or equivalent async frameworks.
Solid SQL and hands-on experience with columnar or analytical databases; ClickHouse strongly preferred; BigQuery, Snowflake, or DuckDB accepted.
Production model deployment experience: you have shipped models to live environments and maintained them, not just trained and handed off.
Familiarity with Kafka or event-driven architectures is a meaningful plus.
Experience with dbt or equivalent transformation tooling and an understanding of medallion or lakehouse architecture patterns.
Experience in health data - FHIR, claims, supply chain, or logistics is a strong differentiator but not a hard requirement.
Comfort working in a B2G context where data sovereignty, auditability, and schema compliance are first-class concerns
Nice to Have:
Exposure to GS1 or EPCIS data formats for pharmaceutical or medical device traceability.
Experience with Dagster, Prefect, or Airflow for pipeline orchestration.
Familiarity with lakeFS, Apache Iceberg, or OpenTable format ecosystems.
Experience
6-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.