Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design and operate batch and streaming data pipelines (ETL/ELT) at production scale.
Build and maintain training-ready datasets for ML/AI feature pipelines, labeling workflows, and data quality for model training.
Own the data warehouse/lakehouse stack (e. g., Snowflake, BigQuery, Databricks, Spark) and transformations (dbt).
Orchestrate workflows (Airflow or equivalent) with strong monitoring and lineage.
Implement data security and compliance: PII classification, encryption, access controls, retention, and governance (GDPR and similar).
Partner with AI/data science teams to make data discoverable, trustworthy, and fast.
Requirements:
5-7 years in data engineering, ideally within a SaaS environment.
Hands-on experience preparing data for model/ML training (not just BI/reporting pipelines).
Strong SQL and Python; deep experience with a modern warehouse/lakehouse and Spark.
Comfortable integrating data services with a Node.js / TypeScript platform (APIs, events, schemas).
Pipeline orchestration (Airflow) and transformation (dbt) at scale.
Demonstrated work under data security and compliance requirements.
Strong data modeling and data quality fundamentals.
Bonus:
Streaming (Kafka, Flink, Kinesis) and real-time feature stores.
MLOps / feature store tooling (Feast, Tecton).
Experience in a multi-tenant or regulated data environment.
Non-negotiables: 5-7 years of SaaS data engineering hands-on ML/training data prep, data security, and compliance experience.
Experience
5-7 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.