Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Design, develop, deploy and maintain reliable ETL/ELT batch & streaming data pipelines, ingest structured/semi-structured/unstructured data from multiple source systems (DB, API, log, Kafka etc.)
Perform dimensional data modeling, build and iterate data warehouse / data lake, define table schema, partition strategy, data layers (ODS/DWD/DWS/ADS)51job
Write & optimize complex SQL, tune query performance, reduce resource cost and improve pipeline stability and SLA
Implement data quality rules, anomaly detection, monitoring & alerting, troubleshoot data lineage and data consistency issues
Use workflow orchestration tools to schedule, manage and monitor data jobs (Airflow/Dagster)
Collaborate with analysts, DS and business stakeholders to translate business requirements into data solutions, build reusable datasets and data APIs
Participate in data governance: metadata management, data security, access control, documentation for schema and pipelines51job
Evaluate and adopt cloud data technologies (Snowflake/BigQuery/Redshift), continuously optimize storage and compute cost
Experience on AWS/Azure/GCP cloud data stack
Real-time streaming data pipeline development (Kafka, Flink, Pulsar)
Data governance, data lineage, metadata platform construction experience
Experience supporting ML feature platform / feature engineering
English working proficiency
Domain knowledge: automotive, manufacturing, finance, retail etc.
Employment type
Full-time
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.