Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We're looking for a Senior Data Engineer with strong data architecture and system design skills to design, build, and own end-to-end production pipelines powering analytics and decision-making across the org. The role centres on deep Python + SQL expertise, robust transaction-level fact tables, and rigorous data quality and auditability, ideally within a finance, risk, or compliance data domain. Internal candidates familiar with Uber's data platform stack (Piper/uWorc, Databook, uSecret, DSW, SourceGraph, Query Builder, OneETL) will be prioritised.
Responsibilities:
Own data architecture and system design decisions for pipelines and data models - grain, partitioning, schema evolution, and scalability trade-offs.
Take end-to-end ownership of pipelines in production: design, build, deploy, monitor, and operate.
Design and build transaction tables (append-only, immutable event/transaction-grain fact tables) alongside standard fact/dimension models using medallion (bronze/silver/gold) architecture.
Build and support pipelines for finance, risk, or compliance use cases where accuracy, auditability, and data lineage are critical.
Implement data quality and auditability controls: validation checks, reconciliation logic, anomaly detection, and audit trails for every pipeline you own.
Write complex, cross-dialect SQL (MySQL + PostgreSQL): window functions, multi-layered CTEs, CASE-driven logic, COALESCE/NULLIF, type casting, and date/timestamp handling.
Build idempotent reload patterns (DELETE+INSERT), UNION ALL/set operations, and templated (Jinja-style) SQL including handling for late-arriving/corrected transaction records without double-counting.
Orchestrate Airflow-style DAGs (Pipeline/BaseTask, ExternalTaskSensor for cross-pipeline deps) with secure credential handling via uSecret.
Build Python ETL tooling: pandas (CSV DB), SQLAlchemy + raw drivers (MySQLdb, psycopg2), type hints/dataclasses, class-based pipeline design.
Integrate with Google Drive/Sheets APIs; build YAML-driven pipeline configs; handle Piper staging/user_staging quirks.
Own CI/CD for pipelines via standard Git workflows.
Experience
5-8 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.