Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
• Design, build, and maintain robust cloud data pipelines using Azure Data Factory (ADF), Azure Databricks, PySpark, and Spark SQL.
• Implement and manage the Medallion Architecture — moving and transforming data through Bronze (raw/audit), Silver (cleansing/dedup/SCD), and Gold (business aggregates) layers.
• Perform complex data transformations, cleansing, deduplication, and incremental loads using Delta MERGE, supporting both Slowly Changing Dimensions (SCD Type 1 and Type 2).
• Optimize Spark workloads by tuning shuffle partitions, managing memory to prevent OOM errors, and leveraging Adaptive Query Execution (AQE).
• Apply optimized join strategies, including broadcast joins for small datasets and salting techniques to handle data skew.
• Implement robust exception handling, file dependency validation, and asynchronous batch processing to ensure pipeline reliability.
• Ensure high data quality through schema enforcement, schema evolution handling, and validation against expected target criteria.
• Monitor, troubleshoot, and resolve production job failures by analysing cluster scaling behaviour, Spark UI metrics, and physical query plans.
• Build and maintain CI/CD pipelines for Databricks using Git, Azure DevOps, and Databricks Asset Bundles (DABs), with environment-specific parameterization for Dev, Test, and Prod.
Work arrangement
Hybrid
More openings worth a look
Recently tracked roles with full details and direct application links.