Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Description: Senior Databricks Data Engineer
Role Overview
We are looking for an experienced Senior Databricks Data Engineer to design, develop, and maintain scalable ETL/ELT data pipelines using Databricks, PySpark, Spark SQL, and Delta Lake. The candidate should have strong hands-on experience in data transformation, pipeline orchestration, performance optimization, data quality, and production support within an AWS-based data engineering environment.
The role also requires practical experience using Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools to accelerate development, testing, debugging, documentation, and code optimization.
Key Responsibilities
Design, develop, test, and maintain scalable ETL/ELT data pipelines using Databricks and PySpark.
Develop and optimize data transformation logic using PySpark, Spark SQL, and SQL.
Build and maintain Delta Lake tables following Bronze, Silver, and Gold layers of the Medallion Architecture.
Implement Delta Lake capabilities, including:
MERGE and upsert operations
Schema enforcement and schema evolution
Time Travel and version management
OPTIMIZE and VACUUM
Incremental and change-based data processing
Develop and manage Databricks Workflows and Jobs for pipeline orchestration, scheduling, dependency management, retries, and alerting.
Design efficient incremental-load and restart/recovery mechanisms to ensure reliable pipeline execution.
Troubleshoot pipeline failures, performance bottlenecks, data-quality issues, and production incidents.
Optimize Spark workloads through effective use of partitioning, joins, caching, broadcast strategies, shuffle management, file sizing, and query execution plans.
Develop reusable data engineering utilities, libraries, and frameworks using PySpark and Python.
Write complex SQL queries for data transformation, reconciliation, validation, analysis, and reporting.
Integrate Databricks pipelines with external applications and systems using REST APIs.
Develop and support inbound and outbound file integrations using SFTP.
Process structured and semi-structured data in formats such as CSV, JSON, Parquet, and Delta.
Work with AWS services used in the data engineering ecosystem, particularly Amazon S3.
Implement exception handling, logging, auditing, monitoring, alerting, and operational recovery mechanisms.
Implement data-quality validations and reconciliation controls across different pipeline stages.
Participate in code reviews and ensure adherence to coding, security, performance, and data engineering best practices.
Collaborate with architects, business analysts, source-system teams, QA teams, DevOps teams, and other engineering stakeholders.
Support production deployments, release validation, operational monitoring, and troubleshooting of production data pipelines.
Create and maintain technical documentation, pipeline specifications, operational procedures, and support runbooks.
Use Cursor or similar agentic AI-enabled IDEs to improve development productivity, including:
Generating and refactoring PySpark, Python, and SQL code
Creating unit tests and data-validation scripts
Troubleshooting errors and performance issues
Generating technical documentation and code explanations
Reviewing and validating AI-generated code for correctness, security, maintainability, and performance
Must-Have Skills
Strong hands-on experience in Databricks-based data engineering.
Advanced proficiency in PySpark, Spark SQL, and SQL.
Strong understanding of Apache Spark architecture and distributed data processing.
Hands-on experience building ETL/ELT pipelines at enterprise scale.
Strong experience with Delta Lake and Medallion Architecture.
Practical experience with Delta Lake features such as MERGE, schema evolution, Time Travel, OPTIMIZE, and VACUUM.
Hands-on experience developing and managing Databricks Workflows and Jobs.
Strong understanding of Spark performance optimization, including:
Data partitioning
Join strategies
Caching and persistence
Shuffle optimization
Data skew handling
Query execution plans
Small-file management
Strong SQL skills, including complex joins, window functions, aggregations, reconciliation, and data-validation queries.
Working knowledge of Python for developing reusable utilities, frameworks, and automation scripts.
Experience working with CSV, JSON, Parquet, and Delta file formats.
Hands-on experience integrating data pipelines with REST APIs and SFTP systems.
Experience working with Amazon S3 in a data engineering environment.
Experience implementing logging, exception handling, auditing, monitoring, alerting, and restart/recovery mechanisms.
Experience troubleshooting and supporting production data pipelines.
Hands-on knowledge of Cursor, GitHub Copilot, Windsurf, or another agentic AI-enabled development environment.
Ability to write effective prompts, review AI-generated code, identify hallucinated or incorrect logic, and apply organizational security and data-privacy standards while using AI tools.
Strong problem-solving, debugging, communication, and stakeholder-collaboration skills.
Preferred Skills
Knowledge of Databricks Unity Catalog, access controls, lineage, and data governance.
Experience with Databricks Asset Bundles, Databricks CLI, or CI/CD-based deployment approaches.
Basic understanding of AWS IAM, Secrets Manager, CloudWatch, and Lambda.
Experience with cloud security, secrets management, and credential-handling best practices.
Knowledge of automated testing frameworks for PySpark and data pipelines.
Familiarity with Git-based development, branching strategies, pull requests, and CI/CD pipelines.
Experience working in Agile/Scrum delivery environments.
Databricks or AWS certification would be an advantage.
Education And Experience
Bachelor?s degree in Computer Science, Information Technology, Engineering, or a related discipline.
Relevant professional experience in data engineering, including substantial hands-on experience with Databricks, PySpark, SQL, and Delta Lake.
Experience delivering and supporting enterprise-scale data platforms in production environments.
Databricks, SQL
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.