Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We're in search of a Data Engineer who brings a passion for innovation and problem-solving, not limited to traditional data engineering. This role is for those with deep expertise in building large-scale data pipelines, SQL optimisation, and cloud infrastructure, excited to design and build the backbone of our AI-powered data solutions. The ideal candidate will play a crucial role in enhancing our AI capabilities, developing high-performance data systems, and integrating advanced AI using the product into thousands of data teams' daily workflows.
Responsibilities:
Building PB Scale Data Pipelines: Building highly performant large-scale data infrastructure that can scale to 100K+ jobs and handle PB-scale data per day.
Cloud-Native Data Infrastructure: Design and implement robust, scalable data infrastructure on AWS, utilising Kubernetes and Airflow for efficient resource management and deployment.
Intelligent SQL Ecosystem: Design and develop a comprehensive SQL intelligence system encompassing query optimisation, dynamic pipeline generation, and data lineage tracking. Leverage your expertise in SQL query profiling, AST analysis, and parsing to create a sophisticated engine focused on query performance improvements, building adaptive data pipelines, and implementing granular column-level lineage.
Innovation at the Forefront: Push the boundaries of data engineering by combining traditional techniques with cutting-edge AI technologies.
High Visibility: Directly affect the productivity and capabilities of global data teams, as your contributions will be crucial to the daily operations of thousands of users spread across 100s of countries.
Open Source Contribution: As part of our commitment to the developer community, you will contribute to our open-source initiatives, gaining recognition in the tech community.
Career Growth: This role is a launchpad into the rapidly advancing field of AI-powered data engineering, offering exposure to state-of-the-art technologies and generative AI applications.
Requirements:
2-4 years of experience in data engineering, with a focus on building scalable data pipelines and systems.
Strong proficiency in Python and SQL.
Extensive experience with SQL query profiling, optimisation, and performance tuning, preferably with Snowflake.
Deep understanding of SQL Abstract Syntax Tree (AST) and experience working with SQL parsers (e. g., sqlglot) for generating column-level lineage and dynamic ETLs.
Experience in building data pipelines using Airflow or dbt.
[Optional] Solid understanding of cloud platforms, particularly AWS.
[Optional] Familiarity with Kubernetes (K8s) for containerised deployments.
Experience
2-4 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.