Live opening · Posted 7 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Requirements:
4+ years in a data engineering role with end-to-end pipeline ownership.
Strong Python async patterns, subprocess management, API clients, and data processing at scale.
Hands-on AWS Athena, Glue, S3, DynamoDB, and Lake Formation; production-grade, not just familiarity.
Apache Hudi or Delta Lake schema evolution, partition strategies, and upsert patterns.
SQL proficiency: able to write and optimize complex analytical queries.
Experience with Airflow or an equivalent workflow orchestrator.
Demonstrated experience with large-scale media data pipelines, audio/video format conversion, metadata extraction, and chunking, e. g., egocentric or multimodal datasets (hard requirement, not a plus).
Understanding of how data format and I/O design affect downstream GPU compute workloads (WebDataset, sharded Parquet, tfrecord, or equivalent).
Nice to Have:
Direct experience designing data handoff for GPU clusters (AWS Batch GPU instances, Ray, or SLURM).
Familiarity with ML training data formats and dataset standards used by AI labs (Hugging Face datasets, WebDataset, dataset cards).
Experience with rclone, large-scale file transfer, or cloud-to-cloud sync pipelines.
Exposure to data lineage or provenance tooling (OpenLineage, DataHub, or custom metadata schemas).
Skills
AWS, Delta Lake, Airflow, Data Engineering, Data Analysis, ETL, NoSQL, SQL, Amazon DynamoDB
Experience
4-8 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.