Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design, build, and optimize scalable batch and streaming data pipelines using distributed frameworks like Spark/Databricks/Kafka.
Data and ML Engineering thought leadership (What, Why, and How).
Design and Code - robust data models, feature pipelines, and ETL/ELT frameworks for analytics and ML.
Ensure data quality, observability, lineage, and performance across data platforms.
Build and refine ML models end to end: feature engineering, training, evaluation, and deployment.
Partner with data scientists to convert prototypes into production-grade ML solutions.
Implement CI/CD, model versioning, monitoring, and automation across data and ML workflows.
Product-Driven Mindset: Collaborate with engineering and product teams to deliver data-driven outcomes.
Requirements:
7+ years of experience in ML-Data Engineering development. Strong SQ/NoSQL, Python, PySpark, and ML Models Lifecycle and Frameworks (MlFlow, Spark-ml), and Orchestration (Airflow/Oozie/Dagster, etc. ).
Expertise in Big Data modeling, Distributed processing, and Lake and Warehouse architectures at large operational scale.
Hands-on with ML lifecycle tools (MLflow, Feature Store, model monitoring, Evaluation).
Strong Analytical and Problem-Solving Skills - Data/Process-Intensive Design/Architecture.
Strong debugging and optimization. Basic hold on foundational modeling concepts and algorithms such as - Regression, Classification and Statistical models.
Good Hold on Concepts- Distributed File Formats, Open table Formats, Distributed transaction management, and Workload Parallelizing.
Hands-on: Unix, Hadoop, Object store fundamental operations, and commands.
Basic skills with containerized processing (Docker + K8s).
Experience in data engineering (SQL, Python, Databricks).
Strong understanding of data modeling and data structures.
Basic to intermediate knowledge of semantic technologies (RDF, OWL, knowledge graphs).
Experience with data pipelines and analytics/reporting (e. g., Power BI).
Understanding of data integration and interoperability concepts.
Experience
7-10 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.