Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the job
We are seeking a Data Engineer to design, build, and optimize modern big data pipelines and streaming solutions. This role will focus on real-time data processing, cloud-based data platforms, and scalable data architectures while partnering with data scientists and business stakeholders to support enterprise data initiatives.
Responsibilities
Design and implement high-volume, real-time data streaming solutions using Apache Kafka and Spark Streaming
Develop and optimize data pipelines using Spark Structured Streaming, PySpark, and Scala
Troubleshoot, tune, and enhance Spark applications for performance and scalability
Build, test, and maintain big data ingestion pipelines and datasets
Deploy and support data platforms in AWS and Azure environments
Leverage serverless technologies such as S3, Kinesis/MSK, Lambda, and Glue
Manage Databricks environments, notebooks, Delta Lake, Delta Live Tables, and Unity Catalog
Work with messaging platforms including Kafka, Amazon MSK, TIBCO EMS, and IBM MQ
Ingest and process structured and semi-structured data from JSON, XML, and CSV sources
Work with NoSQL databases and modern data storage technologies
Develop and maintain shell scripts and data processing workflows on Unix/Linux platforms
Collaborate with data scientists, engineers, and stakeholders to meet business data requirements
Required Skills
Hands-on experience with Apache Kafka and Spark Streaming
Strong experience with Spark Structured Streaming and real-time data processing
Expertise troubleshooting and optimizing Spark applications
Strong programming experience with Python and/or Scala (PySpark/Scala-Spark)
Hands-on experience with Databricks
Experience building, testing, and optimizing big data ingestion pipelines and architectures
Experience deploying and supporting data platforms on AWS and/or Azure
Experience with cloud-native and serverless technologies such as S3, Kinesis/MSK, Lambda, and Glue
Strong knowledge of messaging platforms including Kafka, Amazon MSK, TIBCO EMS, or IBM MQ
Experience managing Databricks Notebooks, Delta Lake, Delta Live Tables, and Unity Catalog
Experience processing data from JSON, XML, and CSV formats
Experience with NoSQL databases such as HBase and Cassandra
Strong Unix/Linux and shell scripting experience
Experience with data platforms including Kudu, Impala, or Delta Lake
Preferred Skills
Experience building enterprise-scale streaming and event-driven architectures
Experience supporting cloud-based analytics and data lake solutions
Knowledge of modern data governance and data cataloging practices
Experience working closely with Data Science and Analytics teams
Exposure to large-scale distributed data processing environments
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.