Live opening · Posted 19 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design, develop, and deploy scalable ETL/ELT pipelines using PySpark on GCP.
Utilise GCP services extensively, including BigQuery (data warehousing), Cloud Storage, Dataproc and Dataflow.
Optimise PySpark jobs for performance and reliability, and fine-tune BigQuery queries.
Implement complex transformations and process large volumes of structured/unstructured data using Spark SQL and PySpark.
Build and manage automated workflows using Apache Airflow or Cloud Composer.
Requirements:
Have implemented and architected solutions on Google Cloud Platform using the components of GCP.
Experience with Apache Beam/Google Dataflow/Apache Spark in creating end-to-end data pipelines.
Experience in some of the following: Big Query, Big Table Cloud Storage, Datastore, Spanner, Cloud SQL, and Machine Learning.
Experience programming in Hadoop, PySpark/Python, SQL.
Expertise in at least two of these technologies: Relational Databases, Analytical Databases, and NoSQL databases.
Being certified as a Google Professional Data Engineer/Solution Architect is a major advantage.
Mandatory Skills: Python, Pyspark, SQL, ETL, No SQL, GCP, BigQuery, Pub/Sub, Cloud Storage, Dataproc and Dataflow.
Bachelor's or Master's degree in Computer Science, Engineering or a related field.
Strong proficiency in Python and SQL is essential.
Extensive hands-on experience with PySpark.
Proven experience with Google Cloud Platform (GCP) services.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.