Live opening · Posted 12 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Title: Cloud Data Engineer - Location - Bangalore - (Hybrid - 3Days a Week)
Looking for Immediate Joiners / Serving Notice Peroid.
Experience
5 7 years of relevant experience in Data Engineering, Big Data, or Cloud Data Engineering.
Strong hands-on experience with Hadoop, Hive, HDFS, and PySpark is mandatory.
Key Responsibilities & Technical Skills
Develop, maintain, and optimize scalable data engineering solutions using Python and Apache Spark/PySpark.
Design and build reusable, scalable data ingestion and data transformation frameworks for large-scale data processing.
Work extensively with the Hadoop ecosystem, including Hadoop, Hive, HDFS, and Impala, preferably in Cloudera-based environments.
Demonstrate strong understanding of HDFS architecture and internals, including:
Data storage architecture
Partition management
Compaction strategies
Resource utilization
Cluster performance optimization
Troubleshoot complex platform, storage, compute, and cluster-level issues, perform detailed root cause analysis, and implement performance improvements.
Work with a variety of database technologies and object storage platforms, ensuring optimized connectivity, data access, and processing patterns.
Apply strong knowledge of distributed data processing, Spark optimization, and data platform engineering best practices.
Develop and support applications using Kubernetes, including:
Containerized Spark workloads
Pod management
Application scaling
Troubleshooting
Operational support
Build, package, and deploy Spark workloads using JFrog-managed container images and automated CI/CD pipelines.
Leverage AI Agents, Generative AI, and LLM-powered development tools to improve engineering productivity and accelerate software delivery.
Collaborate with engineering and platform teams to develop reliable, scalable, and high-performance data solutions.
Mandatory Skills
Hadoop
Hive
HDFS
PySpark / Apache Spark
Python
Distributed data processing
Spark optimization
Good to Have / Added Advantage
Cloudera ecosystem
Impala
Kubernetes
Containerized Spark workloads
JFrog
CI/CD pipelines
Object storage platforms
Scala development
Generative AI / AI Agents / LLM-powered development tools
Preferred Profile
Candidates should have 5 7 years of hands-on experience in data engineering or big data platforms, with strong practical expertise in Hadoop, Hive, HDFS, and PySpark. Experience in Cloudera, Kubernetes, containerized Spark, CI/CD, and platform performance optimization will be an added advantage.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.