Live opening · Posted 4 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Role: Mid-level Big Data Engineer
Position Type: Full-Time Contract (40hrs/week)
Contract Duration: 6 months + extendable
Work Schedule: 8 hours/day (Mon-Fri)
Location: Hybrid - Wed/Thursday to be in onsite in Hyderabad, India
We are seeking a Mid-Level Big Data Engineer with 5+ years of experience developing scalable Big Data solutions and distributed data-processing applications in AWS environments.
The ideal candidate will have strong hands-on experience with Scala, Python, Apache Spark/PySpark, AWS, Linux, and shell scripting, along with experience building batch and streaming data pipelines and optimizing large-scale distributed workloads.
Key Responsibilities
Design, develop, test, and deploy scalable Big Data solutions on AWS.
Build and maintain batch and streaming data pipelines using Scala, Python, Spark, and PySpark.
Process and transform large volumes of structured and unstructured data.
Develop and support distributed data-processing and data-ingestion platforms.
Integrate new data sources and technologies into existing Big Data ecosystems.
Optimize Spark, Hadoop, EMR, and distributed computing workloads.
Build cloud-native APIs and microservices supporting data platforms.
Automate data workflows, deployments, monitoring, and operational processes.
Collaborate with architects, product owners, QA teams, and business stakeholders.
Apply security, governance, reliability, and software engineering best practices.
Requirements
5+ years of experience in Big Data development and data/software engineering.
Strong hands-on experience with Scala and Python.
Strong experience with Apache Spark / PySpark.
Experience with the Hadoop ecosystem, including HDFS, MapReduce, and Hive.
Experience with Kafka or similar streaming technologies.
Strong hands-on AWS experience, particularly with EMR, S3, EC2, ECS/EKS, MWAA/Airflow, Step Functions, API Gateway, Lambda, DynamoDB, and/or RDS/Aurora.
Experience building large-scale batch and API-based data-processing solutions.
Strong Linux and shell scripting skills.
Experience with data ingestion, transformation, data modeling, schema design, and data lifecycle management.
Strong understanding of performance tuning for Spark, Hadoop, EMR, and distributed systems.
Experience working in Agile/Scrum environments.
Strong troubleshooting and problem-solving skills.
Preferred
Generative AI / AI-assisted development
LLMs, RAG, Vector Databases, and AI/ML integrations
AI-enabled data platforms
Git, GitHub, GitLab, Bitbucket, Jenkins, Maven, Gradle, Artifactory
Docker / Kubernetes
Automated testing and data quality tools
Enterprise-scale cloud environments
AWS or Databricks certifications
Work arrangement
Hybrid
More openings worth a look
Recently tracked roles with full details and direct application links.