Live opening · Posted 1 day ago

Pyspark

Infosys · Bengaluru East, Karnataka, India (On-site)
Linkedin No
You are 1 day behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 1 day ago
CompanyInfosys
LocationBengaluru East, Karnataka, India (On-site)
Work modeNo
SourceLinkedin
Listed1 day ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
8 min from Linkedin publishing this role to us finding it
7 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,105 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Good to have skills: SQL, Hadoop, Hive, Kafka, Airflow
Key Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
Contribute to code reviews and follow engineering best practices to improve quality and maintainability. Minimum Qualifications:
Education: BTECH, MTECH, MCA, MSC.
2–3 years of experience in data engineering or large-scale data processing roles.
Strong hands-on experience with PySpark for building data pipelines and transformations.
Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
Ability to debug Spark applications and resolve data/job issues effectively. Preferred Qualifications:
Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App