Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Microsoft Fabric + PySpark Performance Engineer
Experience: 6-8 Years
Location: PAN India
Role Summary
Hands-on Senior Data Engineer with strong expertise in Microsoft Fabric and PySpark to build, optimize, and support scalable enterprise data platforms. Experience in designing Fabric-based solutions and tuning Spark workloads for performance, scalability, and reliability across batch and real-time processing environments.
Must-Have Skills
Microsoft Fabric: Data Engineering, Data Factory, Data Warehouse, Lakehouse, OneLake, Real-Time Analytics, Notebooks, Pipelines, Dataflows, Power BI Integration
Strong hands-on experience in PySpark, Spark SQL, Python, and SQL
Experience building Fabric data pipelines, notebooks, Lakehouses, and Data Warehouse solutions
Expertise in Spark Performance Tuning:
Spark UI & execution-plan analysis
Shuffle optimization
Partitioning/repartitioning strategies
Broadcast joins
Caching & persistence
Data skew mitigation
Adaptive Query Execution (AQE)
Executor/driver memory tuning
Parallelism optimization
OOM and long-running job troubleshooting
Delta Lake optimization including partitioning, compaction, OPTIMIZE, VACUUM, file management, and merge/upsert tuning
Strong knowledge of Data Modeling, Dimensional Modeling, ETL/ELT, Data Integration, Data Quality, Validation, and Error Handling
Experience with ADLS, Azure SQL, Event Hubs, and related Azure data services
Experience with real-time ingestion and analytics using Fabric Eventstreams and Spark Structured Streaming
Experience integrating Fabric solutions with Power BI and downstream systems
Experience with Git, Azure DevOps/GitHub Actions, and CI/CD pipelines
Strong understanding of Fabric governance, security, compliance, monitoring, and production support
Agile/Scrum delivery experience
Key Responsibilities
Design, develop, and support scalable Microsoft Fabric solutions
Build and optimize data pipelines, notebooks, Lakehouses, and Data Warehouse workloads
Develop high-performance PySpark and Spark SQL solutions for large-scale batch and streaming workloads
Analyze Spark jobs and eliminate performance bottlenecks using Spark UI and execution plans
Optimize joins, partitions, shuffles, caching, AQE, memory usage, and Delta Lake workloads
Design reliable ingestion and transformation frameworks with robust data-quality controls
Implement real-time analytics and streaming solutions
Integrate Fabric data platforms with Power BI and enterprise applications
Implement CI/CD, monitoring, source control, and operational best practices
Contribute to solution design, code reviews, troubleshooting, and performance optimization initiatives
Good to Have
Fabric Capacity Management
OneLake Shortcuts
Fabric Deployment Pipelines
Semantic Models
Synapse-to-Fabric or Databricks-to-Fabric migration experience
Microsoft Fabric / Azure Data Engineer / Spark certifications
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.