Live opening · Posted 11 hours ago

Senior Data Engineer (Databricks Pyspark Python)

Haparz · Chengalpattu, Tamil Nadu, India (Hybrid)
Linkedin No
You are 11 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 hours ago
CompanyHaparz
LocationChengalpattu, Tamil Nadu, India (Hybrid)
Work modeNo
SourceLinkedin
Listed11 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
4 min from Linkedin publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
20,770 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

CHENNAI LOCATION IS AVAILABLE
NO FRESHER WILL BE PREFERRED
We are looking for Senior Data Engineer (Databricks Pyspark Python) as a Full Time for our IT Company.
Location- Chennai
Experience- 6+ years
Budget- As per company standard
Notice Period- Immediate Joiner
Data Engineer
Project Description
Horizon is a data and AI platform designed to collect, process, and analyze travel-market signals from multiple external sources, including social-listening platforms, travel-data providers, SEO/AEO tools, RSS feeds, web sources, and public APIs.
The platform will use Databricks as its central data-processing and governance environment. It will ingest and normalize structured, semi-structured, and unstructured data, apply data-quality and privacy controls, and provide reliable datasets for signal scoring, travel-corridor analysis, automated reporting, and downstream AI services.
Responsibilities
Design, build, and maintain scalable data pipelines in Databricks.
Develop ingestion pipelines for REST APIs, RSS feeds, web sources, files, and third-party data providers.
Transform heterogeneous source data into standardized travel-signal and corridor data models.
Implement batch and scheduled data-processing workflows using Python, SQL, PySpark, and Databricks Workflows.
Design and maintain Delta Lake tables and appropriate Bronze, Silver, and Gold data layers.
Configure and use Unity Catalog for data governance, access control, metadata management, and lineage.
Implement data-quality controls, schema validation, deduplication, error handling, and reconciliation mechanisms.
Build processing logic for rolling time windows, signal aggregation, scoring inputs, rankings, and reporting datasets.
Implement data-retention and automated data-cleanup processes.
Support PII detection, filtering, masking, and other data-privacy requirements.
Optimize Spark jobs, queries, storage layouts, cluster utilization, and pipeline performance.
Implement monitoring, logging, alerting, and operational support for data pipelines.
Prepare curated datasets for backend services, analytical components, reports, and AI/LLM workloads.
Integrate Databricks with AWS services, external APIs, SharePoint, and the PRODIGY agentic platform.
Create automated unit, integration, and data-quality tests.
Participate in technical design, backlog refinement, estimation, code reviews, and release activities.
Maintain technical documentation, data mappings, pipeline specifications, and operational runbooks.
Requirements
6+ years of professional experience in data engineering.
Strong hands-on experience with Databricks in production environments.
Advanced knowledge of Python, SQL, and PySpark.
Strong understanding of Apache Spark architecture and distributed data processing.
Experience with Delta Lake, Delta tables, schema evolution, partitioning, and performance optimization.
Practical experience with Databricks Workflows, Jobs, notebooks, and cluster configuration.
Experience with Unity Catalog, data lineage, permissions, and data-governance concepts.
Proven experience building production-grade ETL/ELT and data-ingestion pipelines.
Experience processing structured, semi-structured, and unstructured data, including JSON, CSV, text, and API responses.
Experience integrating third-party REST APIs, including authentication, pagination, rate limits, retries, and failure handling.
Strong understanding of data modelling, data quality, metadata management, and data observability.
Experience with automated testing and CI/CD for data pipelines.
Working knowledge of AWS services, cloud security, IAM, networking, and secrets management.
Familiarity with Git, infrastructure automation, and Agile delivery practices.
Understanding of PII protection, data retention, RBAC, and enterprise security requirements.
Good communication skills and the ability to collaborate with architects, backend developers, DevOps engineers, QA engineers, and business stakeholders.
Nice to Have
Experience with Databricks Auto Loader, Lakeflow Declarative Pipelines or Delta Live Tables.
Experience with streaming pipelines or Structured Streaming.
Familiarity with MLflow, feature engineering, vector databases, embeddings, or LLM data pipelines.
Experience preparing data for NLP, generative AI, or agentic AI solutions.
Knowledge of travel, marketing, social-listening, or market-intelligence data.
Experience working in financial services or another regulated enterprise environment.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App