Live opening · Posted 18 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
PUNE LOCATION IS AVAILABLE
NO FRESHER WILL BE PREFERRED
We are looking for Senior Data Engineer (Databricks Python Pyspark) as a Full Time for our IT Company.
Location- Pune
Experience- 6+ years
Budget- As per company standard
Notice Period- Immediate Joiner
Data Engineer
Project Description
Horizon is a data and AI platform designed to collect, process, and analyze travel-market signals from multiple external sources, including social-listening platforms, travel-data providers, SEO/AEO tools, RSS feeds, web sources, and public APIs.
The platform will use Databricks as its central data-processing and governance environment. It will ingest and normalize structured, semi-structured, and unstructured data, apply data-quality and privacy controls, and provide reliable datasets for signal scoring, travel-corridor analysis, automated reporting, and downstream AI services.
Responsibilities
Design, build, and maintain scalable data pipelines in Databricks.
Develop ingestion pipelines for REST APIs, RSS feeds, web sources, files, and third-party data providers.
Transform heterogeneous source data into standardized travel-signal and corridor data models.
Implement batch and scheduled data-processing workflows using Python, SQL, PySpark, and Databricks Workflows.
Design and maintain Delta Lake tables and appropriate Bronze, Silver, and Gold data layers.
Configure and use Unity Catalog for data governance, access control, metadata management, and lineage.
Implement data-quality controls, schema validation, deduplication, error handling, and reconciliation mechanisms.
Build processing logic for rolling time windows, signal aggregation, scoring inputs, rankings, and reporting datasets.
Implement data-retention and automated data-cleanup processes.
Support PII detection, filtering, masking, and other data-privacy requirements.
Optimize Spark jobs, queries, storage layouts, cluster utilization, and pipeline performance.
Implement monitoring, logging, alerting, and operational support for data pipelines.
Prepare curated datasets for backend services, analytical components, reports, and AI/LLM workloads.
Integrate Databricks with AWS services, external APIs, SharePoint, and the PRODIGY agentic platform.
Create automated unit, integration, and data-quality tests.
Participate in technical design, backlog refinement, estimation, code reviews, and release activities.
Maintain technical documentation, data mappings, pipeline specifications, and operational runbooks.
Requirements
6+ years of professional experience in data engineering.
Strong hands-on experience with Databricks in production environments.
Advanced knowledge of Python, SQL, and PySpark.
Strong understanding of Apache Spark architecture and distributed data processing.
Experience with Delta Lake, Delta tables, schema evolution, partitioning, and performance optimization.
Practical experience with Databricks Workflows, Jobs, notebooks, and cluster configuration.
Experience with Unity Catalog, data lineage, permissions, and data-governance concepts.
Proven experience building production-grade ETL/ELT and data-ingestion pipelines.
Experience processing structured, semi-structured, and unstructured data, including JSON, CSV, text, and API responses.
Experience integrating third-party REST APIs, including authentication, pagination, rate limits, retries, and failure handling.
Strong understanding of data modelling, data quality, metadata management, and data observability.
Experience with automated testing and CI/CD for data pipelines.
Working knowledge of AWS services, cloud security, IAM, networking, and secrets management.
Familiarity with Git, infrastructure automation, and Agile delivery practices.
Understanding of PII protection, data retention, RBAC, and enterprise security requirements.
Good communication skills and the ability to collaborate with architects, backend developers, DevOps engineers, QA engineers, and business stakeholders.
Nice to Have
Experience with Databricks Auto Loader, Lakeflow Declarative Pipelines or Delta Live Tables.
Experience with streaming pipelines or Structured Streaming.
Familiarity with MLflow, feature engineering, vector databases, embeddings, or LLM data pipelines.
Experience preparing data for NLP, generative AI, or agentic AI solutions.
Knowledge of travel, marketing, social-listening, or market-intelligence data.
Experience working in financial services or another regulated enterprise environment.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.