Live opening · Posted 3 days ago

Data Engineer

FirstHive | CDP+AI Data Platform · Bengaluru, Karnataka, India (On-site)
Linkedin No
You are 3 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 3 days ago
CompanyFirstHive | CDP+AI Data Platform
LocationBengaluru, Karnataka, India (On-site)
Work modeNo
SourceLinkedin
Listed3 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
12 min from Linkedin publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
57,979 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We're hiring a Data Engineer in Bangalore to build the platform-level data infrastructure that powers FirstHive's Customer Data Platform across every client integration — connector frameworks, CDC ingestion pipelines, transformation services, and data quality tooling.
You'll work in Java and Spring Boot on Kafka, Kafka Connect, MongoDB, and our analytical warehouse stack (StarRocks, Snowflake, BigQuery), deployed on multi-cloud Kubernetes.
The role is framework-building, not client work — you design once for the next ten clients.
FirstHive ingests customer data from ERPs, CRMs, mobile apps, PoS systems, social, voice, and customer care across hundreds of enterprise deployments. The bottleneck is not data volume — it is the variety of source systems and the cost of onboarding each new client. We need someone who builds the systems that turn client onboarding into configuration, not engineering.
What you'll build
Organized by where it sits in the pipeline. Within each group, the most architectural work comes first.
Ingestion frameworks — the first thing every new client hits
• Pluggable connector framework over Kafka and Kafka Connect for databases, APIs, file feeds, and event streams — custom SMTs, DLQ patterns, reusable connector configurations.
• CDC ingestion pipelines from MongoDB and relational sources via Debezium and Kafka Connect — multi database routing, schema change handling, ordering guarantees.
• Automated schema mapping, detection, and inference tooling — so onboarding a new client is configuration, not engineering.
Transformation and quality — what makes the data trustworthy
• Transformation layer as composable Spring Boot modules: cleaning, deduplication, normalization, identity resolution, enrichment. Not glue scripts.
• Data quality framework — profiling, validation gates, anomaly detection, lineage tracking — wired into the pipeline so bad data is caught at ingestion, not at the dashboard.
• Schema evolution handling — backward compatibility across Kafka topics, transformations, and warehouse tables when client source schemas change.
Warehouse layer — where it lands and gets queried
• Data models for StarRocks, Snowflake, and BigQuery — partitioning, clustering / bucketing, materialization strategy, primary-key vs. duplicate vs. aggregate table design.
• Optimized SQL and stored procedures for mixed workloads: point lookups, high-concurrency customer profile dashboards, and large batch ETL.
• Metadata layer driving per-client schema definitions, mapping rules, and transformation logic — controlled by configuration, not code changes.
What we need:
Grouped by where it matters. The first bullet of each group is the non-negotiable.
4+ years building data systems — not running them
You've designed and shipped framework-level data systems in production. You can point to ones still running.
Production-grade Java and Spring Boot
• Real microservices: error handling, observability, testing, lifecycle management. Not scripts.
• Framework-builder instinct — reusable tooling for the next ten clients, not the next ticket.
SQL fluency and data modeling depth
• Complex joins, window functions, CTEs (including recursive), and a real instinct for performance and cost. • Star schema, SCD types, event sourcing, EAV patterns — and judgment on when each is the right answer.
Real depth on the streaming and warehouse stack
• Kafka and Kafka Connect at depth: connector configuration, custom transforms and converters, consumer group design, DLQ patterns, exactly-once vs. at-least-once tradeoffs.
• At least one analytical warehouse at architecture level — StarRocks, Snowflake, or BigQuery — covering data modeling, performance tuning, partitioning / clustering, and cost optimization.
• MongoDB or similar document store — schema design, compound indexing, change streams, CDC tradeoffs.
• Workflow orchestration in production — Airflow, Argo Workflows, dbt, or similar.
Bonus, not gating
These don't decide the hire, but they shape the shortlist:
• Debezium at production scale — buffer / lock tuning, multi-database capture, snapshot strategies.
• StarRocks, ClickHouse, Druid, or similar MPP / OLAP engines.
• Open table formats — Apache Iceberg, Hudi — and lakehouse architectures.
• Multi-cloud Kubernetes (GKE, EKS) and object storage (GCS, S3).
• Identity resolution — deterministic / probabilistic matching, graph-based stitching.
• CDP, MarTech, or AdTech domain exposure.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App