Live opening · Posted 12 days ago

Lead Databricks Engineer

neurogent.ai · United States (Remote)
Linkedin No
You are 12 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 12 days ago
Companyneurogent.ai
LocationUnited States (Remote)
Work modeNo
SourceLinkedin
Listed12 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
9 min from Linkedin publishing this role to us finding it
6 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,911 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Company: Neurogent.ai
Location: Remote, supporting US-based client teams
Engagement Type: Contract-to-hire
Initial Contract Duration: 6 months
Hourly Rate: $80–$90 per hour
Potential Conversion: Full-time employment may be offered at the end of the contract based on performance, business needs, and mutual fit.
About Neurogent.ai
Neurogent.ai builds agentic AI solutions that help organizations automate workflows and improve customer experiences across Banking, Healthcare, Insurance, and Financial Services.
We develop intelligent AI agents, chatbots, and voice assistants that provide 24/7 support, streamline operations, and deliver measurable efficiency gains. Using modern architectures such as Retrieval-Augmented Generation (RAG), LangGraph, machine learning, and data platforms, we help organizations reduce costs, improve productivity, and accelerate growth.
About the Role
Neurogent.ai is seeking a hands-on Databricks Lead — Data Engineering & MLOps to lead the data platform function for a large US enterprise client. The client operates more than 1,000 locations and is building its analytics, machine learning, and data platform capabilities from the ground up.
You will own the lakehouse and MLOps foundation supporting the client’s data science and analytics programs. This includes building governed data pipelines from multiple source systems, creating reliable feature and prediction tables, implementing data-quality controls, and establishing a promotion path that moves models into production with minimal manual intervention.
This is a technical leadership position—not a coordination-only role. You will write production code, establish engineering standards, review designs and pull requests, mentor engineers, and work directly with senior client stakeholders across data science, engineering, IT, and business functions.
The ideal candidate brings strong ownership and “founder mindset” characteristics: comfort with ambiguity, sound judgment with incomplete information, decisive technical leadership, and accountability for platform outcomes.
Contract-to-Hire Details
Initial six-month contract engagement.
Hourly rate of $80–$90 per hour, based on experience, technical depth, and engagement terms.
Remote role supporting US-based client and internal teams.
Full-time conversion may be considered after the initial contract period.
Conversion depends on performance, client and business requirements, budget availability, and mutual agreement.
Conversion terms, including compensation, benefits, work authorization, and employment classification, will be documented separately if an offer is made.
The contractor or staffing partner will be responsible for applicable onboarding, payroll, tax, benefits, and compliance arrangements based on the selected engagement structure.
Key Responsibilities
Technical Leadership
Lead the client’s Databricks and data platform team.
Set technical direction, engineering standards, and implementation patterns.
Plan and sequence delivery across data engineering, MLOps, and analytics initiatives.
Own platform reliability, scalability, security, cost efficiency, and delivery outcomes.
Work directly with client data science, IT, engineering, and business leadership.
Facilitate architecture discussions and working sessions.
Prepare clear technical decision documents and defend architectural recommendations.
Identify risks, make decisions with incomplete information, and keep delivery moving.
Mentor engineers, conduct code reviews, and remove technical blockers.
Create runbooks, operational documentation, and support procedures.
Data Engineering & Lakehouse Architecture
Design and build medallion architecture pipelines using Bronze, Silver, and Gold layers.
Develop scalable data pipelines using Databricks, Apache Spark, Delta Lake, Python, and SQL.
Integrate data from operational systems, CRM platforms, HRIS and payroll systems, web applications, finance systems, and other enterprise sources.
Implement incremental ingestion, change data capture, schema evolution, backfills, replayability, and historical data retention.
Build dimensional and analytical data models for reporting, experimentation, and machine learning.
Design reusable feature, training, and prediction tables with versioning and multi-year history.
Ensure datasets are reproducible, traceable, and suitable for model retraining.
Optimize Spark jobs, Delta tables, clusters, storage, partitioning, and data-processing costs.
Deliver trusted outputs to Power BI, downstream applications, APIs, and other consumer systems.
CI/CD & Software Engineering
Implement CI/CD practices for data engineering and machine learning workloads.
Establish Git-based workflows, pull-request reviews, branching standards, and release processes.
Build automated unit, integration, data-quality, and end-to-end tests.
Parameterize code and configuration across environments.
Use pinned package versions and reproducible runtime environments.
Implement secure secrets handling and eliminate manual deployment steps.
Establish a controlled approval gate for production releases.
Automate deployment of notebooks, jobs, workflows, libraries, configurations, and infrastructure where appropriate.
Use infrastructure-as-code tools and practices to improve repeatability and operational control.
Orchestrations & Operations
Design and operate time-based, trigger-based, event-driven, and manually initiated workflows.
Configure Databricks Workflows, Airflow, or comparable orchestration tools.
Provide DAG-level visibility into pipeline execution and dependencies.
Implement retries, failure handling, recovery procedures, and replay capabilities.
Define and monitor pipeline-level service-level agreements and operational targets.
Build alerts for failures, delays, SLA breaches, and abnormal processing behavior.
Maintain production runbooks and incident-response procedures.
MLOps & Model Operations
Build the MLOps foundation using MLflow and Databricks machine learning capabilities.
Implement experiment tracking, model registration, model versioning, and lifecycle management.
Create batch inference pipelines and production prediction workflows.
Establish repeatable paths from experimentation to development, staging, and production.
Support model rollback and controlled promotion between environments.
Integrate feature tables, training datasets, model artifacts, and prediction outputs.
Monitor model performance, data drift, concept drift, pipeline health, and prediction quality.
Configure alerts for model and data anomalies.
Partner with data scientists to productionize models without unnecessary rewrites.
Prepare the platform for real-time and low-latency model serving as business requirements evolve.
Data Quality & Reliability
Implement data-quality gates directly within ingestion, transformation, feature, and model pipelines.
Block downstream model runs when critical validation checks fail.
Detects schema changes, unexpected null increases, duplicate records, invalid values, and out-of-range metrics.
Establish validation rules for completeness, accuracy, consistency, uniqueness, timeliness, and referential integrity.
Alert stakeholders rather than allowing critical failures to occur silently.
Maintain quality metrics, issue history, and operational dashboards.
Required Qualifications
At least 6 years of experience in data engineering, data platforms, analytics engineering, or related roles.
Demonstrated experience leading a technical team or owning a data platform end to end.
Deep hands-on experience with Databricks, Delta Lake, Apache Spark, Databricks Workflows, Unity Catalog, and MLflow.
Candidates with equivalent depth in Snowflake, Microsoft Fabric, BigQuery, or another modern cloud data platform may be considered if they can become productive on Databricks quickly.
Strong Python and SQL skills.
Production experience building and operating E

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App