Live opening · Posted 6 days ago

Data Engineer

Soteris · United States (Remote)
Linkedin Yes
You are 6 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 6 days ago
CompanySoteris
LocationUnited States (Remote)
Work modeYes
SourceLinkedin
Listed6 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
4 min from Linkedin publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,969 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
Python, SQL, PySpark, pandas, and related data-engineering libraries
Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
Customer databases, backups, SFTP, APIs, and file-based ingestion
Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
MLflow, SageMaker, and production scoring APIs at the model handoff boundary
Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
Experience integrating with policy administration, claims management, rating, or other operational source systems.
Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App