Live opening · Posted 8 days ago

Sr. Data Engineer

SimpliSafe · Work From Home
Instahyre 5-7 yrs
You are 8 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 8 days ago
CompanySimpliSafe
LocationWork From Home
Experience5-7 yrs
SourceInstahyre
Listed8 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
73,703 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are looking for an experienced Data Engineer to join our machine learning team and serve as the vital link between raw data and high-performing ML models. In this role, you will drive the design, implementation, and management of an integrated system dedicated to data curation, annotation workflows, quality assurance, and dataset composition. In this role, you will drive the design and implementation of an integrated data system dedicated to advanced data curation, annotation workflows, quality assurance, and dataset composition. If you have a passion for data infrastructure, automation, and optimizing human-in-the-loop annotation processes at scale, we'd love to hear from you.
Responsibilities:
Data Lake Management: Design, build, and maintain our AWS data lake architecture (focusing on S3 Athena, and Glue) to serve as the highly accessible foundation for all ML data workloads.
ETL and Pipeline Development: Architect and build scalable, robust data pipelines to ingest, transform, and deliver massive datasets seamlessly.
Build Specialized Datasets: Partner closely with ML engineers and modelers to curate raw data and prepare high-quality baseline training sets, experimental datasets for testing new ML hypotheses, and targeted validation sets specifically designed to evaluate "hard cases" and edge cases.
Advanced Data Curation: Collaborate with the modeling team to leverage data embeddings to cluster, filter, and surface the most informative data points for
annotation, ensuring we are spending our annotation budget on the highest-value data.
Platform Integration: Build secure, reliable integrations between our internal data ecosystem and specialized third-party data annotation platforms.
Manage the Annotation Lifecycle: Oversee end-to-end annotation workflows, including job creation, platform configuration, and syncing/maintaining large annotated datasets.
Ensure Data Quality: Partner with ML engineers and modelers to conduct spot annotation quality assurance (QA), establishing a shared workflow to ensure labeled data meets strict accuracy and consistency standards at scale.
Data Governance and Schema Design: Design optimized database schemas for large-scale data exploration and implement best practices in data cataloging, dataset version control, and data stewardship.
Best Practices: Advocate for and implement best practices in data stewardship, dataset version control, and system monitoring.
Requirements:
Experience: 5+ years of experience in data engineering, with at least 3 years explicitly focused on building data pipelines, managing data lakes, and curating large-scale datasets.
Cloud Data Ecosystems: Strong, hands-on experience with the AWS Data stack, specifically Amazon S3 Athena, and Glue (or deep expertise in an equivalent ecosystem like GCP BigQuery).
Programming and Databases: High proficiency in Python and advanced SQL, with a proven track record of designing robust database schemas and complex ETL processes.
Navigating Ambiguity: Proven skill in developing organized, reliable data structures and pipelines within highly unstructured data settings.
Data Lifecycle: Solid understanding of the ML data lifecycle, including data tracking, dataset versioning, and the specific data structures required by modeling teams.
Communication: Exceptional communication skills, specifically the ability to translate complex data requirements into easily understandable instructions for annotation teams.
Nice to Have:
Embedding and Vector Tech: Experience working with embeddings for data curation as well as familiarity with vector databases or similarity search libraries (e. g., FAISS, Pinecone, and Milvus).
Big Data Technologies: Experience with distributed computing and large-scale data processing frameworks (e. g., Spark, Hadoop, Ray, or AWS EMR).
Annotation Platforms: Hands-on experience with specialized data labeling and annotation platforms (e. g., Scale AI, Labelbox, Snorkel, Toloka).
Data Tracking Tools: Familiarity with tools such as MLflow or Weights and Biases for dataset and artifact tracking.

Experience
5-7 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App