Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
1. About Our Client:
The organization operates within the machine learning infrastructure and platform engineering space. It addresses challenges related to managing scalable, reliable, and cost-efficient machine learning workflows and production systems. By focusing on optimizing compute resources, ensuring operational reliability, and automating deployment pipelines, the organization supports high-throughput model workflows and the integration of machine learning models into production environments.
2. About the Opportunity:
The MLOps / ML Platform Engineer role is responsible for designing, operating, and optimizing machine learning infrastructure that supports the end-to-end lifecycle of model development and deployment. This position plays a critical role in ensuring that machine learning models are efficiently trained, evaluated, and served in production while maintaining reliability, security, and cost-effectiveness. The role contributes by enabling smooth collaboration between research and engineering teams and by developing tools and automation that enhance the repeatability and safety of ML workflows.
3. Responsibilities:
Design and manage ML infrastructure including data, training, serving, and inference systems
Build scalable, reproducible training and evaluation pipelines with versioning and scheduling
Optimize compute workloads and costs through tuning and cluster management
Operate production model serving with autoscaling and deployment safety measures
Define and monitor service level objectives (SLOs) and track system observability metrics
Manage security aspects such as IAM, secrets, and container security
Automate deployment pipelines using CI/CD and infrastructure as code
Collaborate with research scientists and AI engineers to facilitate model production
Create documentation, templates, and internal tools to improve ML workflow efficiency
4. Requirements:
Minimum 4 years of experience in ML platform, DevOps, or infrastructure engineering
Strong knowledge of Kubernetes, CI/CD, containers, and cloud platforms (AWS, GCP, or Azure)
Experience managing GPU clusters and ML training/inference pipelines
Familiarity with data orchestration and storage formats such as Delta, Parquet, Polars, and Spark
Proven ability to deploy and operate production ML systems with defined SLOs
Proficient in Python and automation using infrastructure as code
Experience with observability tools and cost optimization at scale
5. Pay Range and Compensation Package:
The pay range and compensation package for this role will be determined based on the candidate’s experience, skills, and other relevant factors.
6. Benefits & Perks:
Comprehensive health insurance plan
Retirement savings plan (401k) with company match
Remote working environment
Flexible, unlimited time off policy
Thirteen paid holidays annually, including the Monday after the Super Bowl
Equal Opportunity Statement:
Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.
Note:
RemoteHunter is a recruitment partner of this role. Please note that all employment decisions, including candidate assessment, interviews, hiring, compensation, and employment terms, are made exclusively by the hiring employer.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.