Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are hiring exceptional engineers to develop a fully autonomous AI ecosystem that can create, execute, and optimise business strategy end-to-end. A single prompt like "Create a new product line and grow revenue by 10%" triggers AI agents that design products, allocate capital, run experiments, learn from results, and deliver measurable business outcomes with minimal human input.
This role centres on strengthening the MLOps portion of AgentOS, the secure, scalable cloud infrastructure; CI/CD and deployment pipelines for the agent framework; model lifecycle management (versioning and rollout/rollback of the LLMs and SLMs); GPU and compute provisioning; monitoring and model-drift tracking; data security; and the auditability of every agent action. The new joiner will own and continuously improve this MLOps layer as the platform scales.
The core responsibilities for the job include the following:
Cloud Infrastructure and Security:
Infrastructure as Code (IaC): Design and deploy robust, multi-region cloud infrastructure (AWS, GCP, or Azure) using tools like Terraform or CloudFormation.
Security Posture: Implement and manage security protocols, including VPC configuration, firewall rules, access control (IAM/RBAC), and ensuring data encryption at rest and in transit (critical for financial data).
Compute Management: Manage and optimize GPU clusters and compute resources for high-performance model training (SLM fine-tuning) and low-latency inference (agent decision-making).
MLOps and Deployment Automation:
CI/CD for Agents: Establish a standardized CI/CD pipeline for the HTP framework, covering the strategic (PAEP), tactical (ReAct), and worker agents.
Model Lifecycle Management: Implement MLOps tools (e. g., MLflow, SageMaker, or Vertex AI) for experiment tracking, model versioning, and automated rollout/rollback capabilities for the LLMs and SLMs.
Containerization: Manage deployment via container technologies (Docker and Kubernetes/ECS), ensuring efficient resource allocation and environment consistency across development and production.
Monitoring and Auditing:
Monitoring: Implement comprehensive logging and monitoring systems (e. g., Prometheus and Grafana) to track agent performance, model drift, and infrastructure health.
Auditability: Ensure all agent actions and decisions are logged and auditable, fulfilling the fiduciary requirements of the private equity firm.
Requirements:
Cloud Expertise: 3+ years of experience managing production infrastructure on a major cloud provider (AWS, GCP, or Azure). GCP is preferred.
MLOps Tooling: Hands-on experience with MLOps platforms (MLflow, Kubeflow) and version control (Git).
Containerization and Orchestration: Strong command of Docker and Kubernetes or similar container orchestration systems.
IaC and CI/CD: Proficiency with Terraform and building automated CI/CD pipelines (e. g., Jenkins and GitHub Actions).
Linux/Scripting: Strong scripting skills (Bash, Python) for automation.
Desirable Domain Experience:
High-Performance Compute: Experience provisioning and managing high-demand compute resources (GPU instances).
Security Focus: Experience hardening cloud environments for sensitive data applications.
Financial Systems: Experience working with financial systems is a strong plus.
Experience
3-6 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.