Live opening · Posted 8 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for a Senior Infrastructure Engineer to lead the evolution of our cloud ecosystem. In this role, you won't just be "keeping the lights on"; you will be architecting the foundation that powers our next generation of AI-driven products. You will bridge the gap between software engineering and systems operations, ensuring our platform is scalable, resilient, and optimized for high-performance workloads.
Responsibilities:
Search Stack Orchestration: Architect, scale, and maintain our core search and NoSQL technologies, including Solr (indexing/sharding), HBase (distributed
storage), and Redis (high-speed caching).
Cloud Architecture and IaC: Use Terraform to manage multi-cloud environments (AWS/GCP), ensuring that the search stack and its supporting resources are fully versioned and reproducible.
Kubernetes Mastery: Oversee the deployment of data services within Kubernetes, focusing on stateful sets, persistent storage performance, and resource isolation for search workloads.
Evolution toward AI Search: Lead the infrastructure integration of vector search databases and high-performance compute to support AI-driven architectures.
Observability and Reliability: Implement deep-stack monitoring and alerting using Datadog and Prometheus to ensure proactive issue detection and resolution.
Automation-First Mindset: Maintain and evolve an active codebase in Python, Go, or Bash to automate repetitive tasks. You will also integrate LLMs into your workflow to accelerate scripting, documentation, and operational efficiency.
Networking and Traffic: Manage complex cloud networking topologies, including VPCs, Load Balancing, Service Meshes, and caching layers (e. g., Redis) to minimize latency.
Incident Response: Lead the debugging of complex, distributed systems issues, performing root cause analysis to prevent recurrence, as part of an on-call rotation.
Requirements:
Experience: 7+ years of experience in infrastructure, DevOps, or SRE, with a significant portion dedicated to managing high-scale distributed data systems.
Search and NoSQL Expertise: Direct experience with Solr, HBase, and Redis is highly preferred. Experience with similar technologies (e. g., Elasticsearch/OpenSearch, Cassandra, or BigTable) is acceptable if you have the senior-level depth to transition quickly.
Linux Internals: Expert-level knowledge of Linux performance tuning, specifically for data-intensive applications (I/O scheduling, memory management, and JVM tuning).
Infrastructure as Code: Proven track record of managing large-scale infrastructure using Terraform or OpenTofu.
Containerization: Deep experience running production-grade stateful workloads on Kubernetes.
Software Engineering: Strong proficiency in Python or Go, with the ability to build custom tools and operators to manage data lifecycles.
AI-Augmented Engineering: Practical experience using LLMs (GitHub Copilot, ChatGPT, Claude) to increase your personal and team productivity.
Abilities Required:
Demonstrated ability to learn new technologies quickly and independently.
Strong technical, organizational, and interpersonal skills.
Strong written and verbal communication skills.
Must be able to read, understand, and communicate complex problems and solutions in English over a textual medium (such as Slack).
Experience
7-11 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.