Live opening · Posted 8 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Develop a deep understanding of the end-to-end configurations, dependencies, customer requirements, and overall characteristics of the production services as the accountable owner for overall service operations.
Implementing best practices, challenging the status quo, and tab on industry and technical trends, changes, and developments to ensure the team is always striving for best-in-class work.
Lead incident response efforts, working closely with cross-functional teams to resolve issues quickly and minimise downtime.
Implement effective incident management processes and post-incident reviews.
Participate in on-call rotation responsibilities, ensuring timely identification and resolution of infrastructure issues.
Possess expertise in designing and implementing capacity plans, accurately estimating costs and efforts for infrastructure needs.
Systems and Infrastructure maintenance and ownership for production environments, with a continued focus on improving efficiencies, availability, and supportability through automation and well-defined runbooks.
Provide mentorship and guidance to a team of DevOps engineers, fostering a collaborative and high-performing work environment.
Mentor team members in best practices, technologies, and methodologies.
Design for Reliability - Architect and implement solutions that keep Navi's Infrastructure running with Always On availability and ensure high uptime SLA for the Infrastructure.
Manage individual project priorities, deadlines, and deliverables related to your technical expertise and assigned domains.
Collaborate with Product and Information Security teams to ensure the integrity and security of the infrastructure and applications.
Implement security best practices and compliance standards
Requirements:
5-7 years of experience as Devops / SRE / Platform Engineer.
Strong expertise in automating Infrastructure provisioning and configuration using tools like Ansible, Packer, Terraform, Docker, Helm Charts, etc.
Strong skills in network services such as DNS, TLS/SSL, HTTP, etc.
Expertise in managing large-scale cloud infrastructure (preferably AWS and Oracle).
Expertise in managing production-grade Kubernetes clusters.
Experience in scripting using programming languages like Bash, Python, etc.
Expertise in skill sets for centralised logging systems, metrics, and tooling frameworks such as ELK, Prometheus/VictoriaMetrics, and Grafana, etc.
Experience in managing and building high-scale API Gateway, Service Mesh, etc.
Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
Have a working knowledge of a backend programming language.
Deep knowledge and experience with Unix / Linux operating systems internals (Eg, filesystems, user management, etc. )
A working knowledge and deep understanding of cloud security concepts.
Proven track record of driving results and delivering high-quality solutions in a fast-paced environment.
Demonstrated ability to communicate clearly with both technical and non-technical project stakeholders, with the ability to work effectively in a cross-functional team environment.
Experience in developing in-house tools to improve our ability to monitor and maintain multi-cloud Infrastructure.
Experience in self-managed Data Centre based Infrastructure.
Strong cloud foundation with certifications in AWS and container orchestration (CKA/CKAD) is a plus.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.