Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Develop a deep understanding of the end-to-end configurations, dependencies, customer requirements, and overall characteristics of the production services as the accountable owner for overall service operations.
Drive incident response efforts, working closely with cross-functional teams to resolve issues quickly and minimize downtime.
Implement effective incident management processes and post-incident reviews.
Participate in on-call rotation responsibilities, ensuring timely identification and resolution of infrastructure issues.
Systems and infrastructure maintenance and ownership for production environments, with a continued focus on improving efficiencies, availability, and supportability through automation and well-defined runbooks.
Continuously optimize infrastructure cost through visibility, right-sizing, automation, and strategic cloud and management approaches.
Design for Reliability: Implement solutions that keep Navi's infrastructure running with Always.
Focus on availability and ensure a high uptime SLA for the infrastructure.
Manage individual project priorities, deadlines, and deliverables related to your technical expertise and assigned domains.
Collaborate with Product and Information Security teams to ensure the integrity and security of infrastructure and applications.
Implement security best practices and compliance standards.
Requirements:
3-5 years of experience as a DevPlatform Engineer.
Experience with Unix/Linux operating systems fundamentals and internals (e. g., filesystems, user management, DNS, networking protocols, etc. ).
Strong fundamentals of network services such as DNS, TLS/SSL, HTTP, etc.
Strong expertise in automating infrastructure provisioning and configuration using tools like Ansible, Packer, Terraform, Docker, Helm Charts, etc.
Expertise in managing large-scale cloud infrastructure (preferably AWS and Oracle).
Expertise in managing production-grade Kubernetes clusters.
Experience in scripting using programming languages like Bash, Python, Ruby, etc.
Hands-on experience in industry-standard monitoring and logging tools.
Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
A working knowledge of cloud security concepts.
Good Understanding of API Gateway, Service Mesh, etc.
Demonstrated ability to communicate clearly with both technical and non-technical people.
Project stakeholders, with the ability to work effectively in a cross-functional team environment.
Experience
3-5 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.