Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.
Career Level - IC4
Key Responsibilities
Capacity Ingestion and
Management:
-
Designs
and architects infrastructure and/or service according to terms for reliability
and functionality.
-
Forecasts
demands for infrastructure and responds to capacity needs, ensuring systems have
sufficient resources to handle current and future workloads and identifying
resource gaps.
-
Collaborates
with the software development team to develop infrastructures, ensuring
features are reliable and scalable according to deployment requirements.
-
Proactively
identifies opportunities for prototyping and drives prototyping initiatives
(e.g., testing new applications or infrastructures, assisting in onboarding) to
explore novel approaches.
Incident and Service
Lifecycle Management:
-
Exercises
judgment when performing data collection, triage, technical analysis, and
redirection to maintain and optimize operations and infrastructure reliability.
-
Takes
proactive steps to monitor services, maintain up-to-date knowledge of their
performance, and document their condition.
-
Leverages
advanced knowledge to perform incident response, root cause analyses, and/or
maintenance on assigned services (e.g., software installs, version upgrades,
security updates, backup and recovery).
-
Provides
comprehensive health and performance reporting and takes appropriate actions
based on trends in data.
-
May
perform provisioning to support infrastructure, applications, and services.
-
May
experiment with new approaches for and performs decommissioning (e.g., shutting
down servers, removing data from databases) to remove objects that are no
longer needed.
Automation:
-
Identifies
and recommends opportunities for automation and assesses potential benefits to
enhance operational efficiency.
-
Develops
and implements design, automation tools, or scripts to provide solutions,
gather metrics, monitor, analyze, mitigate, or remediate issues/defects within
infrastructures.
-
Conducts
testing on moderately complex automations to ensure they perform tasks
correctly and produce expected results.
Technical Communication and
Guidance:
-
Writes
release notes and/or communicates comprehensive inform
More openings worth a look
Recently tracked roles with full details and direct application links.