Live opening · Posted 11 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Key ResponsibilitiesCloud Infrastructure & Platform Engineering
Design, deploy, and maintain scalable, secure, and highly available cloud infrastructure on AWS.
Apply the AWS Well-Architected Framework across platform architecture and infrastructure.
Build and maintain Infrastructure-as-Code using Terraform.
Manage and optimize containerized workloads using Docker and Amazon ECS Fargate.
Ensure infrastructure is secure, well maintained, reliable, and cost-effective.
Identify and resolve platform-level and network-level latency and performance issues.
Support production services and infrastructure as part of operational responsibilities and the on-call rotation.
Performance, Scalability & Reliability
Design and implement solutions that improve scalability, reliability, availability, and performance.
Optimize high-throughput APIs and distributed systems.
Implement caching strategies and edge technologies such as CDN to improve application performance.
Implement deployment strategies including blue-green deployments and canary releases.
Design systems that gracefully degrade and remain resilient during failures.
Implement auto-scaling and high-availability architectures.
Define and improve SLIs, SLOs, and SLAs in line with SRE principles.
DevOps & Automation
Build and maintain robust CI/CD pipelines using GitHub Actions.
Automate infrastructure provisioning, application deployment, monitoring, and operational processes.
Develop automation and internal tooling using Python, Bash, or PowerShell.
Continuously improve developer productivity through better tooling, processes, and platform capabilities.
Promote a strong DevOps culture across engineering teams.
Security & Access Management
Implement cloud security best practices across infrastructure and applications.
Secure APIs using appropriate authentication and authorization mechanisms.
Implement RBAC/ABAC for fine-grained access control.
Apply Zero Trust Architecture principles where appropriate.
Manage sensitive credentials and application secrets securely using AWS Secrets Manager.
Ensure platform designs meet appropriate security, compliance, and operational requirements.
Observability & Monitoring
Implement comprehensive monitoring and observability across cloud infrastructure and applications.
Work with AWS CloudWatch Logs, OpenTelemetry, SigNoz, and Grafana.
Build and maintain meaningful Grafana dashboards and monitoring integrations.
Proactively identify performance, reliability, and availability issues.
Use metrics, logs, and traces to troubleshoot production incidents and improve system performance.
Collaboration & Technical Leadership
Work closely with software developers to build and operate reliable production services.
Provide technical guidance across architecture, designs, implementations, and operational practices.
Evangelize DevOps, SRE, cloud engineering, and security best practices.
Make strong technical decisions while remaining open to alternative approaches and ideas.
Collaborate with third-party vendors and evaluate emerging technologies.
Mentor engineers and contribute to the technical growth of the wider team.
Engage with cross-functional teams to automate and improve tools and processes.
Create, review, and maintain high-quality technical documentation.
Participate in Agile processes and use tools such as Jira.
Required Skills & Experience
5+ years of industry experience, preferably with a systems engineering background and strong software development/coding capabilities.
Expert-level experience with AWS and the AWS Well-Architected Framework.
Proven experience designing, deploying, maintaining, and operating production workloads at scale on public cloud platforms.
Strong experience working within DevOps, Platform Engineering, or SRE teams.
Extensive hands-on experience with:
AWS ECS / ECS Fargate
Docker
Terraform
GitHub Actions
AWS Secrets Manager
Strong understanding of cloud infrastructure, networking, scalability, availability, and resilience.
Experience designing and operating RESTful APIs, JSON-based services, distributed systems, microservices, and event-driven architectures.
Strong understanding of operational and non-functional requirements, including:
Monitoring
Performance testing
Scalability
Availability
Resilience
Security
Proficiency with observability and monitoring tools such as CloudWatch, OpenTelemetry, SigNoz, and Grafana.
Expert-level experience designing and managing CI/CD pipelines using GitHub Actions.
Strong knowledge of API security, Zero Trust Architecture, RBAC, and ABAC.
Strong understanding of high-availability architecture, auto-scaling, and performance tuning.
Strong knowledge of SRE principles, including SLA, SLO, and SLI.
Strong scripting and automation skills using Python, Bash, or PowerShell.
Familiarity with Agile methodologies and tools such as Jira.
Strong communication, collaboration, leadership, and mentoring skills.
AWS DevOps certification is preferred/required.
Preferred Qualifications
Experience optimizing cloud and network infrastructure for low latency.
Experience with SaaS platform engineering and large-scale production environments.
Experience implementing CDN, caching, blue-green deployments, and canary release strategies.
Experience working with third-party technology vendors.
Previous experience mentoring engineers or leading platform engineering initiatives.
Key Technology Stack
Cloud: AWS
Containers: Docker, ECS Fargate
Infrastructure as Code: Terraform
CI/CD: GitHub Actions
Monitoring & Observability: CloudWatch, OpenTelemetry, SigNoz, Grafana
Security: AWS Secrets Manager, RBAC, ABAC, Zero Trust
Automation: Python, Bash, PowerShell
Architecture: Microservices, REST APIs, Distributed Systems, Event-Driven Architecture
Practices: DevOps, SRE, Agile, Well-Architected Framework
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.