Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Position: Sr SRE
Location: LATAM Remote
Shift: 5 PM – 1 AM EST, Monday–Friday
Experience: 8+ years relevant SRE/Production Operations experience
Openings: 2
Role Overview
We are looking for Senior Site Reliability Engineers to support mission-critical enterprise applications in a high-volume production environment. The role focuses on incident response, production troubleshooting, cloud infrastructure, Kubernetes, automation, and reliability engineering.
Key Responsibilities
• Act as first responder for production alerts and P1/P2 incidents
• Participate in incident bridge calls and serve as Incident Commander when required
• Troubleshoot and resolve critical production issues; drive root-cause isolation within 30 minutes whenever possible
• Clearly present issues, impact, root cause, and solutions to engineering and leadership teams during incidents
• Perform proactive monitoring, alert tuning, reliability improvements, and toil reduction
• Automate operational tasks, ticket handling, and repetitive workflows
• Troubleshoot applications across cloud, on-prem, Kubernetes, networking, CDN, and databases
Must-Have Technical Skills
AWS / Cloud
Hands-on AWS: S3, Lambda, EC2, ECS, Load Balancers
Experience with GCP and/or Azure
Multi-cloud & hybrid environments
Kubernetes / Containers
Kubernetes cluster troubleshooting
Pods, Ingress, environment/configuration issues
Containerized application architecture
Incident Management
Production incident management & triage
P1/P2 incident response
Incident Commander / bridge-call experience
Troubleshooting distributed systems under pressure
Strong communication with leadership during incidents
Application Deployment & Troubleshooting
Hands-on deployment and troubleshooting of:
Java
Node.js
React-based applications
Experience supporting applications running in cloud environments
DevOps / CI/CD
Strong DevOps & CI/CD knowledge
Harness
GitHub and/or GitLab
Automation / scripting / operational tooling
Linux / Systems
Strong Linux administration
Performance troubleshooting
Networking and NTP
On-prem Linux/Windows VMs & VMware
Networking
TCP/IP
DNS
HTTP/HTTPS
Load balancing
Service-to-service & database connectivity troubleshooting
Databases
Troubleshooting connectivity/dependencies involving:
Oracle
MariaDB
MSSQL
DBA ownership is not required
CDN / Traffic Management
Akamai CDN experience is a strong plus
Traffic management / waiting rooms
Automation / AI
Automation mindset and toil reduction
AI-assisted development tools such as Cursor / Claude
Preferred
Fortune 500 / large-scale enterprise production support
Mission-critical, high-availability environments
Previous Incident Commander experience
Akamai experience
AWS / Terraform certifications
Top 3 Experiences Required
AWS: S3, Lambda, EC2, ECS & Load Balancers
P1/P2 Incident Response: First responder, bridge calls, troubleshooting, Incident Command, and communicating solutions to leadership
Java/Node.js/React: Cloud deployment and production troubleshooting
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.