Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.
Responsibilities
Provide on-call support for Java backend identity services during business hours
Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution
Prepare and deploy patches to address issues in cloud infrastructure
Improve service reliability by implementing practical changes that reduce errors and instability
Build and refine metrics and dashboards to surface platform health and service behavior
Monitor SLOs and propose code changes that improve SLO attainment as issues arise
Create and improve runbooks to standardize operational response and reduce time to recovery
Communicate incidents and operational risks clearly in writing during live response
Collaborate closely with engineers to align operational practices with service ownership
Requirements
3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems
Strong on-call support experience for production services and incident response during business hours
Proven experience with Amazon Web Services in production environments
Hands-on experience with Amazon DynamoDB and Amazon ElastiCache
Strong Git skills for collaborating on operational and reliability code changes
Solid Gradle knowledge for building and maintaining Java-based services
Strong troubleshooting skills using logs and telemetry to identify root causes
Clear written communication skills for documenting and reporting operational issues during incidents
Proactive learning mindset to absorb complex information quickly and apply it under pressure
Upper-Intermediate English proficiency (B2)
Nice to have
Kubernetes
Terraform
Grafana
New Relic
Apache Kafka
We offer
International projects with top brands
Work with global teams of highly skilled, diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling, reskilling and certification courses
Unlimited access to the LinkedIn Learning library and 22,000+ courses
Global career opportunities
Volunteer and community involvement opportunities
EPAM Employee Groups
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.