Live opening · Posted 7 hours ago

Senior Staff Engineer - Site Reliability

Freshworks · Hyderabad, TS, India
Smartrecruiters No Full-time
You are 7 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 hours ago
CompanyFreshworks
LocationHyderabad, TS, India
Job typeFull-time
Work modeNo
SourceSmartrecruiters
Listed7 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
22 min from Smartrecruiters publishing this role to us finding it
3 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,165 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

● Design, write, and deliver software to improve the availability, latency, and efficiency of Freshworks’ Products & Platforms.
● Develop scalable, cloud-native architectures that support business growth.
● Design and implement self-healing and auto-scaling mechanisms.
● Manage availability, latency and performance of mission critical services and build automation to prevent problem recurrence.
● Independently determine and develop architectural approaches and Infrastructure solutions.
● Define strategy, vision, and roadmap adhering to well architected principles of performance, availability, scalability, performance and resilience
● Experience with AI-driven system optimization or predictive analytics for IT operations.
● Drive blameless postmortems for large scale incidents.
● Define and drive automation and orchestration strategies.
● Strategize cost optimization across Freshworks Cloud environment.
10-15 years of experience in SRE handling performance, architecture and design of applications
● Strong understanding of cloud computing, networking, Linux systems administration, containerization (e.g., Docker, Kubernetes), and infrastructure as code (e.g., Terraform, Ansible)
● Understanding of SRE principles, including SLOs, SLIs, SLAs, and error budgets.
● Experience in managing incident and retrospectives
● Experience in cloud cost management, cloud architecture
● In-depth knowledge of cloud computing platforms (e.g., AWS)
● Experience with infrastructure as code (IaC) tools and practices
● Experience with monitoring, logging & telemetry tools like New Relic, Splunk, ELK, Nagios, SolarWinds, Prometheus, AWS Cloudwatch, Datadog, Opentelemetry
● Expert in designing, creating and supporting Automation and Identify opportunities for self-healing systems, automated deployments, and other scalable solutions.
● Experience in performance engineering and identify opportunities for performance tuning and profiling
● Experience in prioritizing and managing technical roadmaps.
● Strong skills in stakeholder communication, requirements gathering, and documentation.
● Ability to work with cross-functional teams and build consensus around reliability goals
● Improve operational processes and team practices
● Problem-solving: Ability to analyze complex systems, troubleshoot issues, and devise effective solutions
● Excellent communication skills, with the ability to inspire and motivate cross-functional teams.
● Experience in dealing with the intricacies of large-scale distributed systems and ensuring their reliability and performance.

Employment type
Full-time

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App