Live opening · Posted 13 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Role Description The Analyst – Data Center Operations role is a full-time, on-site position based in Bengaluru South. This role involves monitoring data center infrastructure, systems, and networks to ensure high availability, performance, and security. The analyst will respond to incidents, perform root-cause analysis, and coordinate with technical teams to resolve issues and minimize downtime. Daily tasks include tracking operational metrics, maintaining documentation, supporting audits and compliance activities, and following standard operating procedures. The role also requires collaborating with cross-functional teams, providing status updates, and contributing to continuous improvement of data center processes and controls.
Qualifications
Incident Monitoring & Response
Continuously monitor the Pantomath Operation Center dashboard for pipeline failures, stale or missing data, and anomalies across the enterprise data stack
Perform first-line triage using Pantomath’s automated root-cause analysis and cross-platform lineage tracing to identify affected pipelines and downstream impact
Classify and prioritize incidents by business severity and follow defined escalation playbooks
Drive incidents from detection through resolution, escalating to engineering or platform teams when root cause requires deeper remediation
Role and Responsibilities:
Administer and maintain the Pantomath platform or similar platform including monitor configuration, alert thresholds, connectors, and lineage coverage
Onboard new pipelines and data sources into observability coverage as the data estate grows
Tune monitors and reduce alert noise by refining thresholds based on historical incident patterns
Maintain platform health, user access, and integration with upstream and downstream enterprise systems
Technical Skills
Hands-on experience with Pantomath or a comparable data observability platform (Monte Carlo, Bigeye, or similar) strongly preferred
Familiarity with cloud and lakehouse data platforms such as Databricks, Snowflake, Azure, or AWS
Proficiency in SQL for incident investigation and root cause analysis
Experience with incident management and ticketing tools (ServiceNow, Jira, or similar)
Communication & Stakeholder Management
Communicate outage status, root cause, and resolution timelines clearly to business stakeholders and data consumers across the enterprise
Draft and distribute incident notifications and status updates through approved communication channels
Maintain incident logs and root cause analysis (RCA) documentation, and contribute to post-incident reviews
Qualifications:
Bachelor’s degree in Computer Science, Information Systems, Data Science, or related field, or equivalent practical experience
2–5 years of experience in data engineering, data operations, IT operations, or a related technical support role
Demonstrated experience monitoring and triaging incidents in a production data or IT environment
Working Model and Expectation:
· Shift coverage aligned to U.S. business hours, with rotation to extend the monitoring window as the DOC matures
· Adherence to defined incident severity matrices and communication service-level agreements
· Active participation in shift handoff, daily DOC stand-up, and weekly RCA review
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.