Live opening · Posted 17 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Company Description
We are looking for an experienced SRE Guild Lead to own the practice, standards, and craft of Site Reliability Engineering across our Factory pods. As guild lead, the person will be responsible for the technical and cultural anchor for reliability at TCG Digital – setting the bar for how we define, measure, and defend service reliability, and ensuring every pod applies consistent SRE practices regardless of which product or client engagement they sit under. The role will work closely with the DevOps & Infra guild, Architecture & Design guild, and pod leads to embed observability, automation, and operational excellence into how we build and run systems. To succeed in this role, the person should have hands-on production engineering depth combined with the ability to mentor and set standards across a distributed group of engineers who do not report to the role directly. This is a guild leadership role, not a pure people management role: influence, technical credibility, documentation, and cross pod coordination are your primary tools. The person will also be the escalation point for major incidents and a key voice in capacity planning for reliability roles across the delivery organization.
Scope of work
6+ years working in Information Technology, with 4+ years in a Site Reliability Engineering or production operations role supporting large-scale, customer-facing systems.
• Own service reliability end to end: define and maintain SLIs, SLOs and error budgets, establish steady-state behaviour for critical user journeys, and drive the reliability backlog that comes out of it.
• Performance and load engineering with k6 – author JavaScript/TypeScript test scripts, build scenarios and
executors (constant-VUs, ramping-VUs, ramping-arrival- rate), define custom metrics, checks and thresholds, parameterise test data, and cover HTTP, WebSocket and gRPC protocols.
• Integrate k6 into CI/CD pipelines (Harness, GitHub Actions, Jenkins) as automated performance regression gates; run distributed and containerised tests via the k6 Operator on Kubernetes or Grafana Cloud k6; stream results to Prometheus, InfluxDB, Grafana or Dynatrace for trend analysis.
• Chaos engineering – design and run hypothesis-driven experiments with an explicit steady state, controlled blast radius, abort conditions and rollback plan; plan and facilitate GameDays with application, infrastructure and client teams; convert every finding into a tracked remediation item.
• Fault injection using AWS Fault Injection Service (FIS), Gremlin, Chaos Mesh, LitmusChaos or Chaos Toolkit – CPU and memory pressure, pod and node termination, network latency and packet loss, dependency and third-party API failure, AZ evacuation, and database failover.
• Dynatrace administration and engineering – OneAgent deployment and lifecycle, management zones, entity model and tagging strategy, alerting profiles, SLO definitions, dashboards and notebooks, and day-to-day administration of tooling in the APM space.
• Dynatrace Davis AI and the Davis / Dynatrace API v2 – programmatic access to the problems, metrics, events, entities and SLO endpoints; ingest custom metrics and deployment / chaos events; tune Davis anomaly detection, alerting sensitivity and root-cause behaviour; correlate load tests and chaos experiments with Davis-detected problems to validate detection and MTTR.
• Automate release validation and quality gates using Dynatrace Site Reliability Guardian / Cloud Automation (or equivalent) so that performance and resilience evidence is evaluated automatically as part of every deployment.
• Run production workloads on AWS – EKS (node groups, clusters, autoscaling), EC2, ALB/NLB, Route 53, RDS, Lambda and other AWS native services, with a strong grasp of container monitoring best practices.
• Infrastructure and observability as code with Terraform; scripting in Python, Go, Bash and JavaScript/TypeScript for automation, API integration and internal tooling.
• Participate in the on-call rotation; drive incident triage, root cause analysis and corrective actions, and work with cross-functional teams and Problem Management on escalations.
• Manage uptime and availability reporting, capacity planning and performance tuning, using both load-test evidence and production telemetry.
REQUIRED QUALIFICATIONS - KNOWLEDGE/SKILLS
• Demonstrable hands-on experience building and maintaining a k6 performance testing suite, not just running someone else’s scripts.
• Practical chaos engineering experience in a production or production-like environment, with evidence of the reliability defects it surfaced.
• Working knowledge of the Dynatrace API v2 and Davis AI, including authentication and token scopes, rate limits, and consuming problem and metric data from automation.
• Ability to define alert standards for production environments and implement them, including tuning to
reduce false positives and alert fatigue.
• Strong understanding of distributed systems, networking and troubleshooting techniques – latency, timeouts, retries, connection pooling, DNS, TLS and load balancing.
• Experience with automated build pipelines and continuous integration. Source control, branching and merging: git/svn/etc (Repository Management).
• Familiarity with configuration management software and observability standards such as OpenTelemetry.
• Provide support to teams for alarms and outages on an as- needed basis, and work with development teams and management to ensure high availability.
• Communication Skills- The ability to communicate verbally and in writing with all levels of employees and management, speaks and writes clearly and understandably at the right level.
• Integrity and Trust- Involves being widely trusted, being seen as a direct, truthful individual, can present the unvarnished truth in an appropriate and helpful manner, keeps confidences, admits mistakes, and doesn’t misrepresent him/herself for personal gain.
• Teamwork- Works well in a collaborative setting, volunteering for and completing assignments, acting as a positive team member by contributing to discussions.
Skills Required
Essential Skills SRE practices (SLI/SLO/error budgets, incident and problem management); k6 performance and load engineering; Chaos Engineering; Dynatrace including Davis AI and the Davis /
Dynatrace API v2. Supp orting stack: AWS (EKS, EC2, native services), Kubernetes, Terraform and CI/CD pipeline automation. Refer to the scope of work for the detail.
Desired Skills
• Grafana, Prometheus and OpenTelemetry; experience with other load testing tools (JMeter, Gatling, Locust) in addition to k6.
• Harness for CI/CD pipeline automation, deployment orchestration and release management; progressive delivery patterns (canary, blue-green, feature flags).
• Certifications such as Dynatrace Associate/Professional, Certified Kubernetes Administrator (CKA), or AWS Solutions Architect / DevOps Engineer.
• Hands-on experience and familiarity with AI-assisted “vibe coding” using tools such as GitHub Copilot, Claude Code, Cursor, or similar AI development platforms is preferred. Given the evolving nature of these technologies, practical exposure and a clear understanding of AI-assisted development concepts are acceptable.
Other Information
Education Qualification
B.E. / B.Tech / MCA / M.Sc. in Computer Science,
Information Technology or an equivalent discipline
Experience 6+ Years
*Experience Level
(Must Have Skills)
• 4+ years in SRE / production operations, including on-call ownership
• 2+ years hands-on performance engineering with k6
• 2+ years with Dynatrace, including Davis AI and API-based automation
• 1+ year running chaos engineering experiments or GameDays
• 3+ years running production systems on AWS (EKS/EC2)
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.