Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Raiffeisen Bank is the largest Ukrainian bank with foreign capital. For more than 30 years, we have been creating and building the banking system of our country.
Raiffeisen employs more than 5,000 employees, including one of the largest product IT teams, which includes 900+ specialists. Every day, we work side by side so that more than 2.5 million of our clients can receive quality service, use the bank’s products and services, and develop their business, because we are #TogetherWithUkraine.
We are looking for a Middle SRE to join our team to help ensure the stability, scalability, and predictable performance of our services in production. In this role, you will work at the intersection of software development, infrastructure, monitoring, and reliability engineering practices.
Your future responsibilities:
Maintain and evolve production infrastructure
Ensure high availability, reliability, and performance of services
Automate routine operations and minimize manual effort (toil reduction)
Set up and maintain monitoring, alerting, and observability systems
Analyze incidents, actively participate in troubleshooting, and conduct Root Cause Analysis (RCA)
Conduct post-incident reviews and implement preventive measures to avoid recurring issues
Participate in capacity planning, system scaling, and resource utilization optimization
Enhance disaster recovery (DR) processes, backup strategies, and operational readiness
Implement and support a GitOps approach to deployment using Argo CD; manage declarative application configurations in Kubernetes
Maintain comprehensive documentation of standard operating procedures, runbooks, and architectural decisions
Collaborate closely with software development, QA, security, and IT support teams
Participate in on-call rotations according to an agreed schedule
Adopt and drive AIOps practices: leverage ML/AI for anomaly detection across metrics and logs, alert correlation, and automated root-cause hypothesizing
Integrate LLM-powered tools into daily operational workflows: incident diagnostics, reporting, log analysis, and technical knowledge base search
Tech stack:
OS: Linux
Containers & Orchestration: Kubernetes, Docker
Cloud Platform: AWS
Infrastructure as Code (IaC): Terraform
CI/CD: GitHub Actions, Jenkins
Monitoring & Observability: Prometheus, Grafana, Alertmanager, and related tooling
Logging: OpenSearch, Elasticsearch
Scripting & Automation: Python
Version Control & Quality: Git, code review best practices
Your skills and experience:
2+ years of experience in SRE, DevOps, Platform Engineering, or a related role
Solid knowledge of Linux and networking fundamentals: TCP/IP, DNS, HTTP, TLS
Hands-on experience with Kubernetes and containerization
Proven track record of building and maintaining CI/CD pipelines
Experience with Terraform or other Infrastructure as Code (IaC) tools
Strong understanding of observability principles: monitoring, logging, metrics collection, and distributed tracing
Ability to independently diagnose and troubleshoot production issues
Clear understanding of core SRE practices: SLIs, SLOs, SLAs, and error budgets
Proficiency in Python
Ability to read, interpret, and analyze application logs and metrics
Strong and clear communication skills during incident management
Nice to have:
Hands-on experience with or a strong interest in adopting AI/ML-driven approaches within SRE practices
Experience using LLM tools to accelerate diagnostics, script writing, and operational automation
Understanding of AI agent architectures and principles
Experience applying AI for anomaly detection, alert correlation, or incident root cause analysis (RCA)
Experience designing, developing, or deploying AI agents to automate routine SRE workflows and reduce toil
We offer what matters most to you:
Competitive salary: we guarantee a stable income and annual bonuses for your personal contribution. Additionally, we have a referral reward program for attracting new colleagues to Raiffeisen Bank
Social package: official employment, 28 days of paid leave, additional “maternity leave” for fathers, and financial assistance for parents upon the birth of children
Comfortable working conditions: the possibility of a hybrid work format, offices equipped with shelters and generators, provision with modern equipment
Wellbeing program: all employees have access to medical insurance from the first working day; consultations with a psychologist, nutritionist or lawyer; discount program for sports and shopping; family days for children and adults; massage in the office
Learning and development: access to over 130 online educational resources; corporate training programs, online library, mentoring program
A great team: our colleagues are a community where curiosity, talent and innovation are welcomed. We support each other, learn together and grow. You can find like-minded people in over 15 professional communities, reading or sports clubs
Career opportunities: we encourage advancement within the bank between functions
Innovation and technology. Infrastructure: AWS, Kubernetes, Docker, GitHub, GitHub actions, ArgoCD, Prometheus, Victoria, Vault, OpenTelemetry, ElasticSearch, Crossplain, Grafana. Languages: Java (main), Python (data), Go (infra, security), Swift (IOS), Kotlin (Andorid)Datastores: Sql-Oracle, PgSql, MsSql, Sybase. Data management: Kafka, AirFlow, Spark, Flink, we develop expertise in AI and actively integrate it into processes
Support program for defenders: we preserve jobs and pay the average salary to mobilized people. We have a support program for veterans, and the Bank’s veteran community is developing. We are working to raise awareness among managers and teams on the issues of veterans’ return to civilian life. Raiffeisen Bank is recognized as one of the best employers for veterans (Forbes)
Why Raiffeisen Bank?
Our main value is people, and we support and recognize them, educate them and involve them in changes. Join Raif’s team because for us YOU matter!
One of the largest lenders to the economy and agricultural business among private banks
Recognized as the best employer by EY, Forbes, Randstad, Franklin Covey, and Delo.UA
The largest humanitarian aid donor among banks (Ukrainian Red Cross, UNITED24, Superhumans, СМІЛИВІ)
One of the largest IT product teams among the country’s banks. • One of the largest taxpayers in Ukraine; 6.6 billion UAH were paid in taxes in 2023.
Opportunities for Everyone:
Raif is guided by principles that focus on people and their development, with 5,500 employees and more than 2.7 million customers at the center of attention
We support the principles of diversity, equality and inclusiveness
We are open to hiring veterans and people with disabilities and are ready to adapt the work environment to your special needs
We cooperate with students and older people, creating conditions for growth at any career stage
You matter at Raif!
Want to learn more? Follow us on social media:
Facebook, Instagram, LinkedIn
__________________________________________________________________________________________
Райффайзен Банк — найбільший український банк з іноземним капіталом. Більше 30 років ми створюємо та вибудовуємо банківську систему нашої держави.
У Райфі працює понад 5 тисяч співробітників, серед них одна із найбільших продуктових ІТ-команд, що налічує 900+ фахівців. Щодня пліч-о-пліч ми працюємо, щоб більш ніж 2,5 мільйона наших клієнтів могли отримати якісне обслуговування, користуватися продуктами і сервісами банку, розвивати бізнес, адже ми #Разом_з_Україною.
Шукаємо Middle SRE до нашої команди, який допоможе забезпечувати стабільність, масштабованість і передбачувану роботу сервісів у production. На цій позиції ви працюватимете на перетині розробки, інфраструктури, моніторингу та практик reliability engineering.
Твої майбутні обов’язки:
Підтримувати та розвивати production-інфраструктуру
Забезпечувати високу доступність і стабільність сервісів
Автоматизувати рутинні операції та зменшувати кількість ручної роботи
Налаштовувати й підтримувати моніторинг, алертинг та observability
Аналізувати інциденти, брати участь у troubleshooting і root cause analysis (RCA)
Проводити post-incident review та допомагати запобігати повторенню проблем.
Брати участь у capacity plannin
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.