Live opening · Posted 9 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About Position:
We are looking for an experienced Lead ClickHouse Database Administrator to own and operate our ClickHouse analytics database platform that powers large-scale, real-time network analytics solutions. This role is responsible for ensuring the availability, scalability, performance, and security of distributed ClickHouse environments deployed across Kubernetes-based infrastructure. he ideal candidate will have deep expertise in ClickHouse administration, Kubernetes operations, Kafka integrations, database observability, backup and recovery, infrastructure automation, and large-scale analytics platform management. You will work closely with Engineering, Product Management, DevOps, and Support teams to deliver highly reliable, geo-redundant analytics platforms that support mission-critical business operations.
Role: Lead ClickHouse Database Administrator
Location: All Persistent Locations
Experience: 10 to 15 Years
Job Type: Full-Time Employment
What You'll Do:
Own and execute ClickHouse version upgrades, including compatibility validation, rolling deployments, rollback planning, and post-upgrade verification.
Administer and optimize ClickHouse cluster topology, shard and replica configurations, distributed DDL operations, and node management.
Lead migration initiatives from ZooKeeper to ClickHouse Keeper, including planning, execution, validation, and risk mitigation.
Manage ClickHouse maintenance activities such as merges, mutations, storage optimization, partition management, and TTL administration.
Design, implement, and validate backup and disaster recovery strategies using ClickHouse backup tools and S3-compatible storage solutions.
Administer database security controls including RBAC, authentication, TLS configuration, certificate management, user provisioning, and audit logging.
Modernize aggregation architectures and optimize analytical workloads using advanced ClickHouse engine capabilities.
Manage WAN replication between geographically distributed ClickHouse clusters and ensure replication consistency.
Build and maintain observability dashboards using Prometheus and Grafana to monitor platform health, performance, and reliability.
Troubleshoot complex production incidents using ClickHouse logs, system tables, replication metrics, and query profiling tools.
Support Kubernetes infrastructure operations including StatefulSets, storage management, Helm upgrades, and workload optimization.
Manage CI/CD pipelines, deployment automation, integration testing, and release processes.
Diagnose and resolve Kafka-to-ClickHouse ingestion issues, including consumer lag, schema mismatches, serialization failures, and data consistency concerns.
Collaborate with engineering teams to improve database reliability, operational excellence, and infrastructure automation.
Develop operational runbooks, architecture documentation, troubleshooting guides, and database standards.
Provide advanced L3 production support for critical platform incidents and customer escalations.
Expertise You'll Bring:
10 to 15 years of experience in database administration, data platform engineering, or database reliability engineering.
Strong expertise in ClickHouse internals, MergeTree engines, columnar storage architecture, partition management, and lifecycle operations.
Hands-on experience managing large-scale ClickHouse clusters with sharding, replication, distributed query processing, and high-availability architectures.
Extensive experience with ClickHouse upgrades, compatibility assessments, rollback planning, and change management.
Strong knowledge of ClickHouse Keeper deployment, migration, configuration, monitoring, and operational support.
Experience implementing backup, recovery, disaster recovery, and business continuity solutions.
Expertise in SQL optimization, analytical query tuning, execution-plan analysis, and performance troubleshooting.
Strong understanding of ClickHouse security controls, RBAC, TLS encryption, authentication mechanisms, and audit logging.
Strong experience managing stateful database workloads on Kubernetes platforms.
Expertise in StatefulSets, Persistent Volumes, storage classes, storage migrations, and resource optimization.
Hands-on experience deploying and managing applications through Helm charts.
Experience with Kubernetes upgrades, infrastructure operations, and distributed platform management.
Knowledge of certificate automation tools such as cert-manager and related security infrastructure.
Expertise with Prometheus, Grafana, and enterprise monitoring solutions.
Experience building observability frameworks for database performance and platform monitoring.
Ability to diagnose issues using ClickHouse system tables, replication queues, merge operations, and query logs.
Familiarity with Loki or similar logging and log aggregation platforms.
Strong root-cause analysis and troubleshooting skills within distributed environments.
Strong understanding of Kafka architecture including topics, partitions, consumer groups, offsets, and retention policies.
Experience troubleshooting Kafka-to-ClickHouse ingestion pipelines and streaming architectures.
Ability to identify ingestion bottlenecks, replication issues, and data consistency problems.
Knowledge of schema evolution, serialization formats, dead-letter queue handling, and offset recovery mechanisms.
Experience monitoring Kafka environments using observability and performance management tools.
Strong Bash scripting and automation experience.
Familiarity with Java for client compatibility analysis and platform integration support.
Exposure to Go programming and ClickHouse Go drivers is desirable.
Experience with Git, CI/CD pipelines, infrastructure automation, and DevOps practices.
Strong documentation, knowledge-sharing, and operational governance capabilities.
Excellent analytical, troubleshooting, and problem-solving skills.
Ability to independently manage complex production environments and critical escalations.
Strong stakeholder management and cross-functional collaboration skills.
Experience working with globally distributed engineering, DevOps, product, and support teams.
Strong communication skills with the ability to present technical solutions to leadership and customers.
Benefits:
Competitive salary and benefits package
Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
Opportunity to work with cutting-edge technologies
Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
Annual health check-ups
Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:
Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.
We support hybrid work and flexible hours to fit diverse lifestyles.
Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment
Let's unleash your full potential at Persistent - persistent.com/careers
"Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind."
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.