ClickHouse
Posted 4mo ago

Database Reliability Engineer - Core Team

ClickHouse
Netherlands or United Kingdom or Germany
RemoteFull Time
Responsibilities
  • improving reliability
  • managing incidents
  • creating metrics
Requirements
  • 5+ years reliability or related engineering experience
  • Degree in computer science or related field
  • Production experience operating ClickHouse or other SQL databases
  • Scripting with Shell/Python
  • Familiarity reading C++ code
  • Cloud experience (AWS/Azure/GCP)
  • Strong debugging and communication skills
Technical tools mentioned
ClickHouseSQLShellPythonC++AWSAzureGoogle Cloud Platform

Job description

Note: This position can be based remotely in the Netherlands, UK, or Germany.

We are committed to providing our customers with reliable and secure services at ClickHouse. To continue this, we are building out our Site Reliability Engineering team in ClickHouse Core. As one of the first members of our Reliability Engineering Team at Core, you will be responsible for building and leading processes to ensure and improve the reliability, availability, scalability, and performance of ClickHouse. You will collaborate with different teams like Control Plane, Dataplane,Security, Support and Operations and guide them to implement ClickHouse in the best way for our customers. You will also own the areas of managing engineering escalation management and response, investigations, post-mortem analysis including running blameless postmortems, and continuous improvement of how Clickhouse is run and optimized in the cloud. This role is a unique opportunity to make a significant impact on our elastic, limitless scale, high-performance ClickHouse in ClickHouse Cloud.

What will you do?

  • Continuously improve the reliability and performance of ClickHouse core.

  • Improve and create metrics and alerts for ClickHouse to be able to identify and prevent problems in production before they affect customers. 

  • Dig deeper into the most common problems encountered by customers in Clickhouse Core to identify the root cause of problems and submit bug fixes, issue reports and suggest improvements.

  • Enhance and refine incident response processes and post-mortem analysis for ClickHouse core related outages including working with support and Cloud teams to communicate to the impacted customers.

  • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities.

  • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize customer impact.

About you:

  • Bachelor’s or Master’s degree in Computer Science or a related field.

  • At least 5 years of experience in Reliability Engineering, QA or customer facing engineering.

  • Previous experience operating ClickHouse or other SQL databases in production. 

  • Excellent understanding of distributed database internals and SQL, particularly ClickHouse is a major plus.

  • Scripting experience with Shell or Python,and ability to read and understand C++ code.

  • Knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.

  • You are a strong problem-solver and have solid production debugging skills.

  • You thrive in a fast-paced environment as part of a global team, and you see yourself as a partner with the business with the shared goal of moving the business forward.

  • You have a high level of responsibility, ownership, and accountability.

  • Excellent communication skills

About ClickHouse

Real-time, column-oriented database management system.

Year founded
2021
Employees
500
Organization type
Private
Latest investment
Raised $400.00M Series D (2026) — led by Dragoneer Investment Group
Subsidiaries
Headquarters
US

Similar jobs

Database Reliability Engineer roles
1mo
Save
Mark Applied
Hide
Database Reliability Engineer
London or United Kingdom
RemoteFull Time
Fuse Energy
Fuse Energy: Vertically integrated renewable energy supplier and decentralized grid developer.
3+ YOE3+ years backend or data engineering experience; strong Python and SQL; hands-on Postgres; DBT familiarity; Infrastructure-as-Code experience (Pulumi, AWS CDK); data validation and testing practices.
Python, SQL, Postgres, ClickHouse, DBT, Pulumi, AWS CDK, Dagster, Step Functions, CI/CD
1mo
Save
Mark Applied
Hide
DBRE SME
Bucharest or Dublin or Ireland or United Kingdom or Europe
RemoteFull Time
Amach
Amach: Provides cloud transformation and technology services for airlines.
5+ YOE5+ years MongoDB administration/DBRE experience including replica sets, sharding, backup/DR, performance tuning, profiling (Performance Advisor), observability/SLOs, and Percona tooling.
MongoDB, Performance Advisor, Percona, MongoDB Atlas
4mo
Save
Mark Applied
Hide
Database Reliability Engineer
Southampton, England, United Kingdom
HybridFull Time
Starling
Starling: Digital bank providing personal and business current accounts.
PostgreSQL & Kubernetes expert; infrastructure as code; distributed systems; security and observability mindset; strong Java backend skills.
PostgreSQL, Kubernetes, CNPG, Terraform, Prometheus, Grafana, OpenTelemetry, Humio, Java
5mo
Save
Mark Applied
Hide
Database Reliability Engineer
Germany
RemoteFull Time
Sporty Group
Sporty Group: Global sports entertainment and digital gaming technology.
4+ YOE4+ years of experience with MySQL and MongoDB; strong automation; production-scale database systems; familiarity with SLO/SLI concepts.
MySQL, MongoDB, Relational Databases, NoSQL Databases, Caching, Message Queue, Search Engine, OLAP, Monitoring, Backup and Recovery, Database Proxy, Infrastructure as Code, CI/CD, Network Security, Programming Languages: Go, Python, Containerization: Kubernetes
2w
Save
Mark Applied
Hide
(Senior) DevOps / Database Reliability Engineer (m/w/d)
Bad Friedrichshall, Baden-Württemberg, Germany
OnsiteFull Time
Schwarz Group
Schwarz Group: International retail group operating Lidl, Kaufland, and cloud services.
Experience administering PostgreSQL, Microsoft SQL and Oracle; Docker and Kubernetes; infrastructure provisioning with Terraform and ArgoCD; backend coding (PHP, Go or Kotlin); degree in computer science or equivalent; fluent German and English.
PostgreSQL, Microsoft SQL, Oracle, Docker, Kubernetes, Terraform, ArgoCD, PHP, Go, Kotlin, StackIT, GCP
2mo
Save
Mark Applied
Hide
Database Site Reliability Engineer
London or Peterborough
HybridFull Time
Compare the Market
Compare the Market: A technology team building and operating reliable data platforms for comparison services.
Experience operating and improving production database/data platform systems; end-to-end service ownership; strong understanding of database design, performance optimisation, monitoring, observability, and incident management; collaborative problem solving.