Doctolib
Posted 6mo ago

Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m)

Doctolib
Berlin, Berlin, Germany
OnsiteFull Time
Responsibilities
  • leading team
  • defining strategy
  • ensuring reliability
Requirements
  • 5+ years software engineering/SRE experience
  • 3+ years engineering management
  • Observability tooling expertise
  • Experience with cloud-native environments, and people leadership skills
Technical tools mentioned
RailsTypeScriptJavaPythonKotlinSwiftReact NativeFluent BitOpenTelemetryLokiElasticsearchPrometheusThanosDatadogTerraformOpenTofuTerraform EnterpriseHashiCorp VaultAWS Secrets ManagerGoRubyKubernetesAWSGCP

Job description

We are looking for an Engineering Manager to join the OREO (Observability Reliability Engineering Obsession) team in Platform Engineering.

As an Engineering Manager, your mission will be to lead the Reliability & Observability team and drive the evolution of Doctolib's observability platform, supporting the exponential growth of Doctolib services while building and empowering a world-class SRE team.

Working in the tech team at Doctolib involves building innovative products and features to improve the daily lives of care teams and patients. We work in feature teams in an agile environment, while collaborating with product, design, and business teams.

You will lead a team of Site Reliability Engineers who are responsible for shaping Doctolib's observability strategy and ensuring our platform remains reliable, debuggable, and scalable. This role sits at the intersection of people management, technical leadership, and strategic planning with a particular focus on building organizational capabilities around logging, metrics, tracing, and alerting.

Your team also owns and operates critical transversal services that enable secure, scalable infrastructure management across the organization, including HashiCorp Vault for secrets management and Terraform Enterprise for infrastructure as code.

Your responsibilities include but are not limited to:

 

People Leadership:

  • Lead, coach, and grow a team of Site Reliability Engineers, supporting their technical development and career progression

  • Create a culture of operational excellence, continuous improvement, and psychological safety within the team

  • Conduct regular 1:1s, performance reviews, and career development conversations

  • Recruit, onboard, and retain top SRE talent aligned with Doctolib's mission and values

Technical Strategy:

  • Partner with SREs and senior engineers to define and evolve the observability strategy across the platform, focusing on logging, metrics, tracing, and alerting

  • Own the strategy and evolution of critical transversal services including HashiCorp Vault and Terraform Enterprise

  • Drive prioritization and roadmap planning for large-scale reliability and observability initiatives

  • Ensure alignment between team objectives and broader engineering and business goals

  • Advocate for and allocate resources toward reducing technical debt and improving developer experience

Operational Excellence:

  • Own the team's on-call experience and contribute to the incident response processes, ensuring sustainable practices and continuous improvement

  • Ensure high availability and reliability of transversal services that are critical to the entire engineering organization

  • Lead postmortem reviews and drive systemic improvements to prevent recurring issues

Cross-functional Collaboration:

  • Work closely with Product Managers, Engineering Managers, and architects to align observability capabilities with product and platform needs

  • Partner with security and infrastructure teams to evolve secrets management and IaC practices across the organization

  • Represent the OREO team in engineering leadership forums, architectural reviews, and strategic planning sessions

  • Foster strong partnerships with software engineering teams to improve instrumentation quality and adoption of observability best practices

About our tech environment

  • Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to the country and healthcare specialty requirements. To address these challenges, we are modularizing our platform run in a distributed architecture through reusable components.

  • Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native.

  • We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here and learn about our first AI hackathon here!

Who you are

Before you read on — if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.

  • You have at least 5+ years of software engineering or SRE experience, with a strong technical background in cloud-native environments (preferably AWS, GCP, and/or Kubernetes-based)

  • You have 3+ years of engineering management experience, leading technical teams (ideally SRE, platform, or infrastructure teams)

  • You have deep understanding of observability tooling and architecture (Fluent Bit, OpenTelemetry, Loki, Elasticsearch, Prometheus, Thanos, Datadog)

  • You have experience with infrastructure as code (Terraform, OpenTofu) and secrets management systems (Vault, AWS Secrets Manager)

  • You have proven ability to balance technical depth with people leadership, able to mentor engineers, review technical designs, and guide architectural decisions

Now it would be fantastic if you:

  • Have experience scaling SRE or platform teams in fast-growing, high-traffic environments

  • Have background in designing and operating high-scale telemetry pipelines

  • Have hands-on experience with HashiCorp Vault and Terraform Enterprise in production environments

  • Have hands-on experience with backend programming languages (e.g., Go, Python, Ruby)

  • Have experience driving cultural and technical transformations

What we offer

  • Free comprehensive health insurance for you and your children

  • Parent Care Program: receive one additional month of leave on top of the legal parental leave

  • Free mental health and coaching services through our partner Moka.care

  • For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support

  • Work from EU countries and the UK for up to 10 days per year, thanks to our flexibility days policy

  • Work Council subsidy to refund part of sport club membership or creative class

  • Up to 14 days of RTT

  • A subsidy from the work council to refund part of the membership to a sport club or a creative class

  • Lunch voucher with Swile card

The interview process

  • 30-min phone screen with a Tech Recruiter

  • 1h30 technical interview (SRE System Design & Architecture)

  • 1h15 behavioral interview (Leadership & People Management)

  • 1h30 Engineering Management case study (team scenarios, prioritization, and conflict resolution)

  • 1h manager interview with Senior Engineering Leadership

  • At least one reference check

Job details

  • Permanent position

  • Full-time

  • Berlin

  • Start date: as soon as possible

If you would like to find out more about tech life at Doctolib, feel free to read our latest Medium blog articles!

At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination.

The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, disability.

To ensure equal opportunities, we invite you to exclude personal information (e.g. pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process.

Join us in building the healthcare we all dream of!

All information provided is processed by Doctolib for application management. For data processing details, click here.

Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights.

About Doctolib

Digital health platform for appointment booking and practice management.

Year founded
2013
Employees
3500
Organization type
Private
Latest investment
Raised $549.00M Series G (2022) — led by Bpifrance, General Atlantic, Eurazeo
Subsidiaries
Headquarters
FR

Similar jobs

Engineering Manager roles near Berlin, Berlin
17h
Save
Mark Applied
Hide
Engineering Manager - AI Engineering & AI Platform (m/f/x)
Berlin, Berlin, Germany
HybridFull Time
Scalable Capital
Scalable Capital: Digital investment platform for self-directed trading and wealth management.
Requires engineering leadership experience, scalable backend and cloud architecture knowledge, AI system integration expertise, and hands-on Python, Docker, CI/CD, cloud infrastructure, and Infrastructure as Code experience.
Python, Docker, CI/CD, AWS, Terraform, LangGraph, LiteLLM, Langfuse, Datadog, MCP servers
18h
Save
Mark Applied
Hide
Engineering Manager, Intelligent Platforms (all genders)
Berlin, Berlin, Germany
HybridFull Time
HelloFresh
HelloFreshFrankfurt Stock Exchange: HFG: Global meal kit delivery service and food solutions provider.
Experience managing distributed engineering teams, cloud infrastructure expertise, Kubernetes and major cloud platform experience, production ownership, incident response, stakeholder management, and strong written communication.
Kubernetes, service mesh, API gateways, infrastructure as code, GitOps, secrets management, Headspace, Spill, Urban Sports Club, HelloFresh Academy
21h
Save
Mark Applied
Hide
Engineering Manager (m/f/d) - Berlin
Berlin, Berlin, Germany
OnsiteFull Time
FREENOW
FREENOW: Multi-modal platform for urban taxi and transport services.
5+ YOE2+ MgmtBachelor's degree preferred in computer science or a related technical field; 5+ years in software engineering and 2+ years in leadership or management, with software delivery, agile, coaching, and stakeholder management expertise.
Java, Kotlin, AWS, Docker, Git, Kafka, Elasticsearch
1d
Save
Mark Applied
Hide
Senior Engineering Manager (m/f/d)
Berlin, Berlin, Germany
HybridFull Time
Affinidi
Affinidi: Provides decentralized digital identity and data ownership solutions.
Experience leading engineering organizations and high-performing teams; TypeScript/Node.js, React/Next.js or Flutter, AWS, microservices, cloud systems, and technical and people leadership expertise required.
TypeScript, Node.js, Dart, C#, Go, React, Next.js, Flutter, Amazon Web Services (AWS), Rust
2d
Save
Mark Applied
Hide
Engineering Manager Ads Backend
Berlin, Berlin, Germany
HybridFull Time
Zattoo
Zattoo: Provider of live TV streaming and IPTV infrastructure services.
5+ MgmtRequires 5+ years leading engineering teams, strong hands-on engineering and technical leadership, C++/Go, scalable services, distributed systems, event-driven architectures, message queues, databases, and fluent English.
C++, Go, Kafka, Scylla
2d
Save
Mark Applied
Hide
Senior Engineering Manager, Central Data Products
Berlin, Berlin, Germany
HybridFull Time
GetYourGuide
GetYourGuide: Online marketplace for booking tours and travel activities.
Strong machine learning and engineering background; experience with data architecture, evaluation methods, production pipelines, team development, stakeholder alignment, and excellent English communication.
LLMs, MLOps
3d
Save
Mark Applied
Hide
Engineering Manager, FDE Agentic Platform
Toronto or New York City or London or San Francisco or Montreal or Paris or Berlin or Seoul
RemoteFull Time
Cohere
Cohere: Provides enterprise-grade large language models and AI software platforms.
8+ YOE2+ MgmtRequires 8+ years in software engineering, 2+ years managing engineering teams, enterprise software deployment at scale, agentic applications experience, high agency, execution, and excellent English communication.
North
4d
Save
Mark Applied
Hide
Senior Engineering Manager, Ad Performance Recommendation
Berlin, Berlin, Germany
HybridFull Time
Smartly
Smartly: Automates social media advertising and creative production for brands.
7+ YOE3+ MgmtRequires 7+ years building services and products, 5+ years applying machine learning, 3+ years leading teams, ML expertise with PyTorch, TensorFlow, Python, and SQL, and an M.Sc. or Ph.D. in a relevant field.
PyTorch, TensorFlow, Python, SQL, Stan, Meta, Pinterest, Snap, TikTok, Google, Facebook, Instagram