Man Group
Posted 3mo ago

Site Reliability Engineer

Man Group
Sofia, Sofia, Bulgaria
HybridFull Time
Responsibilities
  • ensure reliability
  • design observability
  • automate tasks
Requirements
  • Strong SRE knowledge with observability
  • Automation, and incident management
  • Hands-on with Prometheus/Grafana/ELK/Loki
  • Ansible/Terraform
  • Python/Go/PowerShell
  • Kubernetes experience
  • On-call and incident response
Technical tools mentioned
PrometheusGrafanaELKLokiAnsibleTerraformPythonGoPowerShellKubernetes

Job description

About Man Group

Man Group is a global alternative investment management firm focused on pursuing outperformance for sophisticated clients via our Systematic, Discretionary and Solutions offerings. Powered by talent and advanced technology, our single and multi-manager investment strategies are underpinned by deep research and span public and private markets, across all major asset classes, with a significant focus on alternatives. Man Group takes a partnership approach to working with clients, establishing deep connections and creating tailored solutions to meet their investment goals and those of the millions of retirees and savers they represent.

Headquartered in London, we manage $227.6 billion* and operate across multiple offices globally. Man Group plc is listed on the London Stock Exchange under the ticker EMG.LN and is a constituent of the FTSE 250 Index. Further information be found at www.man.com

* As at 31 December 2025

The Role

Join our high-performing Site Reliability Engineering (SRE) team and play a pivotal role in ensuring the reliability, scalability, and performance of the technology powering Man Group’s hedge funds. You’ll have the autonomy, tools, and support to innovate and shape the future of our platform. This is an opportunity to work on cutting-edge projects, gain mentorship from senior leaders, and develop a deep understanding of both technology and the business.

As an SRE, you’ll take ownership of service reliability and deliver solutions that make a real impact. Your initial focus will include leveraging AI to accelerate incident diagnosis and resolution, improving observability, capacity planning, and automation. Over time, you’ll work across our entire infrastructure stack, operating at scale and driving continuous improvement.

Role Responsibilities

  • Ensure reliability and performance of critical systems across global infrastructure through proactive monitoring and rapid incident response.
  • Design and implement observability solutions using tools like Prometheus, Grafana, ELK, and Loki to provide deep insights into system health.
  • Automate operational tasks and build self-service capabilities to eliminate toil and improve efficiency.
  • Develop and maintain SLIs, SLOs, and error budgets to guide reliability improvements and inform engineering priorities.

Key competencies

Required

  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and reliability best practices.
  • Hands-on experience with observability and monitoring tools (Prometheus, Grafana, ELK, Loki, or similar).
  • Proficiency with automation tools (Ansible, Terraform) and scripting/programming languages (Python, Go, PowerShell).
  • Strong troubleshooting and debugging skills across distributed systems, with the ability to diagnose complex production issues under pressure.
  • Experience with incident management, on-call rotations, and post-incident reviews.
  • Familiarity with Kubernetes

Advantageous

  • Experience with CI/CD pipelines and source control workflows (Git, Jenkins, TeamCity).
  • Administration of Linux and Windows systems and exposure to cloud technologies (AWS/Azure).
  • Understanding of networking concepts, load balancing, and distributed architectures.
  • Knowledge of AI/LLM concepts (context windows, prompt tuning, MCP servers).

Benefits

  • Modern office located in the OfficeX campus with easy access to transport and amenities.
  • Hybrid working model
  • Competitive compensation package
  • 25 days holiday allowance
  • Premium Health insurance
  • Employee Assistance program
  • Referral Bonus
  • Additional days off for long service and volunteering
  • Multisport card
  • Opportunities for professional development including internal tech talks
  • Conference attendance, and engagement with the open-source community.

Inclusion, Work-Life Balance and Benefits at Man Group
You'll thrive in our working environment that champions equality of opportunity. Your unique perspective will contribute to our success, joining a workplace where inclusion is fundamental and deeply embedded in our culture and values. Through our external and internal initiatives, partnerships and programmes, you'll find opportunities to grow, develop your talents, and help foster an inclusive environment for all across our firm and industry. Learn more at www.man.com/diversity.
You'll have opportunities to make a difference through our charitable and global initiatives, while advancing your career through professional development, and with flexible working arrangements available too. Like all our people, you'll receive two annual 'Mankind' days of paid leave for community volunteering.

Our comprehensive benefits package includes competitive holiday entitlements, pension/401k, life and long-term disability coverage, group sick pay, enhanced parental leave and long-service leave. Depending on your location, you may also enjoy additional benefits such as private medical coverage, discounted gym membership options and pet insurance.

Equal Employment Opportunity Policy

Man Group provides equal employment opportunities to all applicants and all employees without regard to race, color, creed, national origin, ancestry, religion, disability, sex, gender identity and expression, marital status, sexual orientation, military or veteran status, age or any other legally protected category or status in accordance with applicable federal, state and local laws.

Man Group is a Disability Confident Committed employer; if you require help or information on reasonable adjustments as you apply for roles with us, please contact .

About Man Group

Global alternative investment management firm serving institutional clients.

Similar jobs

Site Reliability Engineer roles near Sofia, Sofia
14h
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Sr. SRE)
Sofia, Sofia City Province, Bulgaria
HybridFull Time
hosting.com
hosting.com: Global provider of web hosting and cloud infrastructure solutions.
5+ YOEFive years of Linux systems administration experience, Bash scripting, virtualization, Ansible, databases, web servers, complex incident resolution, and fluent written and spoken English.
cPanel, Plesk, Ansible, Linux, KVM, Proxmox, AWX, Salt, MariaDB, MySQL, Apache, LiteSpeed, NGINX, Bash, Docker, Podman, Python, PHP, Go, Xen, Hyper-V, RAID, Windows, IIS, Microsoft DNS, PowerShell
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Sofia, Sofia City Province, Bulgaria
HybridFull Time
Experian
ExperianLondon Stock Exchange: EXPN: Provides data and analytical tools to manage credit risk.
1+ YOE1+ years SRE experience, strong English, Kubernetes and cloud familiarity, Linux and networking troubleshooting, incident management, observability and IaC experience.
Kubernetes, EKS, Splunk, Dynatrace, Thousand Eyes, ServiceNow, Jira, Jenkins, Python, Java, Cassandra, Redis, Apigee, Okta, Postgres, AWS, Microsoft Azure, GCP, Infrastructure as Code, Git Ops
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sofia or Macedonia or Poland or Portugal
HybridFull Time
Valtech
Valtech: Experience innovation providing digital transformation and consulting services.
5+ years engineering experience (minimum 2 years as SRE), production incident management, programming/scripting, cloud and monitoring expertise, English C1+, experience with CI/CD and Kubernetes.
Datadog, New Relic, Dynatrace, Prometheus, Grafana, GitHub, Azure DevOps, GitLab, Jenkins, Docker, Kubernetes, EKS, Argo CI/CD, Java, Springboot, Kafka, AWS, Azure, GCP
3w
Save
Mark Applied
Hide
Sr Site Reliability Engineer
Sofia, Sofia City Province, Bulgaria
RemoteFull Time
Quickbase
Quickbase: Provides a no-code platform for building custom business applications.
7+ YOE7+ years SRE/Cloud/Platform experience with 4+ years in Microsoft Azure, IaC (Terraform/Azure Bicep/ARM), CI/CD, scripting, observability, security, and technical leadership.
Microsoft Azure, Azure Monitor, Log Analytics, Application Insights, Azure Key Vault, Microsoft Entra ID, Terraform, Azure Bicep, ARM templates, Azure DevOps, GitHub Actions, Jenkins, PowerShell, Python, Bash, Docker, Kubernetes, Azure Kubernetes Service (AKS), AWS, IAM, VPC, EC2, S3, CloudWatch, Azure Policy
1mo
Save
Mark Applied
Hide
Site Reliability Engineer – Middle
Poznań or Warszawa or Sofia
RemoteFull Time
SOFTSWISS
SOFTSWISS: Provider of comprehensive software solutions for the iGaming industry.
3+ YOE3+ years engineering experience; scripting/OOP (Ruby, Python, Java); config management (Ansible/Saltstack/Terraform); monitoring (DataDog/ELK/GrayLog); Kubernetes, DB internals, debugging; English and Russian proficiency.
Ruby, Python, Java, Ansible, Saltstack, Terraform, DataDog, ELK, GrayLog, K8S, Ruby on Rails, SQL, CI/CD
1mo
Save
Mark Applied
Hide
Sr Staff Site Reliability Engineer
Sofia, Sofia City Province, Bulgaria
OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
5+ YOE5+ years SRE experience, strong Kubernetes and Terraform skills, hands-on experience with a major cloud (GCP or AWS), Prometheus/Grafana, CI/CD and GitOps tooling, Python proficiency, incident response and distributed systems troubleshooting.
GCP, AWS, Azure, PagerDuty, Prometheus, Grafana, Kubernetes, Terraform, CI/CD, GitOps, Python, GitLab CI, GitHub Actions, Jenkins, Flux
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center
Sofia, Sofia City, Bulgaria
HybridFull Time
A1
A1Vienna Stock Exchange: TKA: Provides mobile, fixed-line, internet, and digital television services.
SRE with Kubernetes and observability, multi-cloud, and AIOps experience.
Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK/PLG, Datadog, Dynatrace, New Relic, Python, GitHub Actions
5mo
Save
Mark Applied
Hide
Site Reliability Engineer
Sofia, Sofia, Bulgaria
HybridFull Time
Flutter Entertainment
Flutter EntertainmentNYSE: FLUT: Operates online sports betting and digital casino gaming platforms.
5+ YOEOwn and maintain observability, monitoring, and reliability for AWS-based systems; strong SRE experience; cloud platforms AWS/Azure/GCP; on-call; 24/7 support.
Prometheus, Grafana, ELK, AWS, Azure, GCP, Jenkins, GitLab CI, Azure DevOps, Docker, Kubernetes, Terraform