This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

Exadel
Posted 1mo ago

Senior Site Reliability Engineer (DevOps, Java)

Exadel
Bulgaria or Georgia or Lithuania or Poland or Romania or Uzbekistan
HybridFull Time
Responsibilities
  • designing systems
  • improving availability
  • automating deployments
Requirements
  • 7+ years experience with Kubernetes and AWS
  • Strong Java skills
  • Docker/container experience
  • CI/CD and IaC (Terraform/CloudFormation/Pulumi)
  • Messaging systems
  • Relational and NoSQL databases
  • Monitoring/observability, and SRE practices
  • English upper-intermediate
Technical tools mentioned
KubernetesAWSJavaRubyDockerHelmEC2EKSRDSS3IAMVPCLoad BalancersCloudWatchTerraformCloudFormationPulumiRabbitMQKafkaSQSPulsarMySQLPostgreSQLRedisDynamoDBMongoDBPrometheusGrafanaDatadogELK/OpenSearchOpenTelemetryRuby on RailsSpring BootPayPalMasterPassStripeTwilioLinuxTCP/IPDNSHTTP/HTTPSTLS

Job description

Why Join Exadel 

We’re an AI-first global tech company with 25+ years of engineering leadership, 2,000+ team members, and 500+ active projects powering Fortune 500 clients, including HBO, Microsoft, Google, and Starbucks.

From AI platforms to digital transformation, we partner with enterprise leaders to build what’s next.
What powers it all? Our people are ambitious, collaborative, and constantly evolving.

About the Client  

The company has been building solutions for mobile apps, effortless payment, business travel, and advertising since 1992. The customer is developing a mobility platform that allows operators to manage their vehicles and drivers efficiently, regulators to be informed and establish guidelines, service providers to deliver sustainable solutions, and riders to have an effortless transit experience.

Project Team

The project is a taxi ordering service and it contains the following components:

  • Ride server (all data processing)
  • Payment server (PCI DSS-compliant), which performs a transaction with the passenger's digital wallet and payment gateways
  • Mobile application (hail taxi, geocoding, map, payments)
  • Taxi terminal (3rd party)

The project integrates with 3rd party services, including PayPal, MasterPass, Stripe, and Twilio.

The team itself consists of 1 Team Lead, 1 UI Developers, 6 Back-End Developers, 9 Mobile Developers, 5 QA, 2 BA and 2 Designer.

What You’ll Do  

  • Design, build, and operate reliable, scalable distributed systems
  • Improve system availability, performance, and resilience
  • Automate infrastructure, deployments, and operational processes
  • Diagnose and resolve production issues
  • Lead upgrades and migrations with minimal or zero downtime
  • Participate in on-call rotations and incident response
  • Collaborate closely with development teams to improve operability
  • Drive best practices around monitoring, alerting, and capacity planning
  • Reduce operational toil through automation
  • Contribute to incident management, post-mortems, disaster recovery strategies, and continuous reliability improvements

What You Bring  

  • 7+ years of experience, specializing in Kubernetes and AWS
  • Strong programming skills in Java, with willingness to learn Ruby
  • Solid understanding of concurrency, runtime behavior, and performance optimization
  • Hands-on experience with Docker and containerized workloads
  • Strong Kubernetes expertise (Deployments, StatefulSets, Services, Ingress, Helm, troubleshooting, autoscaling)
  • Strong AWS experience (EC2, EKS, RDS, S3, IAM, VPC, Load Balancers, CloudWatch)
  • Experience designing infrastructure for high availability and disaster recovery
  • Experience with CI/CD pipelines and Infrastructure as Code (Terraform, CloudFormation, Pulumi, or similar
  • Experience with RabbitMQ or similar messaging systems (Kafka, SQS, Pulsar, etc.
  • Strong understanding of relational databases (MySQL/PostgreSQL), including query optimization, replication, and failover strategies
  • Familiarity with NoSQL and in-memory databases (Redis, DynamoDB, MongoDB)
  • Experience with distributed systems, microservices, capacity planning, and fault tolerance
  • Experience with monitoring and observability tools (Prometheus, Grafana, Datadog, ELK/OpenSearch, OpenTelemetry)
  • Strong understanding of Linux systems and networking fundamentals (TCP/IP, DNS, HTTP/HTTPS, TLS, load balancing)
  • Experience with SRE practices, including SLOs/SLIs/SLAs, load testing, resilience testing, and incident management
  • Strong communication skills and ability to collaborate across engineering teams
  • Calm and effective during incidents with an ownership mindset

Nice to Have

  • Experience operating production systems written in Ruby, Java, or other major platforms
  • Framework experience such as Ruby on Rails, Spring Boot, or similar
  • Experience operating high-traffic SaaS platforms
  • Cost optimization in cloud environments
  • Chaos engineering practices
  • Experience mentoring junior engineers

English level

Upper-Intermediate

Legal & Hiring Information 

  • Exadel is proud to be an Equal Opportunity Employer committed to inclusion across minority, gender identity, sexual orientation, disability, age, and more
  • Reasonable accommodations are available to enable individuals with disabilities to perform essential functions
  • Please note: this job description is not exhaustive. Duties and responsibilities may evolve based on business needs

Your Benefits at Exadel 

Exadel benefits vary by location and contract type. Your recruiter will fill you in on the details.

  • International projects
  • In-office, hybrid, or remote flexibility
  • Medical healthcare
  • Recognition program
  • Ongoing learning & reimbursement 
  • Well-being program
  • Team events & local benefits 
  • Sports compensation 
  • Referral bonuses 
  • Top-tier equipment provision

Exadel Culture

We lead with trust, respect, and purpose. We believe in open dialogue, creative freedom, and mentorship that helps you grow, lead, and make a real difference. Ours is a culture where ideas are challenged, voices are heard, and your impact matters.

About Exadel

Provides digital engineering and AI-focused software consulting services.

Year founded
1998

Similar jobs

Site Reliability Engineer roles
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Sr. SRE)
Plovdiv, Plovdiv Province, Bulgaria
HybridFull Time
hosting.com
hosting.com: Global provider of web hosting and cloud infrastructure solutions.
5+ YOERequires 5+ years of Linux systems administration, Bash scripting, virtualization, automation and configuration management experience, especially Ansible, plus databases, web servers, incident resolution, and fluent English.
cPanel, Plesk, Ansible, Linux, KVM, Proxmox, AWX, Salt, MariaDB, MySQL, Apache, LiteSpeed, NGINX, Docker, Podman, Python, PHP, Go, Xen, Hyper-V, RAID, Windows, IIS, Microsoft DNS, PowerShell, Bash
1d
Save
Mark Applied
Hide
Site Reliability Engineer - Storage
Geneva or Paris or Zurich or Prague or Barcelona or London or Vilnius or Skopje or Taipei
€50k-€77k/yr HybridFull Time
Proton
Proton: Providing end-to-end encrypted communication and storage tools.
Experience building complex production systems, diagnosing critical-system issues, software development, security best practices, and cryptography concepts; storage systems experience is a bonus.
Ceph, Seaweed, Tape, Greenhouse
1d
Save
Mark Applied
Hide
Senior SRE
Warsaw, Masovian Voivodeship, Poland
HybridFull Time
Tango
Tango: Global live-streaming platform for creators and social interaction
Extensive Kubernetes, GCP or cloud, Terraform, Prometheus, GitLab, Linux, and Python or Bash experience; strong troubleshooting, automation, communication, and technical documentation skills.
Kubernetes, GCP, Terraform, Prometheus, GitLab, Linux, Python, Bash, ArgoCD, Argo Rollouts, Envoy, VictoriaMetrics, Go, GKE, Kafka, Redis, Aerospike, Envoy Gateway, OpenTelemetry
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Spain or Ukraine or United Kingdom or Poland or Ireland or Slovakia or Portugal or Romania
RemoteFull Time
Tempo Software
Tempo Software: Provides software for time tracking and resource management.
5+ YOERequires 5+ years of relevant experience and strong Kubernetes, Docker, Git, Linux, Bash, and Terraform skills, plus cloud deployment, monitoring, troubleshooting, communication, and collaboration abilities.
Kubernetes, Docker, Git, Linux, Bash, Terraform, AWS, Helm, FluxCD, GitHub Actions, Ansible, Datadog, Java, Kotlin
2d
Save
Mark Applied
Hide
[MLA] Senior Site Reliability Engineer - AI Experience Framework
Kraków, Lesser Poland Voivodeship, Poland
RemoteFull Time
Software Mind
Software Mind: Provides software engineering and digital transformation services to businesses.
Requires SRE experience with Kubernetes, Node.js and JVM troubleshooting, Prometheus, Grafana, Linux, networking, CI/CD, infrastructure as code, on-call operations, Splunk, Valkey/Redis, mTLS, JWT, and fluent English.
Kubernetes, Grafana, Prometheus, Node.js, JVM, Linux, Helm, GitOps, ArgoCD, Flux, Splunk, Valkey, Redis, mTLS, JWT, KEDA, PromQL, HPA, HTTP/2, ServiceNow, Lit, V8
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Traffic Routing (Limited-time Relocation Bonus)
Kaunas, Kaunas County, Lithuania
€5k-€9k/mo HybridFull Time
Vinted
Vinted: Online marketplace for buying and selling second-hand items.
Production software experience in Go, Python, C++, or similar; strong Linux and performance-debugging skills; infrastructure automation with Ansible or Chef; reliability practices; and fluent written and spoken English.
Go, Python, C++, Linux, Ansible, Chef, Istio, Envoy, SONiC, Cumulus, Postfix, TCP/IP, BGP, DNS
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE Team)
Cyprus or Czech Republic or Poland or Serbia or Spain
HybridFull Time
Semrush
Semrush: Brand visibility and digital marketing management platform.
3+ YOERequires 3+ years as an SRE, Kubernetes and cloud provider experience, Python or Go engineering, application debugging with metrics, observability knowledge, communication skills, and on-call availability.
Kubernetes, Python, Go, GCP
3d
Save
Mark Applied
Hide
DevOps & Site Reliability Engineer (GCP)
Romania
OnsiteFull Time
Qureos
Qureos: AI-powered recruitment platform automating candidate sourcing and screening.
5+ YOEBachelor's degree or equivalent experience, 5+ years in DevOps or a similar role, Docker and Kubernetes expertise, GCP experience, MongoDB administration, scripting, infrastructure as code, and CI/CD experience.
Google Cloud Platform (GCP), Docker, Kubernetes, GitHub Actions, Google Cloud Run, Elasticsearch, Kibana, Prometheus, New Relic, MongoDB, Python, Bash, Terraform, Ansible, CircleCI, Compute Engine, Cloud Storage, GKE, AWS, Azure, Agile, Scrum
This job has expired