Snapp
Posted 9mo ago

Site Reliability Engineer

Snapp
Tehran, Tehrān, IR
OnsiteFull Time
Responsibilities
  • Automate processes
  • Maintain observability
  • Manage test/staging environments
Requirements
  • SRE concepts
  • Python or scripting
  • Observability tools
  • Kubernetes
  • Databases
  • Incident management
Technical tools mentioned
PythonPrometheusGrafanaELKLokiJaegerTempoKubernetesHelmRedisRDBMS

Job description

Job description

In this role, you will strengthen the SRE Platform team’s mission by advancing the foundational platforms that automate manual workflows and elevate system reliability. Your work will ensure our staging environments remain stable and production-like, empowering QA and development teams to test, validate, and deploy their applications with confidence. You will also contribute to operational excellence through active participation in the weekly on-call rotation, supporting consistent and dependable infrastructure performance.

  • Automate and optimize operational processes

  • Enhance and maintain the observability stack

  • Oversee test/staging environments management

  • Develop and support critical production components

  • Handle and resolve production incidents

  • Participate in the on-call rotation

Job requirements

  • Strong teamwork and collaboration skills

  • Solid understanding of SRE concepts, including SLIs, SLOs, SLAs, and Error Budgets

  • Proficiency in Python or another scripting language

  • Strong grasp of software engineering principles

  • Hands-on experience with observability and monitoring tools such as Prometheus and Grafana

  • Familiarity with logging stacks (e.g., ELK, Loki) and tracing systems (e.g., Jaeger, Tempo)

  • Understanding of RDBMS and Redis

  • Experience working with Kubernetes and related tooling (e.g., Helm)

About Snapp

Iranian private technology operating a super-app for rides, delivery, shopping, travel, payments, insurance, and healthcare.

Similar jobs

Site Reliability Engineer roles near Tehran, Tehrān
1y
Save
Mark Applied
Hide
Site Reliability Engineer
Tehran, Tehran, IR
HybridFull Time
ArvanCloud
ArvanCloud: Iranian cloud infrastructure and AI provider serving online businesses in Iran and worldwide.
3+ YOELinux administration, large-scale maintenance, networking (TCP/IP, TLS, DNS, HTTPS), Git/GitFlow, Golang or Python, Bash, CM tools (Ansible, SaltStack), CI/CD with GitLab, databases and HA, Kafka/RabbitMQ, monitoring/logging (ELK, Prometheus, Grafana), containers and orchestration (Docker, Kubernetes), SLI/SLO, Ceph knowledge a plus.
Linux, Git, GitFlow, Go, Python, Bash, Ansible, SaltStack, GitLab, Kafka, RabbitMQ, ELK Stack, Prometheus, Grafana, Docker, Kubernetes, Ceph