HeyMilo is redefining the recruiting industry with AI-powered hiring solutions. Our platform operates at massive scale, processing thousands of AI-driven interviews daily, and we must strive to maintain a 100% SLA for our customers.
We’re looking for a Senior DevOps Engineer to own platform reliability, observability, and infrastructure security. This role is mission-critical — you’ll ensure that our systems stay up, run fast, and remain secure.
Who You Are
Cloud & DevOps Expert – You have deep experience with AWS, Containerization, and DigitalOcean.
Observability-Driven – You believe that nothing should happen in production without being tracked.
Reliability-Focused – You take 100% SLA uptime seriously and proactively prevent failures.
Security-Conscious – You lock down access, secrets, and API keys to ensure secure deployments.
CS Fundamentals – You have strong systems engineering principles to optimize infrastructure performance.
What You’ll Do
Monitor and maintain AWS and DigitalOcean infrastructure to ensure 100% uptime.
Build observability systems to track every aspect of platform health (Prometheus, Grafana).
Automate deployments using ECR, EKS, Terraform, and CI/CD pipelines (Github Actions).
Secure access tokens, API keys, and environment variables while managing secrets.
Optimize serverless functions (Lambdas) and database performance.
Proactively detect and mitigate system failures before they impact customers.
Ensure S3 storage, Redis, and ClickHouse databases operate at peak efficiency.
Overlook the infrastructure for long term success. Consistently ask yourself - “Can we scale to handle 1 million interviews TODAY?”
Technical Expertise:
Cloud: AWS (EC2, S3, IAM, RDS, Lambda, EKS, SQS), DigitalOcean
Infrastructure: Kubernetes, Docker, Terraform
Observability: Prometheus, Grafana, OpenTelemetry, InfluxDB, Opensearch, Betterstack
Security: Secrets management, access control, compliance
Databases: Redis, ClickHouse, MongoDB,
CI/CD: GitHub Actions, Terraform pipelines
What’s In It for You
Own the DevOps stack for a high-scale AI platform processing thousands of interviews per day.
Work on cutting-edge observability and reliability engineering challenges.
Competitive salary and an opportunity to work in a mission-critical AI infrastructure role.