Job Details
- Location: Select Location
- Experience: 0 - 0 Years
- Job Type: Permanent
- Openings: 1
Role Description
Role-Kubernetes and Cloud SRE
Type-FTE/Perm
Mode-5 Days from Basildon
Site Reliability Engineer SRE Kubernetes Cloud
Position Summary
We are seeking a highly skilled Site Reliability Engineer SRE with deep expertise in Kubernetes
and cloud technologies AWS Azure or GCP The SRE will be responsible for designing deploying
automating and supporting highly available scalable and secure containerized applications in
cloudnative environments You will work closely with development operations and security teams
to ensure the reliability performance and efficiency of our production systems
Key Responsibilities
Design deploy and manage Kubernetes clusters onpremises andor cloudmanaged
such as EKS AKS GKE to support scalable microservices architectures
Automate infrastructure provisioning and application deployment using Infrastructure
as Code IaC tools such as Terraform Helm or CloudFormation
Monitor troubleshoot and optimize system performance using observability tools
Prometheus Grafana ELK Datadog etc
Implement and manage CICD pipelines to ensure rapid repeatable and reliable software
delivery
Ensure system reliability availability and security through proactive monitoring incident
response and root cause analysis
Develop and maintain runbooks dashboards and documentation for operational
procedures and system architectures
Participate in oncall rotations and respond to production incidents ensuring minimal
downtime and fast recovery
Collaborate with development and operations teams to drive DevOps and SRE best
practices including capacity planning scaling and cost optimization
Continuously improve automation tooling and processes to reduce manual work and
increase system reliability
Required Skills Experience
3 years experience as an SRE DevOps Engineer or similar role supporting largescale
productiongrade environments
Expertise in Kubernetes deployment scaling upgrades troubleshooting networking
RBAC etc
Handson experience with at least one major cloud provider AWS Azure or GCP
Proficiency in scriptingprogramming Python Bash Go etc
Experience with IaC tools Terraform Helm CloudFormation ARM etc
Strong knowledge of Linux systems administration and networking concepts
Familiarity with monitoring logging and ing tools Prometheus Grafana ELKEFK
Datadog etc
Experience with CICD tools Jenkins GitLab CI ArgoCD etc
Understanding of security best practices in cloud and containerized environments
Excellent troubleshooting and problemsolving skills
Strong communication and collaboration skills
Preferred Qualifications
Certified Kubernetes Administrator CKA or similar certification
Experience with service mesh Istio Linkerd ingress controllers and API gateways
Experience in a multicloud or hybrid cloud environment
Familiarity with GitOps practices and tools ArgoCD Flux
Experience with disaster recovery backup and business continuity planning
Education
Bachelors degree in Computer Science Engineering or related field or equivalent
experience
This role is ideal for engineers who are passionate about automation reliability and modern
cloudnative architectures and who thrive in fastpaced collaborative environments
Skills
Mandatory Skills : Monitoring & Observability, Service Level & Error Budget Management, Automation & Scripting, Critical Incident Response