This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

LTIMindtree
Posted 10mo ago

Specialist - Software Engineering

LTIMindtree
Basildon, England, United Kingdom
OnsiteFull Time
Responsibilities
  • managing clusters
  • automating deployments
  • responding incidents
Requirements
  • Requires 3+ years in SRE or DevOps
  • Kubernetes and major cloud expertise
  • Scripting, IaC, Linux
  • Networking
  • Monitoring, CI/CD
  • Troubleshooting, and collaboration skills
Technical tools mentioned
KubernetesAWSAzureGCPTerraformHelmCloudFormationPrometheusGrafanaELKDatadogJenkinsGitLab CIArgoCDPythonBashGoLinuxIstioLinkerdFlux

Job description

Job Details

  • Location: Select Location
  • Experience: 0 - 0 Years
  • Job Type: Permanent
  • Openings: 1

Role Description

Role-Kubernetes and Cloud SRE

Type-FTE/Perm

Mode-5 Days from Basildon

Site Reliability Engineer SRE Kubernetes Cloud

Position Summary

We are seeking a highly skilled Site Reliability Engineer SRE with deep expertise in Kubernetes

and cloud technologies AWS Azure or GCP The SRE will be responsible for designing deploying

automating and supporting highly available scalable and secure containerized applications in

cloudnative environments You will work closely with development operations and security teams

to ensure the reliability performance and efficiency of our production systems

Key Responsibilities

Design deploy and manage Kubernetes clusters onpremises andor cloudmanaged

such as EKS AKS GKE to support scalable microservices architectures

Automate infrastructure provisioning and application deployment using Infrastructure

as Code IaC tools such as Terraform Helm or CloudFormation

Monitor troubleshoot and optimize system performance using observability tools

Prometheus Grafana ELK Datadog etc

Implement and manage CICD pipelines to ensure rapid repeatable and reliable software

delivery

Ensure system reliability availability and security through proactive monitoring incident

response and root cause analysis

Develop and maintain runbooks dashboards and documentation for operational

procedures and system architectures

Participate in oncall rotations and respond to production incidents ensuring minimal

downtime and fast recovery

Collaborate with development and operations teams to drive DevOps and SRE best

practices including capacity planning scaling and cost optimization

Continuously improve automation tooling and processes to reduce manual work and

increase system reliability

Required Skills Experience

3 years experience as an SRE DevOps Engineer or similar role supporting largescale

productiongrade environments

Expertise in Kubernetes deployment scaling upgrades troubleshooting networking

RBAC etc

Handson experience with at least one major cloud provider AWS Azure or GCP

Proficiency in scriptingprogramming Python Bash Go etc

Experience with IaC tools Terraform Helm CloudFormation ARM etc

Strong knowledge of Linux systems administration and networking concepts

Familiarity with monitoring logging and ing tools Prometheus Grafana ELKEFK

Datadog etc

Experience with CICD tools Jenkins GitLab CI ArgoCD etc

Understanding of security best practices in cloud and containerized environments

Excellent troubleshooting and problemsolving skills

Strong communication and collaboration skills

Preferred Qualifications

Certified Kubernetes Administrator CKA or similar certification

Experience with service mesh Istio Linkerd ingress controllers and API gateways

Experience in a multicloud or hybrid cloud environment

Familiarity with GitOps practices and tools ArgoCD Flux

Experience with disaster recovery backup and business continuity planning

Education

Bachelors degree in Computer Science Engineering or related field or equivalent

experience

This role is ideal for engineers who are passionate about automation reliability and modern

cloudnative architectures and who thrive in fastpaced collaborative environments

Skills

Mandatory Skills : Monitoring & Observability, Service Level & Error Budget Management, Automation & Scripting, Critical Incident Response

About LTIMindtree

AI-centric global technology services and consulting.

Similar jobs

Site Reliability Engineer roles near Basildon, England
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, ANZ
Christchurch or Auckland or London or San Francisco or Austin
OnsiteFull Time
Partly
Partly: Automotive AI infrastructure helping vehicle-repair businesses identify, source, and manage parts.
Senior SRE experience building scalable cloud infrastructure, Kubernetes clusters, CI/CD systems, Linux environments, and production software; leadership, ownership, communication, and troubleshooting skills required.
Kubernetes, Terraform, GCP, ArgoCD, Python, Bash, Docker, GitOps, GitLab CI, Kafka, Apache Cassandra, Postgres, Rust, Linux
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Hinxton, England, United Kingdom
£4k/mo HybridFull Time, Contract
EMBL-EBI
EMBL-EBI: Intergovernmental bioinformatics research institute providing open biological data, computational research, training, and industry services to scientists.
5+ YOERequires 5+ years with Linux production systems, automation and orchestration experience, troubleshooting skills, directory services knowledge, and ability to review Python, Bash, and Puppet code.
Linux, Active Directory, Entra-ID, Red Hat IDP, Postfix, Cyrus, Roundcube, Mailman, O365, Globus, Aspera, FTP, HTTP, Check_mk, Gerrit, Foreman, RPM, Puppet, NTP, SSSD, Red Hat OS, tcpdump, strace, 389 DS, OpenLDAP, Python, Bash, Zoom
4d
Save
Mark Applied
Hide
Site Reliability Engineer - Private Cloud Compute
London, England, United Kingdom
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
4+ YOEBachelor's degree in computer science or equivalent experience; 4+ years managing distributed systems in cloud environments; programming experience and expertise with service operations, troubleshooting, automation, and scalability.
Java, Go, Python, Perl, Kubernetes, Nginx, Envoy, Prometheus, Docker, HTTP, DNS, ECMP, TCP/IP, ICMP, Linux, Puppet, Chef, Ansible, Salt
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Byfleet or Hounslow
HybridFull Time
Looper Insights
Looper Insights: SaaS analytics helping streamers, broadcasters, studios, and regulators measure and optimize content visibility across connected-TV platforms.
Requires scripting or development experience, network troubleshooting, Unix CLI, physical device troubleshooting, systems administration, DevOps, automation, and regular travel to Byfleet and Hounslow data centres.
Unix CLI, gstreamer, VPN
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Vancouver or Toronto or Tel Aviv or Plano or London or Amsterdam or Tbilisi or Medellín or Foster City
$100k-$125k/yr HybridFull Time
Tipalti
Tipalti: Private fintech providing accounts-payable, payments, procurement, expense, and treasury automation for mid-market businesses.
4+ YOERequires 4+ years of software engineering experience, production deployment and application lifecycle ownership, architecture knowledge, troubleshooting skills, and English communication. SRE, cloud, monitoring, and scripting experience preferred.
.NET, TypeScript, OpenTelemetry, Prometheus, Bash, PowerShell, Python, AWS, GCP, Azure
5d
Save
Mark Applied
Hide
Site Reliability Engineer, Security Engineering
Sydney or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteFull Time
TikTok
TikTok: Short-form mobile video and social media platform.
3+ YOEBachelor's degree in computer science or related field, 3+ years relevant experience, programming in Go, Java, or Python, web framework experience, Linux and networking knowledge, Kubernetes and SRE tooling experience.
Go, Java, Python, Gin, Django, Spring, Linux, TCP/IP, HTTP, Kubernetes, Ansible, Argo CD, Prometheus, Grafana
6d
Save
Mark Applied
Hide
Site Reliability Engineer
Cambridge or United Kingdom
£65k/yr HybridFull Time
Royal Society of Chemistry
Royal Society of Chemistry: UK nonprofit learned society and professional body advancing chemistry through publishing, membership, education, events and policy.
Experience designing, deploying, and operating AWS infrastructure with Terraform; large-scale AWS data delivery; DevOps practices, security vulnerability mitigation, automation, and stakeholder collaboration.
Amazon Web Services (AWS), Terraform, Amazon S3, AWS DataSync, AWS Transfer Family, Amazon CloudFront, AWS Direct Connect, Continuous Integration/Continuous Deployment (CI/CD)
1w
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Infrastructure Observability
London, England, United Kingdom
HybridFull Time
T. Rowe Price
T. Rowe PriceNASDAQ: TROW: Global asset management firm providing investment and retirement services.
10+ YOEBachelor's degree or equivalent, 10+ years designing and operating cloud infrastructure, 5+ years with AWS and DevOps/SRE functions, programming, databases, observability, automation, SLOs, SLIs, and incident recovery.
Amazon AWS, Python, Java, GO, Node.js, .Net Core, SQL Server, PostgreSQL, MySQL, New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, Ansible, Terraform, Vault, Vagrant, Azure
This job has expired