This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

H-E-B
Posted 1mo ago

Senior Cloud Engineer

H-E-B
Austin, Texas, United States
OnsiteFull Time
Responsibilities
  • ensuring reliability
  • building infrastructure
  • automating operations
Requirements
  • 5+ years in SRE/cloud/platform engineering with Databricks, AWS
  • Terraform
  • Python, and large-scale distributed data systems experience
Technical tools mentioned
DatabricksAWS EC2AWS S3AWS VPCAWS IAMAWS LambdaCloudFormationPythonSQLTerraformApache SparkDockerKubernetesCI/CD

Job description

Responsibilities:

Since H-E-B Digital Technology's inception, we've been investing heavily in our customers' digital experience, reinventing how they find inspiration from food, how they make food decisions, and how they ultimately get food into their homes. This is an exciting time to join H-E-B Digital. We're using the best available technologies to deliver modern, engaging, reliable, and scalable experiences that meet the needs of our growing audience. If you enjoy solving complex technical challenges, working in a rapidly changing environment, learning new skills, and building platforms that power enterprise-scale data and digital experiences, we want you on our team.

Our Partners thrive The H-E-B Way. In the Senior Site Reliability Engineer, Data Platform role, that means you have a:

HEART FOR PEOPLE... you collaborate across engineering, data, and product teams, mentor others, and advocate for reliability and operational excellence.

HEAD FOR BUSINESS... you align reliability, scalability, and platform investments to business objectives while driving engineering best practices.

PASSION FOR RESULTS... you build resilient systems, automate operations, improve developer productivity, and ensure the availability of critical data platforms.


What You'll Do

As a Senior Site Reliability Engineer supporting H-E-B's Data Platform, you will be responsible for the reliability, scalability, performance, and operational excellence of cloud-native data infrastructure and services.

Key Responsibilities

  • Design, implement, and maintain highly available, resilient, and scalable data platform infrastructure.
  • Develop and manage Infrastructure as Code (IaC) solutions using Terraform and other automation tools.
  • Partner closely with Data Engineers, Data Scientists, Analysts, and Software Engineers to optimize platform reliability and performance.
  • Develop comprehensive monitoring, alerting, observability, SLO, and capacity planning strategies aligned with business objectives.
  • Monitor, troubleshoot, and optimize distributed storage, compute, and streaming systems across cloud-based environments.
  • Lead root cause analysis efforts, identify systemic risks, and implement preventative solutions.
  • Establish and champion best practices for platform engineering, reliability engineering, security, and operational excellence.
  • Improve system resiliency through architecture reviews, fault tolerance strategies, and performance optimization initiatives.
  • Build and maintain CI/CD pipelines that support rapid, reliable deployments.
  • Implement security best practices and ensure compliance with enterprise and industry standards.
  • Drive automation initiatives that reduce operational overhead and improve platform scalability.
  • Contribute to long-term platform and reliability roadmaps.
  • Stay current with emerging technologies and recommend innovative solutions that enhance platform capabilities.

Who You Are

Minimum Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience).
  • 5+ years of experience in Software Engineering, Platform Engineering, Site Reliability Engineering, Cloud Engineering, or related disciplines.
  • Strong experience managing and supporting large-scale cloud platforms and distributed systems.
  • Extensive experience with Databricks and large-scale data processing environments.
  • Deep expertise in AWS services, including EC2, S3, VPC, IAM, Lambda, CloudFormation, and related cloud technologies.
  • Strong programming and automation experience with Python; experience with SQL is preferred.
  • Hands-on experience with Terraform and Infrastructure as Code practices.
  • Experience with distributed data and compute technologies such as Apache Spark, streaming platforms, and ETL workflows.
  • Strong understanding of software engineering principles with an emphasis on reliability, scalability, observability, and performance optimization.
  • Experience implementing monitoring, logging, observability, and incident response practices.
  • Excellent analytical, troubleshooting, and problem-solving abilities.
  • Strong communication skills with the ability to collaborate across technical and business teams.

Preferred Qualifications

  • Experience with additional cloud platforms such as Google Cloud Platform (GCP) or Microsoft Azure.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Familiarity with CI/CD platforms and deployment automation tools.
  • Experience supporting data lake, lakehouse, or modern data platform architectures.
  • AWS, Terraform, Kubernetes, Databricks, or other relevant industry certifications.
  • Experience mentoring engineers and leading cross-functional technical initiatives.

What Makes You Successful

  • You take ownership of reliability and operational outcomes across complex distributed systems.
  • You proactively identify risks and drive long-term solutions rather than short-term fixes.
  • You can balance technical excellence with practical business needs.
  • You thrive in fast-paced environments and effectively manage competing priorities.
  • You influence engineering culture through collaboration, mentorship, and continuous improvement initiatives.
  • You are passionate about automation, scalability, and creating exceptional platform experiences for engineering teams.

Working Conditions

  • Function in a fast-paced retail and technology environment.
  • Travel by car or plane, including occasional overnight stays.
  • Sit for extended periods while working at a computer.
  • Participate in on-call rotations and support critical production systems as needed.
  • Work flexible hours when required to support business-critical initiatives.

The responsibilities listed above describe the general nature and level of work performed and are not intended to be an exhaustive list of all duties, responsibilities, or qualifications associated with this position.

About H-E-B

Operates a major supermarket chain in Texas and Mexi

Similar jobs

Site Reliability Engineer roles near Austin, Texas
3d
Save
Mark Applied
Hide
Site Reliability Engineer - Data, Apple Ads
Austin, Texas, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
2+ YOERequires 2+ years of infrastructure engineering, AWS and distributed data systems expertise, Apache Spark/Flink/Kafka/Iceberg, Kubernetes, ML and GenAI experience, and programming skills.
AWS, Apache Spark, Flink, Kafka, Iceberg, Kubernetes, Java, Scala, Kotlin, Python, Helm, CRD
3d
Save
Mark Applied
Hide
Engineer III, Site Reliability
Austin or Cranberry Township
HybridFull Time
Omnicell
OmnicellNASDAQ: OMCL: Healthcare technology providing medication-management automation for healthcare settings and pharmacies.
5+ YOEBachelor's degree in a technical field; 5+ years in software or platform engineering, including 3+ years in SRE, DevOps, or reliability; cloud, Python, Kubernetes, Docker, Helm, Terraform, Linux, and incident response experience.
AWS, Azure, GCP, Python, Kubernetes, Docker, Helm, Terraform, GitHub Actions, CodeFresh, TeamCity, Octopus Deploy, IBM, HCL, ArgoCD, Flux, Kafka, RabbitMQ
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, ANZ
Christchurch or Auckland or London or San Francisco or Austin
OnsiteFull Time
Partly
Partly: Automotive AI infrastructure helping vehicle-repair businesses identify, source, and manage parts.
Senior SRE experience building scalable cloud infrastructure, Kubernetes clusters, CI/CD systems, Linux environments, and production software; leadership, ownership, communication, and troubleshooting skills required.
Kubernetes, Terraform, GCP, ArgoCD, Python, Bash, Docker, GitOps, GitLab CI, Kafka, Apache Cassandra, Postgres, Rust, Linux
5d
Save
Mark Applied
Hide
Site Reliability Developer 4
Austin, Texas, United States
$102k-$210k/yr RemoteFull Time
Oracle Corporation
Oracle CorporationNYSE: ORCL: Cloud infrastructure and enterprise software solutions provider.
6+ YOEMaster’s in computer science, engineering, or related field plus 6 years’ network/software development experience. Requires optical DWDM, cloud networking, Python, Linux, protocols, automation, and distributed systems expertise.
Python, Object-Oriented programming, ROADM, FOADM, Mux/Demux, Flex Mux/Demux, OTDR, OSA, OCS, BERT, TCP/IP, BGP, OSPF, YANG, OpenConfig, NETCONF, Linux, Jira, Git, Scrum, Agile, REST API, shell
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Austin or Reston
HybridFull Time
Seekr Technologies
Seekr Technologies: Private American enterprise AI providing explainable, secure AI software and hardware to government and critical-infrastructure customers.
5+ YOERequires 5+ years in site reliability engineering and Linux systems, monitoring and logging expertise, programming or scripting skills, Docker, Kubernetes, configuration automation, and incident response experience.
Linux, ELK, Prometheus, InfluxDB, Grafana, Python, Ruby, Bash, Java, Docker, Kubernetes, Puppet, Chef, Ansible, Terraform, Elasticsearch, Kafka, Aerospike, Git, GitHub, GitLab, ArgoCD
2w
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerThe Nasdaq Stock Market LLC: BTDR: Public Singaporean Bitcoin mining and AI cloud infrastructure serving enterprises with computing, datacenters, and mining solutions.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
2w
Save
Mark Applied
Hide
Cleared Senior Site Reliability Engineer
Austin, Texas, United States
$80k-$210k/yr OnsiteFull Time
Gallatin AI
Gallatin AI: Defense technology providing AI-powered logistics decision-support software to military and allied defense organizations.
3+ YOERequires active Secret clearance, 3–5 years in SRE, DevOps, or production infrastructure, Linux, networking, cloud infrastructure, infrastructure-as-code, containers, CI/CD, monitoring, and incident response experience.
Linux, AWS, Azure, Terraform, Ansible, Kubernetes, Docker, Prometheus, Grafana, Datadog, ELK, Microsoft Azure Government, Microsoft Azure Government Secret, Microsoft Azure Government IL5, Microsoft Azure Government IL6
2w
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNYSE: IBM: Global technology and consulting focusing on cloud and AI.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
This job has expired