39 reliability engineering manager jobs at 30 companies in New Braunfels, TX

3w
Save
Mark Applied
Hide
Engineering Manager, Support & Stability
Columbus or Boston or New York City or Chicago or Austin or Los Angeles
$150k-$190k/yr RemoteFull Time
Loop Returns
Loop Returns: Software platform automating e-commerce returns and post-purchase experiences.
3+ YOE3+ years engineering management experience, experience with platform reliability/system health, AI agent adoption, technical depth in high-risk domains, ability to manage and grow engineers.
PHP, Laravel, Vue.js, MySQL, DynamoDB, Kubernetes, AWS, Jira, Claude, Cursor, Datadog, Snowflake
1mo
Save
Mark Applied
Hide
Manager, Engineering Observability
Austin, Texas, United States
$169k-$232k/yr HybridFull Time
Procore
ProcoreNYSE: PCOR: Cloud-based construction management software for projects and teams.
7+ YOE2+ Mgmt7+ years as a software engineer, 2+ years managing teams; hands-on backend/distributed systems experience; experience with observability tools (Datadog, OpenTelemetry, Prometheus/Grafana, Honeycomb, SumoLogic, Bugsnag); SRE/reliability background; strong leadership and roadmap skills.
Datadog, Honeycomb, SumoLogic, OpenTelemetry, Bugsnag, Prometheus, Grafana
3w
Save
Mark Applied
Hide
Engineering Manager, Data Feeds
New York City or Boston or Chicago or Salt Lake City or Austin or Montreal or Canada or Washington or Dallas or United States or Vancouver or Toronto or Charlotte or Denver
$129k-$304k/yr RemoteFull Time
Chainlink Labs
Chainlink Labs: Building decentralized oracle networks for blockchain smart contracts.
Deep experience building and operating production blockchain infrastructure, strong engineering judgment, leadership and coaching experience, stakeholder management, and ability to deliver reliable, scalable systems.
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1w
Save
Mark Applied
Hide
Reliability Engineer, Mechanical, NA
San Antonio or Ashburn or North America or Europe or Asia
$125k-$135k/yr HybridFull Time
Vantage Data Centers
Vantage Data Centers: Provides hyperscale data center campuses for cloud and AI providers.
2+ YOEBachelor's degree in electrical or mechanical engineering preferred, 2–3 years in critical facility operations and maintenance, strong project management and communication skills, and willingness to travel up to 25%.
Root-Cause Failure Analysis, Facility Acceptance Testing
2w
Save
Mark Applied
Hide
Technical Lead Manager - Production Engineering
Austin, Texas, United States
OnsiteFull Time
Apptronik
Apptronik: Designs and manufactures humanoid robots for industrial automation.
8+ YOE2+ MgmtHands-on technical manager to build and lead production engineering for robot deployment, commissioning, and fleet reliability; 8+ years engineering or 4+ years humanoid experience, 2+ years managerial experience, willing to travel up to 25%.
Linux, DDS, UDP, TCP, Ansible, C++, Python, ROS 2, TypeScript, Kubernetes, React, CSS
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
2K
2KNASDAQ: TTWO: Publishes and develops global video game franchises and entertainment.
5+ YOE5+ years SRE/platform engineering experience, deep Kubernetes (EKS/GKE), Terraform/Pulumi and GitOps, observability with Prometheus/Grafana/Datadog, production coding in Go/Python/TypeScript, Linux and networking expertise, incident management.
Terraform, Pulumi, ArgoCD, Flux, Kubernetes, EKS, GKE, Istio, Cilium, Helm, Terragrunt, Prometheus, Grafana, Datadog, OpenTelemetry, GitHub Actions, Jenkins, Go, Python, TypeScript, PasswordState, 1Password, AWS Secrets Manager, OPA/Gatekeeper, AWS, GCP, VMware, Ansible, Puppet, AWS Systems Manager
1mo
Save
Mark Applied
Hide
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems
Austin or Seattle or Cupertino
$208k-$282k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
10+ YOE5+ Mgmt10+ years reliability/quality engineering experience with server or high-volume electronics, 5+ years people management, bachelor's degree in a relevant field, experience with root-cause analysis, quality systems, and multi-vendor manufacturing.
HALT, HTOL, thermal cycling, QRV, DFMEA, SPC, FMEA, 8D, DOE, Weibull analysis
4w
Save
Mark Applied
Hide
Site Reliability Engineer
San Antonio, Texas, United States
OnsiteFull Time
Infosys
InfosysNYSE: INFY: Provides IT consulting, software development, and business outsourcing services.
Experience with reliability engineering, Terraform IaC, observability (Datadog), disaster recovery, vulnerability management, CI/CD (Harness, Helm), plus bachelor's degree or equivalent experience.
Terraform, Datadog, Harness, Helm Charts, AWS
3d
Save
Mark Applied
Hide
Head of Commissioning and Reliability
Austin or California or New York or Denver or Calgary or Toronto
$221k-$260k/yr RemoteFull Time
Intersect
IntersectNASDAQ: GOOGL: Develops-located data centers and renewable energy infrastructure.
12+ YOEBachelor's degree in a related engineering field and 12+ years commissioning and reliability experience with high-voltage infrastructure, utility-scale power, or mission-critical industrial facilities.
Isograph, ReliaSoft, SCADA, CMMS, Reliability-Centered Maintenance (RCM)
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer (SRE)
San Antonio, Texas, United States
HybridFull Time
IPSecure
IPSecure: Provides cybersecurity and incident response services for government agencies.
7+ YOE7+ years SRE/DevOps experience, active Secret clearance, cloud/DevOps certification (AWS), Splunk or Elastic Stack experience, Linux/Windows administration, CI/CD and automation expertise.
Splunk, Elastic Stack, Linux, Windows, Virtual Desktop Infrastructure (VDI), Terraform, CloudFormation, Ansible, Docker, Kubernetes, AWS, Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, Prometheus, Grafana
3d
Save
Mark Applied
Hide
VP of Systems Engineering - USA
United States or Austin
RemoteFull Time
Bankjoy
Bankjoy: Digital banking platform for community banks and credit unions.
10+ YOE4+ MgmtRequires 10+ years of engineering experience and 4+ years of leadership managing QA, DevOps, or platform reliability teams, plus automation, release management, cloud, observability, and regulated-finance experience.
Angular, Swift, Kotlin, Playwright, Cypress, Appium, Maestro, AWS, CI/CD, SRE
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Antonio, Texas, United States
OnsiteFull Time
iHeartMedia
iHeartMediaNASDAQ: IHRT: Provides radio broadcasting, podcasting, and digital audio streaming services.
2+ YOEMaster's in CS/CE/EE/IS (or equivalent) plus 24 months as a Software Engineer or related; experience leading SRE/DevOps teams; strong communication, delegation, troubleshooting, and improvement skills.
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Atlanta or Austin
$100k-$115k/yr OnsiteFull Time
Atlanticus
AtlanticusNASDAQ: ATLC: Provide credit products and financial services to underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog, Splunk, CI/CD, databases, Python or Bash, Linux, networking, and incident management experience.
Amazon Web Services (AWS), Amazon EKS, Amazon EC2, ALB, NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, Amazon VPC, Java, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, AI-assisted development tools, agentic AI systems
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Austin or Atlanta
$100k-$115k/yr OnsiteFull Time
Atlanticus
AtlanticusNASDAQ: ATLC: Provides credit cards and lending solutions for underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog or Splunk, CI/CD, Python or Bash, Linux, cloud troubleshooting, and incident management experience.
AWS, Amazon EKS, Amazon EC2, ALB/NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, VPC, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, Java, JVM
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
2w
Save
Mark Applied
Hide
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Austin, Texas, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Manage and operate a massive multi-cloud data platform, run incident response, provide hands-on support to internal teams, and partner with developers to keep services reliable across AWS, GCP, and on-prem Kubernetes.
Spark, Flink, Airflow, Ray, Notebooks, Kubernetes, AWS, GCP
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald
1mo
Save
Mark Applied
Hide
Principal, Product RE - Grid Scale systems, BESS
Austin, Texas, United States
$130k-$183k/yr OnsiteFull Time
Enphase Energy
Enphase EnergyNASDAQ: ENPH: Manufacturer of microinverter-based solar and energy storage systems.
13+ YOEBachelor's (B.E.) with 15+ years or Master's with 13+ years in power/product engineering; deep expertise in inverter systems, thermal cooling, high-voltage integration, simulation, DFMEA/AFMEA, root-cause analysis, Python and embedded software, and technical leadership.
Python
3w
Save
Mark Applied
Hide
Technical Operations Manager, Third-Party Data Centers
Austin, Texas, United States
$138k-$201k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
9+ YOE5+ MgmtAssociate's degree or equivalent experience;9 years in electrical/mechanical/HVAC/controls;5 years program/project management in mission-critical facilities;experience with reliability and resource forecasting.