50 reliability engineering manager jobs at 33 companies in Blanco, TX

2w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
Reston or Austin
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOE1+ Mgmt8+ years in software engineering or infrastructure (or fewer with relevant degrees), 3–5 years automation/programming experience, data analysis skills, 1 year leadership experience preferred.
2mo
Save
Mark Applied
Hide
Principal Architect, Site Reliability Engineering
Southlake or Austin
$221k-$252k/yr OnsiteFull Time
Charles Schwab
Charles SchwabNYSE: SCHW: Financial services, brokerage, and investment management provider.
5+ YOE3+ Mgmt5+ years in SRE with 3+ years in architect/leadership; design scalable, fault-tolerant systems; strong observability; CI/CD; postmortems; SRE leadership.
Prometheus, Grafana, Datadog, Splunk
2w
Save
Mark Applied
Hide
Engineering Manager, Support & Stability
Columbus or Boston or New York City or Chicago or Austin or Los Angeles
$150k-$190k/yr RemoteFull Time
Loop Returns
Loop Returns: Software platform automating e-commerce returns and post-purchase experiences.
3+ YOE3+ years engineering management experience, experience with platform reliability/system health, AI agent adoption, technical depth in high-risk domains, ability to manage and grow engineers.
PHP, Laravel, Vue.js, MySQL, DynamoDB, Kubernetes, AWS, Jira, Claude, Cursor, Datadog, Snowflake
1mo
Save
Mark Applied
Hide
Manager, Engineering Observability
Austin, Texas, United States
$169k-$232k/yr HybridFull Time
Procore
ProcoreNYSE: PCOR: Cloud-based construction management software for projects and teams.
7+ YOE2+ Mgmt7+ years as a software engineer, 2+ years managing teams; hands-on backend/distributed systems experience; experience with observability tools (Datadog, OpenTelemetry, Prometheus/Grafana, Honeycomb, SumoLogic, Bugsnag); SRE/reliability background; strong leadership and roadmap skills.
Datadog, Honeycomb, SumoLogic, OpenTelemetry, Bugsnag, Prometheus, Grafana
2w
Save
Mark Applied
Hide
Engineering Manager, Data Feeds
New York City or Boston or Chicago or Salt Lake City or Austin or Montreal or Canada or Washington or Dallas or United States or Vancouver or Toronto or Charlotte or Denver
$129k-$304k/yr RemoteFull Time
Chainlink Labs
Chainlink Labs: Building decentralized oracle networks for blockchain smart contracts.
Deep experience building and operating production blockchain infrastructure, strong engineering judgment, leadership and coaching experience, stakeholder management, and ability to deliver reliable, scalable systems.
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
2d
Save
Mark Applied
Hide
Reliability Engineer, Mechanical, NA
San Antonio or Ashburn or North America or Europe or Asia
$125k-$135k/yr HybridFull Time
Vantage Data Centers
Vantage Data Centers: Provides hyperscale data center campuses for cloud and AI providers.
2+ YOEBachelor's degree in electrical or mechanical engineering preferred, 2–3 years in critical facility operations and maintenance, strong project management and communication skills, and willingness to travel up to 25%.
Root-Cause Failure Analysis, Facility Acceptance Testing
1w
Save
Mark Applied
Hide
Technical Lead Manager - Production Engineering
Austin, Texas, United States
OnsiteFull Time
Apptronik
Apptronik: Designs and manufactures humanoid robots for industrial automation.
8+ YOE2+ MgmtHands-on technical manager to build and lead production engineering for robot deployment, commissioning, and fleet reliability; 8+ years engineering or 4+ years humanoid experience, 2+ years managerial experience, willing to travel up to 25%.
Linux, DDS, UDP, TCP, Ansible, C++, Python, ROS 2, TypeScript, Kubernetes, React, CSS
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
2K
2KNASDAQ: TTWO: Publishes and develops global video game franchises and entertainment.
5+ YOE5+ years SRE/platform engineering experience, deep Kubernetes (EKS/GKE), Terraform/Pulumi and GitOps, observability with Prometheus/Grafana/Datadog, production coding in Go/Python/TypeScript, Linux and networking expertise, incident management.
Terraform, Pulumi, ArgoCD, Flux, Kubernetes, EKS, GKE, Istio, Cilium, Helm, Terragrunt, Prometheus, Grafana, Datadog, OpenTelemetry, GitHub Actions, Jenkins, Go, Python, TypeScript, PasswordState, 1Password, AWS Secrets Manager, OPA/Gatekeeper, AWS, GCP, VMware, Ansible, Puppet, AWS Systems Manager
1mo
Save
Mark Applied
Hide
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems
Austin or Seattle or Cupertino
$208k-$282k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
10+ YOE5+ Mgmt10+ years reliability/quality engineering experience with server or high-volume electronics, 5+ years people management, bachelor's degree in a relevant field, experience with root-cause analysis, quality systems, and multi-vendor manufacturing.
HALT, HTOL, thermal cycling, QRV, DFMEA, SPC, FMEA, 8D, DOE, Weibull analysis
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Antonio, Texas, United States
OnsiteFull Time
Infosys
InfosysNYSE: INFY: Provides IT consulting, software development, and business outsourcing services.
Experience with reliability engineering, Terraform IaC, observability (Datadog), disaster recovery, vulnerability management, CI/CD (Harness, Helm), plus bachelor's degree or equivalent experience.
Terraform, Datadog, Harness, Helm Charts, AWS
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer (SRE)
San Antonio, Texas, United States
HybridFull Time
IPSecure
IPSecure: Provides cybersecurity and incident response services for government agencies.
7+ YOE7+ years SRE/DevOps experience, active Secret clearance, cloud/DevOps certification (AWS), Splunk or Elastic Stack experience, Linux/Windows administration, CI/CD and automation expertise.
Splunk, Elastic Stack, Linux, Windows, Virtual Desktop Infrastructure (VDI), Terraform, CloudFormation, Ansible, Docker, Kubernetes, AWS, Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, Prometheus, Grafana
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Antonio, Texas, United States
OnsiteFull Time
iHeartMedia
iHeartMediaNASDAQ: IHRT: Provides radio broadcasting, podcasting, and digital audio streaming services.
2+ YOEMaster's in CS/CE/EE/IS (or equivalent) plus 24 months as a Software Engineer or related; experience leading SRE/DevOps teams; strong communication, delegation, troubleshooting, and improvement skills.
6d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Atlanta or Austin
$100k-$115k/yr OnsiteFull Time
Atlanticus
AtlanticusNASDAQ: ATLC: Provide credit products and financial services to underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog, Splunk, CI/CD, databases, Python or Bash, Linux, networking, and incident management experience.
Amazon Web Services (AWS), Amazon EKS, Amazon EC2, ALB, NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, Amazon VPC, Java, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, AI-assisted development tools, agentic AI systems
6d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Austin or Atlanta
$100k-$115k/yr OnsiteFull Time
Atlanticus
AtlanticusNASDAQ: ATLC: Provides credit cards and lending solutions for underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog or Splunk, CI/CD, Python or Bash, Linux, cloud troubleshooting, and incident management experience.
AWS, Amazon EKS, Amazon EC2, ALB/NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, VPC, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, Java, JVM
2d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
3w
Save
Mark Applied
Hide
Electric Reliability Compliance Analyst Senior - Operations & Planning
Austin, Texas, United States
HybridFull Time
City of Austin
City of Austin: Providing municipal services and public infrastructure to the Austin community.
4+ YOEBachelor's in Business, Engineering, or related plus 4 years energy/electric utility experience; knowledge of NERC/FERC/ERCOT/PUCT reliability requirements; audit, reporting and training experience; ability to travel and obtain required clearances.
1w
Save
Mark Applied
Hide
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Austin, Texas, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Manage and operate a massive multi-cloud data platform, run incident response, provide hands-on support to internal teams, and partner with developers to keep services reliable across AWS, GCP, and on-prem Kubernetes.
Spark, Flink, Airflow, Ray, Notebooks, Kubernetes, AWS, GCP
2w
Save
Mark Applied
Hide
Senior AI Ops & Incident/Site Reliability Engineer
Austin or Plano or Fort Mill or Boston or New York City or Tempe or San Diego
$94k-$171k/yr HybridFull Time
Perficient
Perficient: Provides digital transformation and AI consulting for global enterprises.
8+ YOEBachelor's degree or equivalent,8+ years in IT operations/SRE or production support,5+ years leading incident management,experience with Dynatrace,ServiceNow,cloud platforms and AIOps implementations.
Dynatrace, ServiceNow, AWS, Azure, GCP
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald

Explore Jobs

Expand Your Job Search