43 reliability engineering manager jobs at 31 companies in Wills Point, TX

2mo
Save
Mark Applied
Hide
Operations Engineering Manager, Fleet Reliability
Dallas or Bellevue
$143k-$191k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOE2+ Mgmt7+ years in software or infrastructure engineering with 2+ years leadership; SRE fundamentals, incident management, observability, change management; strong automation and people development skills.
4w
Save
Mark Applied
Hide
Director Reliability, Automation & Performance Engineering
Dallas, Texas, United States
$135k-$225k/yr HybridFull Time
Catalyst Brands
Catalyst Brands: Operates a portfolio of retail brands including JCPenney.
12+ YOE5+ Mgmt12+ years engineering leadership with 5+ years leading reliability, performance or automation functions; deep SRE, performance engineering, automation, cloud and observability experience; BA/BS preferred.
AWS, Azure, Kubernetes/EKS, Jenkins, Kafka, Dynatrace, Splunk, Datadog, New Relic, Catchpoint, CloudWatch, ELK
2mo
Save
Mark Applied
Hide
Software Engineering Manager - Site Reliability Center
Pittsburgh or Cleveland or Birmingham or Dallas or Denver or Phoenix
$100k-$204k/yr OnsiteFull Time
PNC Financial Services
PNC Financial ServicesNYSE: PNC: Provides banking, lending, and investment services to customers.
5+ YOE3+ MgmtLead SRE teams to ensure reliability, incident and change management, production support, automation, observability, and performance; 5+ years related experience with 3+ years management; hands-on with monitoring, cloud/infrastructure, databases and automation.
Dynatrace, BigPanda, Logscale, Linux, Windows, Oracle, SQL, MongoDB, Cassandra, Elasticsearch, Redis, MQ, Kafka, OCP, ShiftPlanning, ETL
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering – Paylo Platform
Alpharetta or Temple or Dallas or Houston
HybridFull Time
PDI Technologies
PDI Technologies: Software solutions for convenience retail and petroleum wholesale operations.
8+ YOE4+ MgmtRequires 8+ years in SRE, DevOps, or infrastructure engineering, 4+ years leading people, manager-management experience, and hands-on AWS, Azure, Kubernetes, Helm, Argo, Terraform/OpenTofu, Jenkins, and Datadog expertise.
AWS, Microsoft Azure, Kubernetes, Helm, Argo CD, Argo Workflows, Terraform, OpenTofu, Jenkins, Datadog, Kafka, SQS, SNS, PagerDuty
1mo
Save
Mark Applied
Hide
Quality & Reliability Manager
Dallas, Texas, United States
OnsiteFull Time
Aligned Data Centers
Aligned Data Centers: Provides scalable data center infrastructure and colocation services.
3+ YOEBachelor's in engineering or equivalent, 3+ years commissioning/reliability experience in mission-critical facilities, familiarity with Level 4/5 testing, Cx Alloy, BAS/EPMS/SCADA, MS365, strong analytical and communication skills, up to 20% travel, US work authorization.
Cx Alloy, BAS, EPMS, SCADA, Microsoft 365, Microsoft Excel, Microsoft Teams, Microsoft SharePoint, Microsoft Word, Microsoft PowerPoint
4w
Save
Mark Applied
Hide
Director Reliability, Automation & Performance Engineering
Dallas, Texas, United States
$135k-$315k/yr HybridFull Time
Catalyst Brands
Catalyst Brands: Operates a portfolio of diverse retail clothing and apparel brands.
12+ YOE5+ Mgmt12+ years progressive engineering leadership with 5+ years leading reliability/performance/automation; deep SRE, performance, and automation experience; cloud and observability tool expertise; ability to lead distributed engineering teams.
AWS, Azure, Kubernetes, EKS, Jenkins, Kafka, Dynatrace, Splunk, Datadog, New Relic, Catchpoint, CloudWatch, ELK
1mo
Save
Mark Applied
Hide
Engineering Manager, Data Feeds
New York City or Boston or Chicago or Salt Lake City or Austin or Montreal or Canada or Washington or Dallas or United States or Vancouver or Toronto or Charlotte or Denver
$129k-$304k/yr RemoteFull Time
Chainlink Labs
Chainlink Labs: Building decentralized oracle networks for blockchain smart contracts.
Deep experience building and operating production blockchain infrastructure, strong engineering judgment, leadership and coaching experience, stakeholder management, and ability to deliver reliable, scalable systems.
2mo
Save
Mark Applied
Hide
Reliability Engineer — Advanced Thermal Management
Saxonburg or Newark or Dallas
$115k-$160k/yr OnsiteFull Time
Coherent
CoherentNYSE: COHR: Manufactures lasers, engineered materials, and optical networking components.
5+ YOE5+ years reliability engineering experience in semiconductor packaging, electronics cooling, or related fields; BS/MS in engineering or equivalent; experience with accelerated life testing, failure analysis, and industry reliability standards.
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Idaho or Plano
$117k-$209k/yr RemoteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
7+ YOEU.S. citizen with 7+ years SRE/platform/cloud experience, strong reliability engineering skills (SLOs/SLIs, observability, incident management), cloud experience (AWS/Azure), automation using Python/Go/Java/PowerShell/Bash, and ability to operate production services in regulated GovCloud environments.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Infrastructure as Code, CI/CD, Splunk, Dynatrace, Datadog, CloudWatch, Kubernetes
1d
Save
Mark Applied
Hide
Director, Technical Program Manager (Resiliency and Reliability Engineering)
McLean or Richmond or New York City or Plano or San Francisco
$210k-$287k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
7+ YOEBachelor's degree and 7+ years managing technical programs required. Preferred experience includes distributed systems, cloud computing, reliability engineering, Agile delivery, complex programs, and regulated environments.
Agile, cloud computing, distributed computing, site reliability engineering, observability, chaos engineering
1d
Save
Mark Applied
Hide
Director, Technical Program Manager (Resiliency and Reliability Engineering)
McLean or San Francisco or Richmond or New York City or Plano
$210k-$287k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
7+ YOE5+ MgmtBachelor's degree and 7+ years managing technical programs required. Preferred experience includes distributed systems, cloud computing, reliability engineering, Agile delivery, large complex programs, and regulated environments.
1w
Save
Mark Applied
Hide
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Dallas, Texas, United States
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
6+ YOERequires 6+ years in site reliability engineering, programming in Java, Python, or Go, cloud and container expertise, IaC and configuration management skills, Linux and distributed systems knowledge, and advanced monitoring experience.
Java, Python, Go, AWS, GCP, Docker, Kubernetes, Terraform, CloudFormation, Puppet, Chef, Ansible, Prometheus, Grafana, ELK, Datadog, PagerDuty, Jenkins, GitLab, Maven, Elastic Search, GCP Big Query, Kafka, Linux, Infrastructure as Code (IaC), Prompt Engineering, Retrieval-Augmented Generation (RAG), CI/CD
1d
Save
Mark Applied
Hide
Director, Technical Program Manager (Resiliency and Reliability Engineering)
McLean or Richmond or New York City or Plano or San Francisco
$210k-$287k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
7+ YOEBachelor's degree and 7+ years managing technical programs required. Preferred: distributed systems, cloud, SRE, resilience engineering, Agile delivery, complex program leadership, and regulated-environment experience.
Agile
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Dallas, Texas, United States
HybridFull Time
Longbridge Group
Longbridge Group: AI-powered online brokerage and global investment platform.
5+ YOE5+ years in SRE, DevOps, or production engineering; AWS/GCP/Azure, Docker, Kubernetes, Linux, CI/CD, incident management, distributed systems, and programming experience required.
Terraform, Ansible, Helm, Kubernetes, Prometheus, AWS, GCP, Docker, Python, Go, Linux, CI/CD
1mo
Save
Mark Applied
Hide
Maintenance and Reliability Technical Director
Dallas, Texas, United States
OnsiteFull Time
Danone
DanoneEuronext Paris: BN: Global food and beverage producing dairy and plant-based goods.
8+ YOEBachelor's in Engineering required; 8+ years manufacturing operational and management experience; program and capital project management; experience with AutoCAD, Microsoft Project, SAP PM/CMMS; strong leadership, reliability and TPM knowledge.
AutoCAD, Microsoft Project, SAP PM, CMMS
1w
Save
Mark Applied
Hide
Site Reliability Engineer Intern
Dallas, Texas, United States
OnsiteFull Time, Internship
Copart
CopartNASDAQ: CPRT: Provides global online vehicle auction and remarketing services.
Experience with production incident management, Linux, Windows, scripting, automation, monitoring, troubleshooting, and observability tools. Strong communication and analytical skills required; programming, virtualization, and cloud experience preferred.
Python, Ansible, Datadog, Kubernetes, Linux, Windows, VMware vSphere, Unix, AWS, GCP
2mo
Save
Mark Applied
Hide
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms
Jersey City or Dallas
$171k-$260k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, strong SLI/SLO/SLA and observability knowledge, experience with Grafana/Dynatrace/Prometheus/Datadog/Splunk, distributed systems expertise, mentoring and leadership experience, familiarity with safe AI usage in operations.
Grafana, Dynatrace, Prometheus, Datadog, Splunk, AWS, Databricks, Spark, Glue, MapReduce, Docker, Kubernetes, Terraform, Python
2w
Save
Mark Applied
Hide
IT Site Reliability Engineer — API Management Platforms
Dallas, Texas, United States
OnsiteFull Time
Texas Instruments
Texas InstrumentsNASDAQ: TXN: Designs and manufactures semiconductors and integrated circuits.
3+ YOEBachelor's degree or equivalent practical experience and 3+ years managing Apigee or similar API platforms. Requires platform administration, automation, troubleshooting, monitoring, and incident response expertise.
Apigee, Apigee Edge, Linux, Unix, RHEL, Rocky Linux, Python, Bash, PowerShell, REST API, OAuth 2.0, JWT, Elastic, Elasticsearch, Logstash, Kibana, Splunk, Prometheus, Grafana, TCP/IP, DNS, TLS/SSL, Terraform, Ansible, Chef, Puppet, Jenkins, GitLab CI, GitHub Actions, JFrog Artifactory, Sonatype Nexus, Docker, Kubernetes, AWS, GCP, Azure, Cassandra, ZooKeeper, Qpid, PostgreSQL
1mo
Save
Mark Applied
Hide
Total Cost of Ownership Product Manager - Blasthole Drills
Garland, Texas, United States
HybridFull Time
Epiroc
EpirocNasdaq Stockholm: EPI A: Manufactures machinery and tools for mining and construction industries.
5+ YOEBachelor's degree in engineering, mining, business or related; 5+ years in mining equipment/product management/aftermarket; strong lifecycle cost, reliability, analytical, and stakeholder-management skills; fluent English.
1w
Save
Mark Applied
Hide
Data Center Controls Manager - Operations
Elk Grove Village or Tolleson or Denver or Dallas or Lockhart
$150k-$175k/yr HybridFull Time
Prime Data Centers
Prime Data Centers: Developer and operator of hyperscale and wholesale data centers.
Manages BMS, EPMS, DCIM, and related controls across data centers; leads reliability, troubleshooting, commissioning, documentation, contractors, and efficiency improvements.
Building Management Systems (BMS), Electrical Power Monitoring Systems (EPMS), Data Center Infrastructure Management (DCIM), Building Automation Systems (BAS)

Explore Jobs

Expand Your Job Search