5+ YOEBachelor's or equivalent,5+ years deploying/supporting distributed systems,cloud and on-prem storage,Kubernetes (EKS/AKS/RKS),CI/CD automation,observability,backup/recovery,Python/NodeJS/Java and scripting.
Shackelford County or Texas or Atlanta or Abilene or Dallas or Phoenix or Ashburn or Wisconsin
HybridFull Time
Vantage Data Centers: Provides hyperscale data center campuses for cloud and AI providers.
2+ YOEMechanical reliability engineer for data center cooling systems; 2–3 years critical facility experience preferred; bachelor’s degree preferred; experience with commissioning, maintenance program design, RCA, and technical support.
Manassas or Sterling or Portland or Chicago or Dallas Fort Worth
OnsiteFull Time
STACK Infrastructure: Developer and operator of sustainable wholesale data center infrastructure.
5+ YOE5–8 years in critical infrastructure; strong fluency in electrical systems; RCA/forensic troubleshooting; bachelor’s in engineering or equivalent.
Power distribution equipment, Waveform analysis, Fault analysis tools
New York or Los Angeles or Chicago or Houston or Tempe or Philadelphia or Dallas or North Miami Beach or Denver
$111k-$145k/yrHybridFull Time
JacobsNYSE: J: Global provider of professional engineering and technical services.
4+ YOEBachelor's in electrical engineering, 4+ years power-system and reliability analysis experience, familiarity with reliability indicators and NERC/FERC standards, strong analytical and communication skills.
OptimumNYSE: OPTU: Provides broadband, television, and mobile connectivity services to customers.
2+ YOEBachelor's in telecommunications/computer engineering, 2+ years systems or mobile network operations experience; deep Unix/Linux administration, GCP experience, Terraform/Ansible, Python/Go scripting, SAN/NAS storage protocols, observability tooling.
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
5+ YOE5+ years systems engineering experience, 4+ years SRE/production support, observability and incident leadership, GCP and data-pipeline familiarity, strong communication and mentorship skills.
Grafana, Splunk, AppDynamics, Google Cloud Platform (GCP), Airflow, Cloud Composer, BigQuery, CI/CD, Infrastructure as Code
American Heart Association: Non-profit organization funding cardiovascular research and public health education.
5+ YOEBachelor's degree or equivalent, 5+ years relevant experience, expertise in multi-cloud operations, IAM, security, automation, DevOps, and infrastructure reliability.
American Heart Association: Nonprofit organization dedicated to fighting heart disease and stroke.
5+ YOEBachelor's degree or equivalent experience, minimum 5 years relevant experience, strong expertise in multi-cloud operations, IAM, security, automation, and infrastructure as code.
Azure, AWS, GCP, Oracle, Entra ID, Structured Query Language (SQL)
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms
Jersey City or Dallas
$171k-$260k/yrOnsiteFull Time
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, strong SLI/SLO/SLA and observability knowledge, experience with Grafana/Dynatrace/Prometheus/Datadog/Splunk, distributed systems expertise, mentoring and leadership experience, familiarity with safe AI usage in operations.
onsemi: Designs and manufactures semiconductor solutions for power and sensing.
3+ YOE3+ years semiconductor lab/reliability testing experience; Associate's in Electronics or related; hands-on with ESD simulators and Latchup systems; familiarity with JEDEC/ESDA standards; able to read schematics and use oscilloscopes, DMMs, curve tracers.
Vizient: Provides performance improvement and supply chain services to hospitals
7+ YOE7+ years in quality engineering or testing, experience with AI/ML/LLM-enabled systems, test automation and validation frameworks, strong analytical and communication skills, US work authorization required.
Charles SchwabNYSE: SCHW: Financial services, brokerage, and investment management provider.
5+ YOE3+ Mgmt5+ years in SRE with 3+ years in architect/leadership; design scalable, fault-tolerant systems; strong observability; CI/CD; postmortems; SRE leadership.
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
Extensive experience architecting and shipping production AI/LLM-powered systems, strong backend and distributed systems skills, production reliability and observability experience, and ability to mentor senior engineers.
San Francisco or Ottawa or Phoenix or Toronto or Los Angeles or Denver or Salt Lake City or Atlanta or Chicago or Houston or Portland or New York City or Vancouver or San Diego or Sacramento or Jacksonville or Seattle or Mexico City or Austin or Miami or Boston or Dallas or Charlotte
$141k-$267k/yrHybridFull Time
Scribd: Subscription-based digital library for e-books, audiobooks, and documents.
Significant backend engineering experience building and scaling distributed systems; strong coding in Ruby/Scala/Go/Python; expertise in reliability, observability, SLOs, and mentoring; cross-functional collaboration experience.
Match GroupNasdaq: MTCH: Provider of a global portfolio of online dating services.
5+ YOE5+ years onsite data center or systems engineering experience; Linux/Windows, AD/DNS/DHCP, HPE and Lenovo hardware, scripting (PowerShell,Bash,Python); able to lift 50 lbs; reliable for after-hours on-call; travel ~10%.
Triumph Financial: Provides financial and technology solutions for the transportation industry.
10+ YOE10+ years software engineering experience, 4+ years production ML, experience designing and operating distributed ML systems, model deployment and data pipeline expertise, strong reliability and production support skills.
Santa Barbara or San Diego or San Francisco or Denver or Dallas or Atlanta or Chicago or Washington, D.C.
HybridFull Time
AppFolioNASDAQ: APPF: Provides cloud-based property and investment management software.
Proven experience building production ML systems at scale, architectural leadership, training/fine-tuning LLMs, experience with LangChain/LangGraph and RAG patterns, AI safety/authorization, and production reliability discipline.