113 customer reliability engineer jobs at 89 companies in United States

1mo
Save
Mark Applied
Hide
Customer Reliability Engineer
San Francisco, California, United States
$204k-$284k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience supporting large-scale compute customers, debugging distributed systems across stacks, strong incident communication, and driving engineering fixes.
InfiniBand, RoCE, Slurm, Kubernetes, NCCL
2mo
Save
Mark Applied
Hide
Customer Reliability Engineer
Chicago, Illinois, United States
$103k-$159k/yr HybridFull Time
iManage
iManage: AI knowledge-work software helping legal, accounting, and financial-services organizations manage documents, email, and governed knowledge.
3+ YOE3+ years in technical escalation/Support/CRE/SRE roles; incident response for P1/P2; experience with distributed cloud services; SQL, Python, Bash/Shell, PowerShell, REST APIs; AKS/Azure services; Splunk, Grafana, Kibana, Prometheus; strong communication and troubleshooting.
SQL, Python, Bash/Shell, PowerShell, REST APIs, Azure Kubernetes Service (AKS), Azure services, Splunk, Grafana, Kibana, Prometheus
2w
Save
Mark Applied
Hide
Customer Reliability Engineer - Airflow
United States or New York City or Boston or San Francisco
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
4+ YOERequires data engineering background, 4 years with Python, 1 year administering Airflow and creating DAGs, Kubernetes, Docker, containers, cloud provider experience, troubleshooting, communication, autonomy, and mentoring experience.
Apache Airflow, Python, Kubernetes, Docker, AWS, GCP, Azure, Zoom, SQL, PostgreSQL, Databricks, Snowflake, Redshift, dbt
2mo
Save
Mark Applied
Hide
Customer Reliability Engineer Intern (PT)
Port Arthur, Texas, United States
OnsitePart Time
John Crane
John Crane: Industrial manufacturer and service provider of mechanical seals, couplings and filtration systems for energy and process industries.
Currently enrolled mechanical engineering student (≥2 years), good academic standing, mechanical aptitude, proficient with Microsoft Office and some CAD, effective communication skills, able to work in a production environment, must be ≥18 and attend school in Port Arthur, TX.
Microsoft Word, Microsoft Excel, Microsoft PowerPoint, CAD
2w
Save
Mark Applied
Hide
Reliability Engineer Senior
Milwaukee, Wisconsin, United States
RemoteFull Time
Regal Rexnord
Regal RexnordNYSE: RRX: Global industrial leader in power transmission and motion control.
10+ YOEBachelor's degree preferred in engineering, 10+ years of industrial reliability and manufacturing experience, predictive maintenance expertise, customer-facing skills, and Certified Vibration Analyst CAT II certification.
RCA, OEE, HSA, FSA, AD&D
3mo
Save
Mark Applied
Hide
Reliability Engineer
Philadelphia, Pennsylvania, United States
OnsiteFull Time
Proscia
Proscia: AI-powered digital pathology software serving diagnostic laboratories, pathologists, scientists, and biopharmaceutical research teams.
Hands-on container deployments, Linux systems, problem solving, and customer-facing collaboration.
1mo
Save
Mark Applied
Hide
Staff Senior Reliability Engineer
United States
$120k-$165k/yr RemoteFull Time
Symbotic
SymboticNASDAQ: SYM: AI-powered robotics and software for warehouse automation.
8+ YOE8+ years supporting production reliability, 5+ years leading technical RCA/Problem Management, strong troubleshooting and data analysis, customer-facing experience, bachelor's in technical field or equivalent.
Kubernetes, VMware, Linux, Grafana, Prometheus, Zabbix, GitLab, Power BI, Tableau
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise software providing cloud-based workflow automation platforms.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
1mo
Save
Mark Applied
Hide
Senior Safety & Reliability Engineer
Everett, Washington, United States
$112k-$178k/yr OnsiteFull Time
Sigma Design
Sigma Design: Product development and engineering firm that designs, builds, tests, manufactures, and staffs solutions for clients worldwide.
7+ YOE7+ years in safety and reliability engineering, experience with aerospace safety standards, FMEA/FTA/FRACAS, systems engineering, strong communication and customer-facing skills.
PTC Windchill, Windchill Risk & Reliability, CAFTA Fault Tree, Relex Fault Tree, DOORs, Minitab, Jira, ADO
1mo
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: -packaged optics providing connectivity for hyperscale AI infrastructure.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
1mo
Save
Mark Applied
Hide
Senior Inference Reliability Engineer
San Mateo, California, United States
OnsiteFull Time
Parasail
Parasail: AI inference cloud providing production-ready open-model services for AI-native startups.
5+ YOE5+ years production engineering experience operating customer-facing systems; strong SRE and production diagnostics skills; Kubernetes, Linux, distributed systems, and software engineering proficiency; ability to lead incident response and build observability.
Kubernetes, Linux, Python, Go, Java, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Hyderabad or New York City or Chicago or London or Singapore or Tokyo or Hong Kong or Europe or United States or Asia-Pacific
HybridFull Time
Pico
Pico: Private financial-markets technology providing trading infrastructure, connectivity, market data, software, and analytics to institutions.
Bachelor's degree or relevant experience; financial markets technology experience; Linux, networking, computer architecture, programming or scripting, customer service, communication, and collaborative teamwork skills.
Linux, Python, C, C++, Java
1mo
Save
Mark Applied
Hide
Senior Systems Reliability Engineer
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Enterprise hybrid multi-cloud computing and software provider.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud and strong customer-facing troubleshooting and communication skills.
VMware ESXi, VMware, Linux, DevOps, Cloud, Citrix, Microsoft
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Private software providing an internal-tools development platform for business and enterprise teams.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3w
Save
Mark Applied
Hide
Site Reliability Engineer - Public Sector
United States
$140k-$170k/yr RemoteFull Time
Blitzy
Blitzy: Private AI software platform that autonomously builds and modernizes enterprise software for development teams.
3+ YOEU.S. citizen with 3+ years SRE/DevOps experience, strong Kubernetes, IaC (Terraform/Pulumi), cloud platform, observability, scripting (Python/Go/Bash), and customer-facing communication.
Kubernetes, Terraform, Pulumi, Python, Go, Bash
1mo
Save
Mark Applied
Hide
Fellow Software Engineer — AI Performance & Reliability
San Jose or Bellevue
$235k-$402k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
PhD or equivalent in AI/ML/CS, strong software engineering, experience profiling and optimizing ML models and AI workloads, proficiency in Python/C++, ML frameworks, customer-facing troubleshooting and performance analysis.
Python, C++, PyTorch, TensorFlow, JAX, ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL
1w
Save
Mark Applied
Hide
R&D Quality & Reliability Engineer
San Jose, California, United States
$122k-$195k/yr OnsiteFull Time
Broadcom Inc.
Broadcom Inc.NASDAQ: AVGO: Global technology leader in semiconductors and infrastructure software.
3+ YOEBachelor's plus 8 years, master's plus 6 years, or PhD plus 3 years in a relevant technical field; silicon photonics or semiconductor reliability expertise, statistical analysis, failure analysis, and customer-facing experience required.
JMP, Python, R, MATLAB
1mo
Save
Mark Applied
Hide
Senior Manager, Foundry Customer Quality & Reliability
San Jose, California, United States
$189k-$301k/yr OnsiteFull Time
Samsung Semiconductor
Samsung SemiconductorKorea Exchange (KRX): 005930: Global leader in semiconductor solutions including memory, system LSI, and foundry services.
8+ YOE15+ years industry experience (Bachelor+), minimum 8 years foundry-related experience, strong communication, analytical and problem-solving skills, QMS and fab process knowledge, ability to lead cross-functional teams.
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Reliability Engineering & SDLC Governance
Carlsbad or Germantown or San Jose or San Francisco or New York City
$165k-$261k/yr OnsiteFull Time
Viasat
ViasatNASDAQ: VSAT: Global communications providing satellite connectivity and defense solutions.
8+ YOE8+ years software engineering experience with SDLC governance, SRE/DevOps knowledge, observability, incident analysis, and cross-team influence to improve reliability and customer experience.
Prometheus, Grafana, OpenTelemetry, Datadog, AS9115
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Golang, Kubernetes)
Calgary or British Columbia or United States
RemoteFull Time
Mirantis
Mirantis: Kubernetes-native AI infrastructure and cloud software serving enterprises with open-source orchestration, automation, security, and support.
5+ YOERequires 5+ years in DevOps or software development, Kubernetes and cloud infrastructure expertise, Golang exposure, Linux, distributed systems, CI/CD, networking, storage, and customer-facing communication.
Kubernetes, Golang, Python, JavaScript, Linux, CI/CD, OpenStack, Rancher, OpenShift, VMware, MIG/vGPU, RDMA/RoCE, InfiniBand, NVLink, DCGM, NVIDIA AI Enterprise