43 software reliability engineer jobs at 24 companies in Lathrop, CA
4w
Save
Mark Applied
Hide
4w
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems; strong software engineering in Go or Python; SLO program experience; chaos engineering and failure-injection experience; ability to lead incident response and influence cross-team.
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Senior Software Engineer, Site Reliability Engineering
New York or San Ramon or Reno
$153k-$210k/yrHybridFull Time
Ridgeline: Cloud-native platform for investment management operations.
3+ YOE3–6 years SRE/DevOps experience, 2+ years on AWS, proficiency with Terraform, observability, CI/CD, Python/Go/Bash, incident response, and strong communication and troubleshooting skills.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$123k-$317k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS or related, strong software engineering and Linux knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; SRE and AI-ops experience preferred.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
MaxInsights: Provides robot data collection for physical AI development.
Experience owning production pipeline systems, workflow orchestration, schedulers, reliability, debugging across infra and data dependencies, strong engineering judgment and clear communication.
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOE5+ years professional software development with production programming, system design/architecture, scalability and reliability expertise; experience mentoring or leading engineering efforts; BS in CS preferred.
Lead Software Engineer - Cloud Storage (GO , Python automation)
San Jose, California, United States
$215k-$245k/yrOnsiteFull Time
NetAppNASDAQ: NTAP: Sells enterprise data storage and cloud management software.
8+ YOE8+ years software/systems engineering experience; 3+ years in data management/storage; strong Go and Python; experience with Kubernetes, cloud (Azure/AWS/GCP), distributed systems, APIs, observability, and production reliability.
Sr. Site Reliability Engineering (Agentic Builders Experience team)
San Jose, California, United States
$159k-$302k/yrOnsiteFull Time
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years in software or infrastructure engineering with deep SRE skills, experience with high-scale production systems, AWS or Azure, microVM/container tech, and coding in Go/Python/Java.
XperiNYSE: XPER: Develops entertainment and audio technology for consumer electronics and automotive.
Design, develop, and maintain scalable Java/goLang/Elixir microservices on AWS; collaborate with cross-functional teams to build cloud-native, performant, reliable, and secure applications.
eBayNASDAQ: EBAY: Global online marketplace for buying and selling diverse products.
5+ YOE5+ years building and operating production backend systems; proficiency with Java, Nodejs, TypeScript or Python; experience designing and shipping AI-enabled services and reliable workflows; strong communication and technical judgment.
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOE5+ years SRE or software development experience; expertise in cloud-based systems, distributed systems, databases (SQL/NoSQL), Kubernetes, and strong communication skills.
San Francisco or San Jose or Seattle or Sacramento
RemoteFull Time
Unstructured: Enterprise data transformation for LLM and AI applications.
8+ YOE7-10+ years in production systems; cloud-native and distributed architectures; auth systems (OAuth/OIDC/SAML/JWT/RBAC/ABAC); enterprise identity integration; mentoring; focus on performance, reliability, and scalable design.
Platform Reliability, Availability, Serviceability (RAS), and Manageability Software Architect. Principal Engineer
Santa Clara or Austin
$212k-$318k/yrOnsiteFull Time
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
6+ YOEExpertise in ARM/ARM64, RAS and manageability, Linux kernel, DDR, PCIe, I2C/SPI/MDIO, 6+ years software experience (varies by degree), proficiency in C/C++/Java/Python, strong documentation and communication skills.
Anyware Robotics: Develops autonomous mobile robots for warehouse and logistics automation.
3+ YOE3+ years building/debugging production software for physical or real-time systems; strong C++ and Python; experience with reliability, incident response, regression testing, and system-level debugging.
15+ YOE7+ Mgmt15+ years software engineering experience with 7+ years leading large distributed engineering organizations; experience architecting high-scale cloud SaaS platforms (AWS/GCP/Azure), distributed databases, microservices, and production reliability; strong leadership and customer-first orientation.
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident management ownership, and absolute fluency in Mandarin and English.