67 cloud reliability engineer jobs at 30 companies in Marina, CA

1w
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
2mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer
San Jose or Palo Alto
HybridFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
3+ YOE3-5+ years SRE/DevOps in public cloud; Bachelor's degree or equivalent; Python/Go/Java; Docker/Kubernetes; monitoring/observability tools; IaC; CI/CD.
Docker, Kubernetes, Prometheus, Grafana, ELK Stack, Datadog, Terraform, CloudFormation, Jenkins, GitHub Actions, ArgoCD, Python, Go, Java
4w
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
4w
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems; strong software engineering in Go or Python; SLO program experience; chaos engineering and failure-injection experience; ability to lead incident response and influence cross-team.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Cloud Reliability Intelligence
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent, 8+ years experience with data structures and algorithms, 3+ years leading distributed systems projects and in technical leadership; experience with full-stack architectures and LLM/Generative AI preferred.
Go, Java, TypeScript, Angular, LLMs, Generative AI
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1w
Save
Mark Applied
Hide
Senior Systems Reliability Engineer
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud and strong customer-facing troubleshooting and communication skills.
VMware ESXi, VMware, Linux, DevOps, Cloud, Citrix, Microsoft
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
2w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Berkeley Heights or Alpharetta or Sunnyvale
$128k-$216k/yr OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
5+ YOE5+ years production experience with AWS, Kubernetes, and Linux; strong Terraform, CI/CD (GitHub Actions), Docker, GitHub, RDBMS/Document storage, and scripting (Python/Bash/Node/Ruby); experience designing scalable cloud systems.
Amazon Web Services, Kubernetes, GitHub Actions, Terraform, New Relic, Dynatrace, Datadog, Docker, GitHub, Python, Bash, Node, Ruby on Rails
1mo
Save
Mark Applied
Hide
Principal, Site Reliability Engineer
Bentonville or Sunnyvale
$110k-$220k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
5+ YOEBachelor's in a related CS field plus 5+ years SRE or 7+ years SRE-related experience. Expertise in monitoring, root cause analysis, disaster recovery, cloud and containerization (Docker), coding in JavaScript and Python, CI/CD automation, and chaos testing.
Docker, JavaScript, Python
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Global E-Commerce
San Jose, California, United States
$213k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's or equivalent,5+ years SRE/infra experience,proficiency in Go/Python/Java,strong Linux,networking and distributed systems knowledge,cloud-native production experience.
Go, Python, Java, Linux
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Sunnyvale, California, United States
$170k-$196k/yr OnsiteFull Time
Illumio
Illumio: Provides zero-trust segmentation software to contain cyberattacks.
5+ YOE5+ years SRE experience with AWS/Azure, automation, scripting (Python/PowerShell/Go), CI/CD, containers, and cloud security.
AWS, Azure, PowerShell, Python, Go, CI/CD, Docker, Kubernetes, Azure DevOps, Jenkins, GitLab CI/CD
6d
Save
Mark Applied
Hide
Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, California, United States
$195k-$285k/yr HybridContract, Full Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
15+ YOE5+ MgmtBachelor's in CS/EE,15+ years SRE/infrastructure engineering,5+ years leading SRE teams,deep Linux,Terraform,Ansible,Kubernetes,Prometheus/Grafana/Datadog,Python or Go,cloud (AWS/Azure/GCP).
Prometheus, Grafana, Datadog, Terraform, Ansible, Kubernetes, Python, Go, AWS, Azure, GCP, Slurm, LSF, InfiniBand, RoCE, NVLink
1mo
Save
Mark Applied
Hide
Senior Manager, Networ Reliability Engineering
Santa Clara or Seattle or United States
$133k-$306k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
5+ YOE3+ Mgmt5+ years network reliability engineering, 3+ years engineering/operations management, strong cloud networking and distributed systems expertise, proven people leadership, excellent communication and organizational skills.
OCI