60 network reliability engineer jobs at 31 companies in Santa Cruz, CA

3w
Save
Mark Applied
Hide
Senior Network Reliability Engineer, Incident Management
Mountain View or Espoo or Bengaluru
HybridFull Time
Skylo
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
5+ YOE5+ years telecom or network operations experience, incident/outage management expertise, observability and Kubernetes literacy, ticketing and on-call tool proficiency, strong written and verbal communication.
Grafana, Prometheus, Loki, kubectl, Jira, ServiceNow, PagerDuty, Python, Bash
2mo
Save
Mark Applied
Hide
Senior Network Reliability Engineer - DGX Cloud
Santa Clara or United States
$136k-$265k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years network operations experience; deep knowledge of TCP/IP, BGP, OSPF, MPLS, EVPN/VxLAN and related protocols; experience in CSPs (AWS, Microsoft Azure, GCP, OCI); scripting/automation and vendor familiarity (Arista, Juniper, Fortinet).
TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, MACsec, AWS, Microsoft Azure, GCP, OCI, Arista, Fortinet, Juniper, Mellanox, Cumulus OS, NetBox, Nautobot, Prometheus, Grafana, Panoptes, Python, Shell
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineering - Network
Palo Alto or Columbus
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE10+ MgmtFormal network engineering training, 5+ years applied experience, 10+ years leading technologists, advanced network reliability skills, SD-WAN and cloud (AWS, Azure) proficiency, major network vendor experience, observability tooling and incident leadership.
SD-WAN, AWS, Azure, Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, Grafana, SevOne, Prometheus, Kibana, ThousandEyes, Splunk, Jenkins, GitLab, Terraform, eBPF, TCP/IP, HTTPS, BGP
3mo
Save
Mark Applied
Hide
Senior Network Reliability Engineer - DGX Cloud
Santa Clara or United States
$136k-$265k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years in network operations; strong TCP/IP, BGP, OSPF, MPLS, EVPN; experience with AWS/Azure/GCP; hands-on with automation; Bachelor’s in CS or related field.
TCP/IP, BGP, OSPF, MPLS, EVPN, VxLAN, GRE, IPsec, DNS, MACsec, Arista, Fortinet, Juniper, Mellanox, Cumulus OS, Infiniband, Netbox, Nautobot, Prometheus, Grafana, Python, Shell, AWS, Azure, GCP, OCI
4d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
1d
Save
Mark Applied
Hide
Site Reliability Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
Experience with large-scale reliability, high-availability architecture, incident response, cross-region disaster recovery, Linux, networking, middleware, cloud-native infrastructure, automation, and Python, Go, or Java.
Linux, MySQL, Redis, Kafka, Kubernetes, Service Mesh, Python, Go, Java
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
1w
Save
Mark Applied
Hide
Systems Reliability Engineer II
San Jose or Durham or Mexico City or Bangalore or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$71k-$143k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
2+ YOEMinimum 2 years relevant experience; degree in IT/Networking/Cloud preferred; expertise troubleshooting Linux, virtualization, networking or storage; excellent written/verbal communication; fluent English.
Linux, VMware ESXi, VMware, Citrix, Microsoft
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS/Engineering, strong Linux, networking, databases, Kubernetes, SRE/DevOps toolset knowledge, experience with ClickHouse/Spark/Presto/Doris/Hadoop, coding in Python/Shell/Java/Go, strong problem-solving and communication.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Python, Shell, Java, Go
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Redwood City, California, United States
HybridFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
5+ YOE5+ years SRE or infrastructure experience, deep Linux and low-level performance debugging, Terraform, Airflow, Ray, AWS or OCI, high-performance networking (InfiniBand/RDMA/RoCE), security/compliance familiarity.
Linux, Python, Go, Bash, Terraform, Airflow, Ray, AWS, OCI, InfiniBand, RDMA, RoCE, DCGM, ROCm, Kubernetes, NVIDIA, AMD
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
1d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Mateo or United States
$240k-$300k/yr RemoteFull Time
Skydio
Skydio: Develops autonomous AI drones for defense and industrial inspection.
8+ YOE8+ years in SRE, platform, DevOps, production engineering, or equivalent; strong Kubernetes and AWS experience; Terraform, CI/CD, networking, observability, and production reliability expertise.
Kubernetes, Amazon Web Services (AWS), Amazon Elastic Kubernetes Service (EKS), Terraform, Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, Jenkins, Linux, Python, Go, Helm, GitOps, Datadog, PostgreSQL
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Experience with Linux, networking, databases, Kubernetes, ClickHouse/Hadoop/Doris/Spark/Presto, scripting or programming (Python, Shell, Java, Go), incident management, and capacity planning.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Linux, Python, Shell, Java, Go
1w
Save
Mark Applied
Hide
Site Reliability Engineer, AI Infrastructure
San Jose, California, United States
$123k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOEBachelor's degree or equivalent experience, 1+ year in SRE, DevOps, or systems engineering, Linux and networking knowledge, distributed systems experience, programming, scripting, CI/CD, and automation skills.
Linux, Go, Python, C, C++, Java, Bash, Kubernetes, AWS, GCP, Azure, Terraform, Prometheus, Grafana, Distributed Tracing, LLMs, Agentic AI
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Jose, California, United States
$119k-$170k/yr HybridFull Time
Zscaler
ZscalerNASDAQ: ZS: Provides cloud-native cybersecurity solutions through zero trust architecture.
5+ YOE5+ years Linux/UNIX admin, Kubernetes/Docker, automation (Ansible), network/security fundamentals, and strong security practices.
Docker, Kubernetes, Ansible, Python, Golang, BASH, Openstack, CEPH, HashiCorp Vault, nftables
3d
Save
Mark Applied
Hide
Mid-Senior Site Reliability Engineer Kubernetes Platform
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOERequires 8+ years in SRE, DevOps, or platform engineering; Kubernetes, cloud, Linux, networking, distributed systems, IaC, scripting, monitoring, and compliance experience. Must obtain U.S. security clearance.
Kubernetes, AWS, Microsoft Azure, Linux, Terraform, Python, Go, Bash, Prometheus, Grafana, NIST 800-53, STIGs, RMF, Istio, Linkerd, OPA/Gatekeeper, Kyverno, ArgoCD, FedRAMP, FIPS, DoD Cloud SRG, CI/CD
1d
Save
Mark Applied
Hide
GPU DC East-West Network SRE Expert (SME)
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years in data center networking, including 3+ years with InfiniBand or RoCE; experience with Nvidia/Mellanox fabrics, UFM, RoCEv2, NCCL, RDMA diagnostics, optics, cabling, and telemetry-driven operations.
InfiniBand, RoCEv2, Nvidia, Arista, Cisco, UFM, Unified Fabric Manager, NCCL, ibdiagnet, perfquery, ibstat, DCQCN, ECN, PFC, RDMA, GDR

Explore Jobs

Expand Your Job Search