50 network reliability engineer jobs at 37 companies in Benicia, CA

1mo
Save
Mark Applied
Hide
Senior Network & Site Reliability Engineer
San Francisco, California, United States
$210k-$240k/yr OnsiteFull Time
Alembic
Alembic: AI-powered marketing attribution and revenue forecasting platform.
8+ YOE8+ years in network or infrastructure engineering (5+ years datacenter ops); strong network security and architecture skills; hands-on with BGP, QoS, MPLS, IPsec, EVPN/VXLAN, ECMP; IaC (Ansible, Terraform, Nornir); NetBox/Infoblox; Kubernetes networking; Linux; monitoring stacks; Python/Bash.
NVIDIA DGX SuperPOD, Grace Blackwell, BGP, VPNs, WAN, QoS, MPLS, IPsec, EVPN, VXLAN, ECMP, Ansible, Terraform, Nornir, NetBox, Infoblox, Kubernetes, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Python, Bash, Cumulus Linux, InfiniBand, Spectrum-X, BlueField, Spark, Airflow, Kafka, NFS, LustreFS, iSCSI, Linux
1mo
Save
Mark Applied
Hide
Senior Network Reliability Engineer - DGX Cloud
Santa Clara or United States
$136k-$265k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years network operations experience; deep knowledge of TCP/IP, BGP, OSPF, MPLS, EVPN/VxLAN and related protocols; experience in CSPs (AWS, Microsoft Azure, GCP, OCI); scripting/automation and vendor familiarity (Arista, Juniper, Fortinet).
TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, MACsec, AWS, Microsoft Azure, GCP, OCI, Arista, Fortinet, Juniper, Mellanox, Cumulus OS, NetBox, Nautobot, Prometheus, Grafana, Panoptes, Python, Shell
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineering - Network
Palo Alto or Columbus
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE10+ MgmtFormal network engineering training, 5+ years applied experience, 10+ years leading technologists, advanced network reliability skills, SD-WAN and cloud (AWS, Azure) proficiency, major network vendor experience, observability tooling and incident leadership.
SD-WAN, AWS, Azure, Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, Grafana, SevOne, Prometheus, Kibana, ThousandEyes, Splunk, Jenkins, GitLab, Terraform, eBPF, TCP/IP, HTTPS, BGP
2mo
Save
Mark Applied
Hide
Senior Network Reliability Engineer - DGX Cloud
Santa Clara or United States
$136k-$265k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years in network operations; strong TCP/IP, BGP, OSPF, MPLS, EVPN; experience with AWS/Azure/GCP; hands-on with automation; Bachelor’s in CS or related field.
TCP/IP, BGP, OSPF, MPLS, EVPN, VxLAN, GRE, IPsec, DNS, MACsec, Arista, Fortinet, Juniper, Mellanox, Cumulus OS, Infiniband, Netbox, Nautobot, Prometheus, Grafana, Python, Shell, AWS, Azure, GCP, OCI
2mo
Save
Mark Applied
Hide
Senior Director, Network Reliability
San Ramon or United States
$136k-$448k/yr HybridFull Time
Five9
Five9NASDAQ: FIVN: Provides cloud-based software for enterprise contact center operations.
10+ YOE10+ MgmtLeads global production networks; builds and mentors teams; drives engineering-led operations and modernization.
OSPF, EVPN, VXLAN, Cisco Nexus, NX-OS, Nexus, GCP, Terraform, Ansible, Python, Git, CI/CD, Kubernetes, ThousandEyes, Grafana, Loki, Netflow, Netscout, IaC
6d
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
OnsiteFull Time
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
AWS, Bash, C, Docker, Go, Kubernetes, Linux, Python, Rust, SSH, DNS, VPN, IAM
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
3mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability
Fremont, California, United States
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Designs and manufactures wafer fabrication equipment for the semiconductor industry.
12+ YOE6+ MgmtSenior SRE leader with 12+ years infrastructure/SRE/DevOps experience and 6+ years leading teams; deep multi-cloud networking, observability, DR/BCP, backup/restore, automation (Ansible/Terraform/Python), Kubernetes, and incident management experience.
Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Kubernetes, AKS, EKS, GKE, ServiceNow
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
1mo
Save
Mark Applied
Hide
Senior Manager, Networ Reliability Engineering
Santa Clara or Seattle or United States
$133k-$306k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
5+ YOE3+ Mgmt5+ years network reliability engineering, 3+ years engineering/operations management, strong cloud networking and distributed systems expertise, proven people leadership, excellent communication and organizational skills.
OCI
2mo
Save
Mark Applied
Hide
Integration Reliability Engineer
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years in SRE, infrastructure, or distributed systems; strong Linux and networking; production experience; cloud platforms; Kubernetes; observability tools; on-call readiness.
Kubernetes, Docker, Prometheus, Grafana, OpenTelemetry, Linux, Networking, AWS, GCP, Azure
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Francisco, California, United States
$200k-$260k/yr HybridFull Time
Sight Machine
Sight Machine: Developer of an AI-powered manufacturing data analytics platform.
10+ YOE10+ years experience with Kubernetes/Docker and major cloud providers, 10+ years coding (Python/Go/Java), IaC/CI-CD expertise, experience operating LLM/agentic AI systems, strong Linux and networking fundamentals.
Kubernetes, Docker, Azure, GCP, AWS, Python, Go, Java, Terraform, OpenTofu, FluxCD, Jenkins, GitHub Actions, Prometheus, Grafana, Loki, Sentry, Signoz, Helm Charts, Elasticsearch, Kafka, Postgres, LLM
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - US
United States or Oakland
$222k-$342k/yr RemoteFull Time
Teleport
Teleport: A remote-first building identity and security solutions for infrastructure and AI workloads.
Strong Linux, networking, containers, Go and Kubernetes experience; observability with Prometheus/Grafana/Loki; AWS or GCP experience; on-call participation; background check required; onboarding week in Oakland, CA.
Go, Kubernetes, AWS, GCP, Prometheus, Grafana, Loki, GitHub, Slack
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - AI Infrastructure
San Francisco or United States
RemoteFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Senior SRE with GPU infra, distributed training, and networking expertise.
NVIDIA GPUs, InfiniBand, RoCE, NVLink, NCCL, CUDA, PyTorch, DeepSpeed, Megatron, FSDP, Linux, Kubernetes, Slurm, Terraform, Helm, Ansible, DCGM, nvidia-smi
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Scottsdale or San Francisco or Chicago or New York
$194k-$237k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
12+ YOESenior-level SRE with 12+ years in software/technical leadership; strong in cloud, microservices, automation, and observability.
Python, Go, Java, Docker, Microservices, Kafka, SQS, JMS, Oracle, DynamoDB, Aurora, Redis, memcached, Linux, GIT, Chef, Maven, Jenkins, Networking, Kubernetes