41 network reliability engineer jobs at 33 companies in San Rafael, CA

1mo
Save
Mark Applied
Hide
Senior Network & Site Reliability Engineer
San Francisco, California, United States
$210k-$240k/yr OnsiteFull Time
Alembic
Alembic: AI-powered marketing attribution and revenue forecasting platform.
8+ YOE8+ years in network or infrastructure engineering (5+ years datacenter ops); strong network security and architecture skills; hands-on with BGP, QoS, MPLS, IPsec, EVPN/VXLAN, ECMP; IaC (Ansible, Terraform, Nornir); NetBox/Infoblox; Kubernetes networking; Linux; monitoring stacks; Python/Bash.
NVIDIA DGX SuperPOD, Grace Blackwell, BGP, VPNs, WAN, QoS, MPLS, IPsec, EVPN, VXLAN, ECMP, Ansible, Terraform, Nornir, NetBox, Infoblox, Kubernetes, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Python, Bash, Cumulus Linux, InfiniBand, Spectrum-X, BlueField, Spark, Airflow, Kafka, NFS, LustreFS, iSCSI, Linux
2w
Save
Mark Applied
Hide
Senior Network Reliability Engineer, Incident Management
Mountain View or Espoo or Bengaluru
HybridFull Time
Skylo
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
5+ YOE5+ years telecom or network operations experience, incident/outage management expertise, observability and Kubernetes literacy, ticketing and on-call tool proficiency, strong written and verbal communication.
Grafana, Prometheus, Loki, kubectl, Jira, ServiceNow, PagerDuty, Python, Bash
2w
Save
Mark Applied
Hide
Site Reliability Engineer (Network)
San Francisco or Golden
$157k-$239k/yr OnsiteFull Time
Loft Orbital
Loft Orbital: Deploy and operate satellite missions for organizations and governments.
4+ YOE4–5 years network engineering experience, hands-on SDN and public-cloud networking (ideally GCP), Kubernetes/Docker familiarity, IaC (Terraform) and GitOps, SRE mindset (SLOs, observability), degree or equivalent experience.
GCP, k8s, Docker, Terraform, Grafana, ArgoCD, FluxCD, Cockpit, GitOps
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineering - Network
Palo Alto or Columbus
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE10+ MgmtFormal network engineering training, 5+ years applied experience, 10+ years leading technologists, advanced network reliability skills, SD-WAN and cloud (AWS, Azure) proficiency, major network vendor experience, observability tooling and incident leadership.
SD-WAN, AWS, Azure, Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, Grafana, SevOne, Prometheus, Kibana, ThousandEyes, Splunk, Jenkins, GitLab, Terraform, eBPF, TCP/IP, HTTPS, BGP
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
OnsiteFull Time
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
AWS, Bash, C, Docker, Go, Kubernetes, Linux, Python, Rust, SSH, DNS, VPN, IAM
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
2w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Redwood City, California, United States
HybridFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
5+ YOE5+ years SRE or infrastructure experience, deep Linux and low-level performance debugging, Terraform, Airflow, Ray, AWS or OCI, high-performance networking (InfiniBand/RDMA/RoCE), security/compliance familiarity.
Linux, Python, Go, Bash, Terraform, Airflow, Ray, AWS, OCI, InfiniBand, RDMA, RoCE, DCGM, ROCm, Kubernetes, NVIDIA, AMD
4d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Francisco, California, United States
$195k-$258k/yr RemoteFull Time
Circle
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
6+ YOE6+ years SRE/DevOps experience preferred; strong Kubernetes, IaC (Terraform/Pulumi), cloud networking, Go or Python, CI/CD, observability, and distributed systems experience; blockchain node operation experience is a plus.
Kubernetes, Helm, Terraform, Pulumi, Go, Python, CI/CD, RBAC, VPCs, DNS, SQL
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - US
United States or Oakland
$222k-$342k/yr RemoteFull Time
Teleport
Teleport: A remote-first building identity and security solutions for infrastructure and AI workloads.
Strong Linux, networking, containers, Go and Kubernetes experience; observability with Prometheus/Grafana/Loki; AWS or GCP experience; on-call participation; background check required; onboarding week in Oakland, CA.
Go, Kubernetes, AWS, GCP, Prometheus, Grafana, Loki, GitHub, Slack
2mo
Save
Mark Applied
Hide
Production Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York
$66k-$82k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's degree in CS/IS; 2-5 years in IT/DevOps; ITIL/ITSM knowledge; Agile; strong problem solving; ability to relate business needs to system capabilities.
ITIL, ITSM, Agile, Networking, Distributed Systems
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
2mo
Save
Mark Applied
Hide
Production Engineer, Network
London or San Francisco
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Design and own network reliability, build debugging tooling and automation, run incidents, and deliver realtime monitoring. Experience with network automation, observability, and production tooling expected.
LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, gNMI, gRPC, NETCONF, SONiC, BGP, ECMP, Go, Python, SSH
3mo
Save
Mark Applied
Hide
Software Engineer, Core Network Engineering
San Francisco, California, United States
$230k-$342k/yr OnsiteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
Design, build, and operate networking systems for large-scale AI training; improve performance and reliability; develop automation and observability tooling.
Linux, RDMA, InfiniBand, RoCE, DPDK, C++, Python, Go, NICs
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides edge-to-cloud IT infrastructure and platform services.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, reliability, SolidWorks/PLM, FEA, GD&T, thermal and cooling technologies; strong leadership and troubleshooting skills.
SolidWorks, EPDM, FEA, ECAD, IDF
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Providing global edge-to-cloud infrastructure and IT solutions for businesses.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, thermal and reliability; SolidWorks and PLM experience; tolerance/GD&T, FEA, ECAD-MCAD, and supplier engagement experience.
SolidWorks, EPDM, PLM, FEA, ECAD-MCAD, IDF, GD&T

Explore Jobs

Expand Your Job Search