30 network reliability engineer jobs at 24 companies in Cotati, CA

1mo
Save
Mark Applied
Hide
Senior Network & Site Reliability Engineer
San Francisco, California, United States
$210k-$240k/yr OnsiteFull Time
Alembic
Alembic: AI-powered marketing attribution and revenue forecasting platform.
8+ YOE8+ years in network or infrastructure engineering (5+ years datacenter ops); strong network security and architecture skills; hands-on with BGP, QoS, MPLS, IPsec, EVPN/VXLAN, ECMP; IaC (Ansible, Terraform, Nornir); NetBox/Infoblox; Kubernetes networking; Linux; monitoring stacks; Python/Bash.
NVIDIA DGX SuperPOD, Grace Blackwell, BGP, VPNs, WAN, QoS, MPLS, IPsec, EVPN, VXLAN, ECMP, Ansible, Terraform, Nornir, NetBox, Infoblox, Kubernetes, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Python, Bash, Cumulus Linux, InfiniBand, Spectrum-X, BlueField, Spark, Airflow, Kafka, NFS, LustreFS, iSCSI, Linux
2w
Save
Mark Applied
Hide
Site Reliability Engineer (Network)
San Francisco or Golden
$157k-$239k/yr OnsiteFull Time
Loft Orbital
Loft Orbital: Deploy and operate satellite missions for organizations and governments.
4+ YOE4–5 years network engineering experience, hands-on SDN and public-cloud networking (ideally GCP), Kubernetes/Docker familiarity, IaC (Terraform) and GitOps, SRE mindset (SLOs, observability), degree or equivalent experience.
GCP, k8s, Docker, Terraform, Grafana, ArgoCD, FluxCD, Cockpit, GitOps
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
OnsiteFull Time
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
AWS, Bash, C, Docker, Go, Kubernetes, Linux, Python, Rust, SSH, DNS, VPN, IAM
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
2w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
5d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
5d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Francisco, California, United States
$195k-$258k/yr RemoteFull Time
Circle
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
6+ YOE6+ years SRE/DevOps experience preferred; strong Kubernetes, IaC (Terraform/Pulumi), cloud networking, Go or Python, CI/CD, observability, and distributed systems experience; blockchain node operation experience is a plus.
Kubernetes, Helm, Terraform, Pulumi, Go, Python, CI/CD, RBAC, VPCs, DNS, SQL
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - US
United States or Oakland
$222k-$342k/yr RemoteFull Time
Teleport
Teleport: A remote-first building identity and security solutions for infrastructure and AI workloads.
Strong Linux, networking, containers, Go and Kubernetes experience; observability with Prometheus/Grafana/Loki; AWS or GCP experience; on-call participation; background check required; onboarding week in Oakland, CA.
Go, Kubernetes, AWS, GCP, Prometheus, Grafana, Loki, GitHub, Slack
2mo
Save
Mark Applied
Hide
Production Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York
$66k-$82k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's degree in CS/IS; 2-5 years in IT/DevOps; ITIL/ITSM knowledge; Agile; strong problem solving; ability to relate business needs to system capabilities.
ITIL, ITSM, Agile, Networking, Distributed Systems
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
2mo
Save
Mark Applied
Hide
Production Engineer, Network
London or San Francisco
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Design and own network reliability, build debugging tooling and automation, run incidents, and deliver realtime monitoring. Experience with network automation, observability, and production tooling expected.
LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, gNMI, gRPC, NETCONF, SONiC, BGP, ECMP, Go, Python, SSH
3mo
Save
Mark Applied
Hide
Software Engineer, Core Network Engineering
San Francisco, California, United States
$230k-$342k/yr OnsiteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
Design, build, and operate networking systems for large-scale AI training; improve performance and reliability; develop automation and observability tooling.
Linux, RDMA, InfiniBand, RoCE, DPDK, C++, Python, Go, NICs
1mo
Save
Mark Applied
Hide
Systems Engineer
San Francisco, California, United States
OnsiteFull Time
Dedalus Labs
Dedalus Labs: Infrastructure for building and deploying AI agent applications.
Fluency in Rust/Go/C/C++; strong software engineering fundamentals and systems knowledge (OS, networking, distributed systems); performance-oriented debugging and reliability focus.
Rust, Go, C/C++, Kubernetes, Firecracker, GitHub
4d
Save
Mark Applied
Hide
Staff Software Engineer, Platform
San Francisco or New York or Washington
$240k-$300k/yr OnsiteFull Time
Peregrine
Peregrine: Data integration and analytics platform for public safety agencies.
8+ YOE8+ years building and operating cloud infrastructure, hands-on ownership of platform/networking/traffic systems, strong security and reliability experience, on-call and incident response experience, degree or equivalent.
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$150k-$300k/yr OnsiteFull Time
Reducto
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure; proficient in Python; strong cloud, Kubernetes, networking, storage, and automation; focus on reliability.
Python, Kubernetes, Cloud platforms, Networking, Storage, Automation
1mo
Save
Mark Applied
Hide
Senior AV Automation Engineer
New York City or Seattle or San Francisco
$149k-$246k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years in systems/site reliability/DevOps/AV automation, proficiency in Python and REST API integrations, experience with automation/configuration tools and networking concepts, related technical degree required.
Splunk, Grafana, Slack, NetBox, Python, Ansible, Terraform, Puppet, Chef, Salt, New Relic, Kentik, Google Meet, Logitech, Neat, Cisco, Q-SYS, Google Workspace, Zoom, WebEx, Git, REST, AWS
1mo
Save
Mark Applied
Hide
Staff Cluster Infrastructure Engineer
San Francisco, California, United States
$224k-$284k/yr OnsiteFull Time
Atoms
Atoms: Building robotics and software to automate physical world industries.
6+ YOE6+ years operating GPU compute on Kubernetes, strong Python/Go programming, familiarity with Terraform or CloudFormation, experience with bare-metal Linux, GPU hardware, and networking; strong automation and reliability focus.
Kubernetes, Python, Go, Terraform, CloudFormation, Linux
2mo
Save
Mark Applied
Hide
Senior/Staff Software Engineer, Core Infrastructure
San Francisco, California, United States
$160k-$210k/yr RemoteFull Time
Zip
Zip: AI-powered intake-to-procure platform for enterprise spend management
6+ YOE6+ years software engineering in infrastructure; BS or higher in CS or related; Kubernetes/EKS, multi-region, observability; experience in a small company; quick learner.
Kubernetes, EKS, Networking, Observability, Reliability, Performance Engineering