47 site operations engineer jobs at 41 companies in Ross, CA
2w
Save
Mark Applied
Hide
2w
Staff Site Reliability Engineer-Production Operations
Palo Alto, California, United States
$186k-$233k/yrHybridFull Time
Rivian and Volkswagen Group Technologies: A joint venture creating software-defined vehicle technology and connected services for electric vehicles.
Senior SRE with incident command experience, strong systems engineering for distributed systems, hands-on coding (Python or Go), observability expertise (Datadog or comparable), and experience with blameless post-incident practices.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
6+ YOE6+ years SRE/DevOps experience preferred; strong Kubernetes, IaC (Terraform/Pulumi), cloud networking, Go or Python, CI/CD, observability, and distributed systems experience; blockchain node operation experience is a plus.
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
Bellevue or Chicago or New York City or San Francisco or Washington
$174k-$267k/yrHybridFull Time
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
8+ YOE8+ years operations experience in cloud and Linux, strong networking and web server knowledge, proficiency with Terraform/Chef and scripting (Bash, Python, Go), experience with automation tools and on-call duty.
Member of Technical Staff, Site Reliablity Engineer
San Francisco, California, United States
$200k-$270k/yrHybridFull Time
Vapi: Build and deploy AI-powered voice agents via flexible APIs.
Experience running incident command and postmortems, operating SLOs/error budgets, capacity planning and load testing, Kubernetes production ops, KEDA autoscaling, and shipping services in Go or TypeScript.
San Francisco or New York or Denver or Austin or Calgary or Toronto
$118k-$135k/yrHybridFull Time
IntersectNASDAQ: GOOGL: Develops-located data centers and renewable energy infrastructure.
1+ YOEB.S. in Electrical Engineering; 1–3 years field experience; interpret electrical diagrams; analyze data from SCADA/CMMS; strong communication; safety-focused.
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOERequires 5+ years operating Linux systems in production or HPC environments, large-scale storage experience, incident response, monitoring, Kubernetes, CI/CD, Python or Go, and Terraform or Ansible.
Genesis AI: Building general-purpose robots with human-level physical intelligence.
Hands-on experience with robotics, hardware, teleoperation, or data-systems operations; comfortable debugging across hardware, software, and data; Linux and Python proficiency; customer-facing and travel-ready.
8+ YOEBA/BS in engineering; 8+ years infrastructure design/operation experience; 3+ years campus environment experience; proficiency with BAS and CMMS (e.g., SAP, Siemens); knowledge of NFPA, ASHRAE, OSHA; reliability engineering and capital project leadership.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
3+ YOE3+ years in data center or infrastructure engineering; strong Linux administration, x86 hardware, enterprise networking, BIOS/firmware and remote management; scripting with Bash or Python; on-site work and ability to lift/move servers.
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco or Seattle or New York City
$405k-$485k/yrHybridFull Time
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEProduction change-management experience, high-stakes release and on-call experience, AWS/GCP operations, Python proficiency, and a bachelor's degree or equivalent experience.
Python, Rust, AWS, GCP, AWS Bedrock, GCP Vertex, Claude
NV5NASDAQ: NVEE: Technical engineering and consulting for infrastructure and energy projects.
1+ YOEDegree in physical sciences/environmental science/engineering,1–3 years site/remediation experience,soil/soil-vapor sampling experience,ability to oversee SVE installation/operation,MS Office proficiency,GIS/AutoCAD a plus,California driver’s license required.
Pronto AI: Develops autonomous haulage technology for mining and industrial equipment.
15+ YOE5+ MgmtUniversity degree in mining engineering, 15+ years mining experience, 5+ years leadership, tech rollout experience, ability to work on-site in rugged conditions, medical clearances, strong communication and operational authority.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
2+ YOEBachelor's degree plus 2 years of software engineering experience, or 4 years of experience. Requires cloud, Linux, Kubernetes, observability, programming, automation, SRE, and incident management expertise.
Microsoft Azure, Google Cloud Platform (GCP), Kubernetes, Docker, Grafana, Prometheus, Splunk, Dynatrace, New Relic, OpenTelemetry, Linux, Unix, Python, Java, JavaScript, Shell, REST APIs, SQL, CI/CD, Infrastructure as Code, AIOps, LLMs, Web Content Accessibility Guidelines (WCAG)