81 platform reliability engineer jobs at 69 companies in Cotati, CA

1mo
Save
Mark Applied
Hide
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yr HybridFull Time
Grow Therapy
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
AWS, Kubernetes, EKS, Terraform, DataDog, PostgreSQL, Gem
2w
Save
Mark Applied
Hide
Senior Reliability Engineer
San Francisco or Oakland
$139k-$205k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years reliability validation or hardware testing experience for robotics/unmanned platforms, Bachelor's in engineering, proficiency with environmental test equipment, Python and CAD, strong communication and analytical skills.
Python, CAD, DAQ, FRACAS, HIL
2mo
Save
Mark Applied
Hide
Director of Platform & Reliability Engineering
San Francisco or New York City
$235k-$245k/yr HybridFull Time
Forge Global
Forge GlobalNYSE: FRGE: Marketplace for trading private shares and pre-IPO stock.
8+ YOE5+ Mgmt8+ years software engineering experience with infrastructure/platform focus, 5+ years people leadership, deep cloud/observability/incident response experience, strong distributed systems judgment, and ability to set platform strategy.
Kubernetes, CI/CD
2mo
Save
Mark Applied
Hide
Senior Platform Engineer
San Francisco, California, United States
$170k-$230k/yr HybridFull Time
Authorium
Authorium: Cloud-based administrative operations platform for government agencies.
6+ YOE6+ years in platform/infrastructure/DevOps engineering with distributed systems, cloud architecture, security, and reliability; strong collaboration in an in-person SF office (Mon–Thu).
Cursor, Claude, AWS ECS, AWS EKS, CloudWatch, IAM, VPC, Parameter Store, Terraform, Pulumi, CDK, Datadog, OpenTelemetry
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1w
Save
Mark Applied
Hide
Platform Engineer, Billing Systems
San Francisco or United States
$220k-$350k/yr HybridFull Time
Wispr Flow
Wispr Flow: Provides AI-powered voice dictation software for computers and mobile devices.
Experienced engineer who has built billing systems, run payments migrations, reconciled billing across platforms, and built testing for high-reliability billing.
RevenueCat, Stripe, Sequence
1mo
Save
Mark Applied
Hide
Platform Engineer
San Francisco, California, United States
OnsiteFull Time
Phonic
Phonic: Building a platform for lifelike, reliable voice AI agents.
Strong experience with cloud infrastructure, containerized deployments, infrastructure-as-code, reliability engineering, SLOs, observability, and incident response; strong ownership and developer-experience focus.
GCP, AWS, Azure, Docker, Kubernetes, Terraform, WebSockets, WebRTC, TypeScript, Python
2mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Ray, Kubernetes, Prometheus, Grafana, Go, Python, AWS, Azure, GCP, Linux kernel
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Plenful
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.
OpenTelemetry, Datadog, CloudWatch, Grafana, Sentry, AWS Lambda, ECS, Aurora Postgres, ClickHouse, GitHub Actions, Python, Bash, Vanta
2w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
3mo
Save
Mark Applied
Hide
Senior Platform Engineer
Santa Monica or Lower Manhattan or San Francisco or Los Angeles
$150k-$200k/yr HybridFull Time
Pivotal Health
Pivotal Health: AI platform automating healthcare insurance claim disputes for providers.
5+ YOE5+ years in platform, infrastructure, or software engineering; strong Python; cloud-native systems (GCP); Terraform; CI/CD; containers; event-driven architectures; security and reliability.
Python, Terraform, GitHub Actions, Kubernetes, Docker, Kafka, Pub/Sub, Kinesis, Google Cloud Platform
1mo
Save
Mark Applied
Hide
Foundations Engineer (Platform Hybrid, SF in office)
San Francisco, California, United States
HybridFull Time
Rox
Rox: AI-powered revenue operating system with autonomous sales agents.
Experience building and operating production infrastructure and distributed systems; strong cloud (AWS/GCP) and Kubernetes experience; CI/CD and developer platform knowledge; observability and reliability engineering; production debugging skills.
Kubernetes, AWS, GCP, CI/CD
1mo
Save
Mark Applied
Hide
Platform Engineer
San Francisco, California, United States
OnsiteFull Time
Engram
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building and scaling production systems, experience designing APIs/platforms, strong product orientation, security/reliability focus, and experience across the stack; early-stage/ML experience a plus.
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Hiring Sprint
San Francisco, California, United States
$196k-$255k/yr HybridFull Time
Airbyte
Airbyte: Open-source data integration platform for automated data movement.
7+ YOE7+ years in infrastructure/platform engineering/SRE/DevOps; hands-on Kubernetes, Helm, Terraform; observability with Prometheus/Grafana/Datadog; CI/CD ownership; ability to read backend code; fluency with LLMs and agentic tools.
Kubernetes, Helm, Terraform, AWS, GCP, Prometheus, Grafana, Datadog, CI/CD, Java, Python, Airbyte, CDKs, LLMs
2d
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Charlotte or Chandler or San Francisco or Columbus
$119k-$224k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
5+ YOERequires 5+ years in systems engineering or architecture, 5+ years SRE, cloud observability experience, hosting platforms, and familiarity with DevOps, Agile, and IT service management.
Elasticsearch, Kibana, Kafka, Airflow, Logstash, Grafana, Elastic APM, Jaeger, Zipkin, AWS, OCP, Kubernetes, PKS, Azure, VMware, Unix, Linux, Windows, Jenkins, Maven, Gradle, Groovy, Artifactory, GIT, Harness IO, Spinnaker, Terraform, UDeploy, AIOPS, ServiceNow, Remedy, Big Panda, Netcool
5d
Save
Mark Applied
Hide
Director, Site Reliability Engineering
New York City or San Francisco or Dallas
$197k-$314k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ YOE5+ MgmtBachelor's in a technical field,10+ years engineering experience with 5+ years leading SRE/Platform teams; experience with observability, incident management, distributed systems, and cloud architecture.
AWS, New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, OpenTelemetry
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Platform
San Francisco, California, United States
OnsiteFull Time
Emanate
Emanate: A building AI-powered revenue platform.
Senior software engineer with platform infrastructure focus, data pipelines, deployment automation, and reliability at scale.