441 observability engineer jobs at 264 companies in Petaluma, CA

2mo
Save
Mark Applied
Hide
Infrastructure Engineer (Observability)
New York or Remote or San Francisco or Seattle
$180k-$200k/yr HybridFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years in infrastructure engineering, SRE, or observability; strong monitoring (Prometheus, Grafana, ELK, VictoriaMetrics); experience with scalable observability platforms; Python/Go/bash automation; Kubernetes observability; streaming telemetry (Kafka, OTEL, Promtail); multi-tenant monitoring; strong communication.
Prometheus, Grafana, ELK, VictoriaMetrics, Python, Go, bash, Kubernetes, OTEL, Promtail, Kafka
3mo
Save
Mark Applied
Hide
Software Engineer, Security Observability
San Francisco or Seattle or New York City
$325k-$405k/yr RemoteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
Strong software engineering experience; Python or Go; IaC (Terraform); data pipelines for security; collaborate with security and engineering teams; proactive problem solving.
Python, Golang, Terraform, Azure
1mo
Save
Mark Applied
Hide
Platform Engineer, APIs & Observability
San Francisco or New York City
$180k-$200k/yr HybridFull Time
StackAI
StackAI: Build and deploy custom AI agents without writing code.
4+ YOE4+ years building backend services and public APIs; strong REST/OpenAPI skills, observability and distributed tracing, Python and FastAPI, familiarity with TypeScript/Node.js, and analytics-driven metrics and reporting.
OpenAPI, Python, FastAPI, TypeScript, Node.js, OpenTelemetry, ClickHouse, Druid
2mo
Save
Mark Applied
Hide
Senior Infra Engineer: Observability
San Francisco or United States or North America
RemoteFull Time
Railway
Railway: Infrastructure platform for automated application deployment and cloud hosting.
Experience building distributed systems, observability tooling, backend services in Golang/Rust, GRPC; familiarity with Terraform and Ansible; strong communication and ownership.
Golang, Rust, GRPC, Terraform, Ansible, TypeScript, GraphQL
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Observability
San Francisco or United States
$177k-$365k/yr HybridFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
7+ YOE7+ years in distributed systems and data engineering; expert in Java, Python, Go, or Scala; strong observability (metrics/logs/traces) with OpenTelemetry/Prometheus/Grafana; experience building scalable observability platforms; cloud-native (Kubernetes); product mindset and collaboration skills.
Java, Python, Go, Scala, Kafka, Flink, OpenTelemetry, Prometheus, Grafana, Kubernetes
1mo
Save
Mark Applied
Hide
Software Engineer | Observability
Lisbon or San Francisco or Sunnyvale or Raleigh or Seattle or Boston or London or Bengaluru or Dublin or Kyiv or United States or Portugal
RemoteFull Time
SingleStore
SingleStore: Provides a distributed SQL database for real-time analytics.
2+ YOE2+ years building distributed systems or backend services; strong Go (Golang) proficiency; familiarity with Kubernetes, cloud providers (AWS/GCP/Azure), observability concepts, and debugging production systems.
OpenTelemetry Collector, Alertmanager, SingleStore DB, Grafana, Loki, Tempo, Kubernetes, AWS, GCP, Azure, OpenTelemetry, OTLP protocol, Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, Apache Kafka, Parquet, Arrow, Flink, Go (Golang), Rust, Python, C++, SQL
1mo
Save
Mark Applied
Hide
Staff Engineer – Observability Platform
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman
Postman: Platform for building, testing, and managing software APIs.
10+ YOE10+ years software engineering experience with distributed systems, observability expertise, production operations, strong programming in Go/Java/Python/Node.js, and experience with monitoring/logging/tracing tools.
Go, Java, Python, Node.js, OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb
1mo
Save
Mark Applied
Hide
Sr. Staff Software Engineer — Observability, Insights & Governance
Mountain View or San Francisco
$229k-$314k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years in distributed systems, observability or governance; strong CS fundamentals; cross-functional communication; BS in CS (MS/PhD a plus).
Distributed Systems, Observability, Telemetry, Query Engines, Data Governance, Time-Series Storage, Logging Systems, LLMs (Bonus)
1mo
Save
Mark Applied
Hide
Full Stack Engineer, Observability
Oakland or United States
$146k-$235k/yr RemoteFull Time
LaunchDarkly
LaunchDarkly: Provides a feature management platform for software development teams.
5+ YOE5+ years of professional software engineering; TypeScript/React frontend and Go backend; experience with integrations, data pipelines or APIs; IaC tooling; RBAC and GitOps; strong communication.
TypeScript, React, Go, Terraform, Pulumi, OpenTelemetry, Datadog, Grafana
1mo
Save
Mark Applied
Hide
Senior Software Engineer - Observability and Reliability
New York City or San Francisco or London or Sydney
$170k-$240k/yr OnsiteFull Time
Sigma Computing
Sigma Computing: Cloud-native analytics platform featuring a spreadsheet-style interface.
5+ YOE5+ years building high-quality software, strong CS fundamentals, experience building observability tools, proficiency with Go, OpenTelemetry, Kubernetes, participation in on-call rotations, cloud service administration (GCP/AWS/Azure) preferred.
Go, OpenTelemetry, Kubernetes, GCP, AWS, Azure, SQL, Python
1mo
Save
Mark Applied
Hide
Engineering Manager, Observability Infrastructure
San Mateo, California, United States
$295k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
3+ Mgmt3+ years engineering management experience; strong background in building and operating large-scale distributed systems and data infrastructure; experience with observability or AI/ML infrastructure preferred; excellent communication and cross-functional collaboration skills.
3w
Save
Mark Applied
Hide
Platform Engineer
San Francisco, California, United States
OnsiteFull Time
HUD
HUD: Platform for reinforcement learning environments and AI agent evaluations.
Production infra and backend experience owning uptime, performance, deployment safety, and cost; strong AWS, Kubernetes/EKS, Terraform, CI/CD, observability, and backend engineering skills.
AWS, Terraform, Kubernetes, EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, CI/CD, Observability
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Evals & Observability
San Francisco or Mountain View
$200k-$300k/yr HybridFull Time
Glean
Glean: AI platform for enterprise search and automated workplace agents
2+ YOE2+ years software engineering; Go and Python; distributed data pipelines; LLM evaluation, RLHF, NLP; strong backend; analytics mindset
Go, Python, LLM evaluation, NLP, data pipelines
2w
Save
Mark Applied
Hide
Software Engineer- GPU Fabric Observability
San Francisco, California, United States
$200k-$380k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Staff-level experience building production infrastructure software, strong distributed systems and telemetry pipeline background, networking and high-performance network knowledge, experience processing high-volume operational data.
Kubernetes
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Triumph
Triumph: Skill-based mobile gaming platform for real money tournaments.
Experience operating and scaling large production systems; deep Postgres knowledge; CI/CD; observability tooling; able to lead a function independently.
PostgreSQL, AlloyDB, CI/CD, Observability, Telemetry, Logging
2mo
Save
Mark Applied
Hide
Principle Software Engineer, AI Observability & Evals Platform
Boston or San Francisco or New York City
$230k-$270k/yr OnsiteFull Time
LangChain
LangChain: Tools for building and deploying production-ready AI agents.
10+ YOE10+ years backend or full-stack engineering; strong Python/Go; TypeScript/React; architectural decisions; reliable, scalable systems; mentoring.
Python, Go, TypeScript, React
3mo
Save
Mark Applied
Hide
Senior Software Engineer - Customer Developer Observability
San Francisco, California, United States
$190k-$290k/yr OnsiteFull Time
Adyen
AdyenEuronext Amsterdam: ADYEN: Unified payment platform for global business commerce.
8+ YOE8+ years software development; strong Java web services; PostgreSQL; Kafka/Elasticsearch a plus; excellent English; SF office-based, no remote.
Java, PostgreSQL, Kafka, Elasticsearch
1mo
Save
Mark Applied
Hide
Forward Deployed Engineer
San Francisco, California, United States
$180k-$200k/yr RemoteFull Time
Felt Technologies
Felt Technologies: Provides embedded telehealth and provider network infrastructure for applications.
Proficiency with NextJS, React, Postgres; experience with serverless environments, observability, and healthtech/PII/PHI best practices; mobile SDK experience (Swift/Kotlin) is a plus. Must work US timezones.
NextJS, React, Postgres, Swift, Kotlin, Observability, serverless
1mo
Save
Mark Applied
Hide
Platform Engineer
San Francisco, California, United States
OnsiteFull Time
Phonic
Phonic: Building a platform for lifelike, reliable voice AI agents.
Strong experience with cloud infrastructure, containerized deployments, infrastructure-as-code, reliability engineering, SLOs, observability, and incident response; strong ownership and developer-experience focus.
GCP, AWS, Azure, Docker, Kubernetes, Terraform, WebSockets, WebRTC, TypeScript, Python
3d
Save
Mark Applied
Hide
Staff Software Engineer
Livingston or New York City or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in platform or infrastructure engineering with Kubernetes, CI/CD, IaC, observability, multi-region systems, and production ownership for high-availability services.
Spark, Airflow, Kafka, Flink, Kubernetes, Argo CD, GitHub Actions, Prometheus, Grafana, OpenTelemetry, Helm, Terraform, Pulumi