914 observability jobs at 442 companies in San Francisco, CA

1mo
Save
Mark Applied
Hide
Engineering Manager Observability
Mountain View or Austin or Sunnyvale or Warren or United States
$219k-$335k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
7+ Mgmt7+ years managing software or SRE teams; deep observability knowledge (logs, metrics, traces) with Prometheus, Grafana, OpenTelemetry; distributed systems architecture; programming in Go, Python, Typescript; cloud (GCP/AWS/Azure), Kubernetes, Docker; strong communication.
Prometheus, Grafana, OpenTelemetry, Go, Python, Typescript, GCP, AWS, Azure, Kubernetes, Docker, Istio, Terraform, TSDBs, CI/CD pipelines
2d
Save
Mark Applied
Hide
Observability Software Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
3+ YOEBachelor's degree or higher, 3+ years of computer science experience, Java or Go proficiency, distributed systems and concurrency knowledge, observability tools experience, and fluent English and Chinese.
Java, Go, OpenTelemetry, CAT, SkyWalking, Prometheus, VictoriaMetrics, ELK, ClickHouse, eBPF, Kubernetes, Linux, PyTorch, Spring AI, Langfuse
2mo
Save
Mark Applied
Hide
Senior Engineer, Network Observability
Warsaw or Bellevue or Ireland or Livingston or London or New York or Sunnyvale
OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Cloud computing infrastructure specialized for large-scale AI workloads.
Experience building and operating network observability: Prometheus/Grafana/Alertmanager, gNMI/SNMP, Python/Go/Bash, Kubernetes, Linux networking; telemetry collectors/exporters, on-call support, and mentoring junior engineers.
Prometheus, Grafana, Alertmanager, gNMI, SNMP, Python, Golang, Go, Bash, Kubernetes, Ansible, Jinja2, TensorFlow, scikit-learn, OpenTelemetry, Jaeger, Zipkin, Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, Linux
1w
Save
Mark Applied
Hide
Sr. Staff Observability Software Engineer
Mountain View or Seattle
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
8+ YOEBachelor's in CS/EE/Math,8+ years building large-scale distributed systems,deep observability experience (metrics,logs,tracing),SLO/KPI definition,cloud and container familiarity,programming in Go/Java/Python/Ruby.
OpenTelemetry, Prometheus, Grafana, Elastic stack, Docker, Kubernetes, Dynatrace, AppDynamics, Jaeger, Zipkin, AWS, Azure, Google Cloud Platform, Go, Java, Python, Ruby
5d
Save
Mark Applied
Hide
Software Engineer - Observability
San Francisco or Toronto or New York City or Montreal
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Deep observability experience, scalable telemetry pipeline knowledge, Prometheus, Grafana, ClickHouse, or OpenTelemetry experience, and proficiency in Python, Rust, or Go.
Prometheus, Grafana, ClickHouse, OpenTelemetry, Python, Rust, Go
3w
Save
Mark Applied
Hide
Engineering Manager, Observability
Sunnyvale, California, United States
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE2+ Mgmt5+ years software engineering, 2+ years engineering management, experience with observability platforms, reliability engineering, scaling telemetry, and hiring/managing teams.
OpenTelemetry, Grafana, Prometheus, Kubernetes
3mo
Save
Mark Applied
Hide
Tech Lead – Network Observability
Palo Alto, California, United States
$180k-$260k/yr OnsiteFull Time
Clockwork Systems
Clockwork Systems: Software-driven network fabrics for GPU cluster performance optimization.
Lead architecture and development of a high-performance network observability platform; strong distributed systems, Linux networking, and observability tooling; mentor engineers.
Prometheus, Grafana, Datadog, OpenTelemetry, NCCL, rdma-core, libibverbs, tcpdump, Wireshark, ethtool, iproute2, Linux kernel networking tools
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Observability
San Francisco or United States
$177k-$365k/yr HybridFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
7+ YOE7+ years in distributed systems and data engineering; expert in Java, Python, Go, or Scala; strong observability (metrics/logs/traces) with OpenTelemetry/Prometheus/Grafana; experience building scalable observability platforms; cloud-native (Kubernetes); product mindset and collaboration skills.
Java, Python, Go, Scala, Kafka, Flink, OpenTelemetry, Prometheus, Grafana, Kubernetes
2mo
Save
Mark Applied
Hide
Sr. Staff Software Engineer — Observability, Insights & Governance
Mountain View or San Francisco
$229k-$314k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years in distributed systems, observability or governance; strong CS fundamentals; cross-functional communication; BS in CS (MS/PhD a plus).
Distributed Systems, Observability, Telemetry, Query Engines, Data Governance, Time-Series Storage, Logging Systems, LLMs (Bonus)
2w
Save
Mark Applied
Hide
Sr. Partner Dev Mgr, DevOps and Observability, AMER, Observability
Austin or Seattle or Santa Monica or New York or East Palo Alto or Chicago or United States or San Francisco
$148k-$220k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
6+ YOE6+ years GTM/business development/sales experience, 5+ years negotiating business agreements, experience with Observability/DevOps/Data/Analytics/SaaS, strong strategic and co-selling skills.
AWS Marketplace
3mo
Save
Mark Applied
Hide
Senior Infra Engineer: Observability
San Francisco or United States or North America
RemoteFull Time
Railway
Railway: Infrastructure platform for automated application deployment and cloud hosting.
Experience building distributed systems, observability tooling, backend services in Golang/Rust, GRPC; familiarity with Terraform and Ansible; strong communication and ownership.
Golang, Rust, GRPC, Terraform, Ansible, TypeScript, GraphQL
1mo
Save
Mark Applied
Hide
Staff Engineer – Observability Platform
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman
Postman: Platform for building, testing, and managing software APIs.
10+ YOE10+ years software engineering experience with distributed systems, observability expertise, production operations, strong programming in Go/Java/Python/Node.js, and experience with monitoring/logging/tracing tools.
Go, Java, Python, Node.js, OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb
2w
Save
Mark Applied
Hide
Staff Software Engineer, Observability
Menlo Park or Toronto
HybridFull Time
Robinhood
RobinhoodNASDAQ: HOOD: Provides a commission-free platform for investing and financial services.
8+ YOE8+ years software engineering experience, deep Kubernetes and public cloud (AWS) expertise, strong coding in Go/Python, experience owning observability/control-plane and telemetry pipelines.
Kubernetes, AWS, Go, Python, Vector, Prometheus, Grafana, Honeycomb, Humio, Sentry
1w
Save
Mark Applied
Hide
Staff Software Engineer, Observability
San Francisco or United States
$177k-$365k/yr HybridFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding and saving ideas.
7+ YOERequires 7+ years designing and operating large-scale distributed systems, data engineering expertise, observability experience, expert Java, Python, Go, or Scala skills, and strong technical leadership.
Kafka, Flink, OpenTelemetry, Prometheus, Grafana, Java, Python, Go, Scala, Kubernetes
1mo
Save
Mark Applied
Hide
Platform Engineer, APIs & Observability
San Francisco or New York City
$180k-$200k/yr HybridFull Time
StackAI
StackAI: Build and deploy custom AI agents without writing code.
4+ YOE4+ years building backend services and public APIs; strong REST/OpenAPI skills, observability and distributed tracing, Python and FastAPI, familiarity with TypeScript/Node.js, and analytics-driven metrics and reporting.
OpenAPI, Python, FastAPI, TypeScript, Node.js, OpenTelemetry, ClickHouse, Druid
1mo
Save
Mark Applied
Hide
Principal Software Development Engineer - Observability
San Jose or Seattle
$249k-$349k/yr OnsiteFull Time
Expedia Group
Expedia GroupNASDAQ: EXPE: Operates a global platform for travel bookings and services.
10+ YOE10+ years building and operating large-scale distributed systems; deep observability expertise (logs, metrics, traces); proficiency in Go/Java/Python, cloud-native tech, OpenTelemetry, and observability tooling; leadership and architecture experience.
Go, Java, Python, AWS, Kubernetes, Docker, OpenTelemetry, Prometheus, Grafana, Datadog, Splunk, Clickhouse, Terraform, Crossplane, Electronic Medical Records (EMR)
1mo
Save
Mark Applied
Hide
Staff Product Manager - Observability
Bellevue or San Francisco
$291k-$430k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years product management experience including 3+ years on technical infrastructure or platform products; experience defining observability, SLIs/SLOs/SLA, partnering with SRE and fleet engineering; strong writing, influence, and execution skills.
NCCL, InfiniBand, PyTorch, DCGM, Datadog, Grafana, Prometheus
2mo
Save
Mark Applied
Hide
Engineering Manager, Observability Infrastructure
San Mateo, California, United States
$295k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
3+ Mgmt3+ years engineering management experience; strong background in building and operating large-scale distributed systems and data infrastructure; experience with observability or AI/ML infrastructure preferred; excellent communication and cross-functional collaboration skills.
3mo
Save
Mark Applied
Hide
Senior Product Manager - Data Observability
Menlo Park, California, United States
$200k-$288k/yr OnsiteFull Time
Snowflake
SnowflakeNYSE: SNOW: Cloud-based platform for data storage, processing, and analytics.
5+ YOE5+ years of product management in data infrastructure or observability; AI/ML familiarity; strong customer empathy; defines strategy and roadmaps.
2w
Save
Mark Applied
Hide
Senior Software Engineer, AIOps and Observability
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years product development experience, 5+ years building/operating observability platforms, Bachelor's in CS/Engineering or equivalent, experience with Prometheus/Grafana/OpenTelemetry, Kubernetes, cloud/on-prem observability, and Go/Python/Java/C#.
Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, BigPanda, PagerDuty, Datadog, Kubernetes, Nomad, Docker, NATS, Kafka, Go, Python, Java, C#