903 observability engineer jobs at 473 companies in California

1mo
Save
Mark Applied
Hide
Senior Observability Engineer
New York City or Atlanta or Jersey City or Los Angeles or United States or Canada or Puerto Rico or United Kingdom or Ireland or Portugal or Romania or Australia
$149k-$186k/yr HybridFull Time
FanDuel
FanDuelNYSE: FLUT: Offers online sports betting and daily fantasy sports services.
Hands-on observability/SRE/platform engineering experience; expertise with monitoring, alerting, SLO/SLI design; proficiency with Datadog, Kubernetes, AWS, Terraform; software engineering skills in Go/Java/Python/TypeScript; mentoring and automation experience.
AWS, Kubernetes, Terraform, Helm, Ansible, Vault, Datadog, PagerDuty, Go, Java, Python, TypeScript
3w
Save
Mark Applied
Hide
Senior Observability Engineer
Woodland Hills, California, United States
$120k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
6+ YOE6+ years experience in observability/telemetry architecture, alerting optimization, dynamic baselining, SLO governance, and telemetry data pipelines.
AWS, Azure, GCP, EKS, AKS, Lambda, RDS, S3, ADLS, Open Telemetry (OTel), Guidewire, Salesforce, Earnix, Uniphore
1mo
Save
Mark Applied
Hide
Senior Engineer, Network Observability
Warsaw or Bellevue or Ireland or Livingston or London or New York or Sunnyvale
OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Cloud computing infrastructure specialized for large-scale AI workloads.
Experience building and operating network observability: Prometheus/Grafana/Alertmanager, gNMI/SNMP, Python/Go/Bash, Kubernetes, Linux networking; telemetry collectors/exporters, on-call support, and mentoring junior engineers.
Prometheus, Grafana, Alertmanager, gNMI, SNMP, Python, Golang, Go, Bash, Kubernetes, Ansible, Jinja2, TensorFlow, scikit-learn, OpenTelemetry, Jaeger, Zipkin, Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, Linux
1mo
Save
Mark Applied
Hide
Staff Engineer, Network Observability
Livingston or New York City or Sunnyvale or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
Deep experience building network observability platforms, proficiency with Python/Go/Bash, Kubernetes, networking (routing/switching), collector and persistence technologies, ability to lead cross-team initiatives and participate on on-call rotation.
gNMI, SNMP, Prometheus, OTEL, Loki, Clickhouse, Grafana, Alertmanager, Python, Go, Bash, Ansible, Jinja2, Kubernetes, Jaeger, Zipkin, OpenTelemetry, SONiC, HPE Junos, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux
1mo
Save
Mark Applied
Hide
Senior Engineer – AI & HPC Observability
San Diego or Austin
$155k-$206k/yr OnsiteFull Time
Cirrascale
Cirrascale: Provides specialized GPU-based cloud infrastructure for AI workloads.
5+ YOEBachelors in CS/CE or equivalent; 5+ years observability and distributed systems experience; 1+ year HPE OpsRamp; strong Bash and Python; experience with OpenTelemetry, Prometheus, Grafana, Datadog, ELK, ThousandEyes; cloud skills (AWS/GCP/OpenStack/Proxmox/k8s).
HPE OpsRamp, Open Telemetry, Prometheus, Grafana, Nagios, Datadog, ELK, Thousand Eyes, Bash, Python, AWS, GCP, OpenStack, Proxmox, k8s, Netbox, MaaS, Redfish
1d
Save
Mark Applied
Hide
Sr. Staff Observability Software Engineer
Mountain View or Seattle
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
8+ YOEBachelor's in CS/EE/Math,8+ years building large-scale distributed systems,deep observability experience (metrics,logs,tracing),SLO/KPI definition,cloud and container familiarity,programming in Go/Java/Python/Ruby.
OpenTelemetry, Prometheus, Grafana, Elastic stack, Docker, Kubernetes, Dynatrace, AppDynamics, Jaeger, Zipkin, AWS, Azure, Google Cloud Platform, Go, Java, Python, Ruby
1mo
Save
Mark Applied
Hide
Engineering Manager Observability
Mountain View or Austin or Sunnyvale or Warren or United States
$219k-$335k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
7+ Mgmt7+ years managing software or SRE teams; deep observability knowledge (logs, metrics, traces) with Prometheus, Grafana, OpenTelemetry; distributed systems architecture; programming in Go, Python, Typescript; cloud (GCP/AWS/Azure), Kubernetes, Docker; strong communication.
Prometheus, Grafana, OpenTelemetry, Go, Python, Typescript, GCP, AWS, Azure, Kubernetes, Docker, Istio, Terraform, TSDBs, CI/CD pipelines
1mo
Save
Mark Applied
Hide
Platform Engineer, APIs & Observability
San Francisco or New York City
$180k-$200k/yr HybridFull Time
StackAI
StackAI: Build and deploy custom AI agents without writing code.
4+ YOE4+ years building backend services and public APIs; strong REST/OpenAPI skills, observability and distributed tracing, Python and FastAPI, familiarity with TypeScript/Node.js, and analytics-driven metrics and reporting.
OpenAPI, Python, FastAPI, TypeScript, Node.js, OpenTelemetry, ClickHouse, Druid
3mo
Save
Mark Applied
Hide
Senior Infra Engineer: Observability
San Francisco or United States or North America
RemoteFull Time
Railway
Railway: Infrastructure platform for automated application deployment and cloud hosting.
Experience building distributed systems, observability tooling, backend services in Golang/Rust, GRPC; familiarity with Terraform and Ansible; strong communication and ownership.
Golang, Rust, GRPC, Terraform, Ansible, TypeScript, GraphQL
1mo
Save
Mark Applied
Hide
Principal Platform Engineer, Observability (CIPE)
California, United States
$147k-$238k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
7+ YOE7+ years platform/SRE or software engineering experience; deep observability expertise (OpenTelemetry, Prometheus, Jaeger); Kubernetes and Python experience; familiarity with AI coding agents (Claude, Codex); ability to design scalable telemetry, SLOs, and incident workflows.
OpenTelemetry, Prometheus, Jaeger, Alertmanager, Chronosphere, PromQL, Grafana, OpenTelemetry Collector, Grafana k6, Prometheus Blackbox Exporter, Playwright, Selenium, Python, Go, Java, Rust, Node.js, Claude, Codex, MCP, Linux, Helm, Terraform, Argo CD, Flux, GitOps, Kubernetes
23h
Save
Mark Applied
Hide
Senior Datadog Security & Observability Engineer
United States or El Dorado Hills or Chicago
RemoteFull Time
Keeper Security
Keeper Security: Provider of password management and security software solutions.
4+ YOERequires 4+ years of production Datadog experience, Cloud SIEM, log pipelines, detection engineering, AWS, scripting, APIs, infrastructure-as-code, MITRE ATT&CK, and cross-functional communication.
Datadog, Datadog Cloud SIEM, SentinelOne, Wiz, MITRE ATT&CK, Claude, ChatGPT, Python, PowerShell, Datadog APIs, Terraform, AWS, Sigma
6d
Save
Mark Applied
Hide
Senior Software Engineer, AIOps and Observability
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years product development experience with 5+ years building and operating observability platforms; proficiency with Prometheus, OpenTelemetry, EM systems; bachelor’s in CS/engineering or equivalent.
Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, BigPanda, PagerDuty, Datadog, Kubernetes, Nomad, Docker, NATS, Kafka, Go, Python, Java, C#
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Observability
San Francisco or United States
$177k-$365k/yr HybridFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
7+ YOE7+ years in distributed systems and data engineering; expert in Java, Python, Go, or Scala; strong observability (metrics/logs/traces) with OpenTelemetry/Prometheus/Grafana; experience building scalable observability platforms; cloud-native (Kubernetes); product mindset and collaboration skills.
Java, Python, Go, Scala, Kafka, Flink, OpenTelemetry, Prometheus, Grafana, Kubernetes
1mo
Save
Mark Applied
Hide
Staff Engineer – Observability Platform
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman
Postman: Platform for building, testing, and managing software APIs.
10+ YOE10+ years software engineering experience with distributed systems, observability expertise, production operations, strong programming in Go/Java/Python/Node.js, and experience with monitoring/logging/tracing tools.
Go, Java, Python, Node.js, OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb
1w
Save
Mark Applied
Hide
Staff Software Engineer, Observability
Menlo Park or Toronto
HybridFull Time
Robinhood
RobinhoodNASDAQ: HOOD: Provides a commission-free platform for investing and financial services.
8+ YOE8+ years software engineering experience, deep Kubernetes and public cloud (AWS) expertise, strong coding in Go/Python, experience owning observability/control-plane and telemetry pipelines.
Kubernetes, AWS, Go, Python, Vector, Prometheus, Grafana, Honeycomb, Humio, Sentry
2mo
Save
Mark Applied
Hide
Sr. Product Engineer- Observability, ArcGIS Enterprise
Redlands, California, United States
$94k-$159k/yr OnsiteFull Time
Esri
Esri: Developing software for digital mapping and spatial data analysis.
5+ YOE5+ years’ experience with ArcGIS Enterprise; strong observability knowledge; Prometheus/Grafana; test automation; Python/Java; degree in CS or related field; strong problem solving and teamwork.
Prometheus, Grafana, Python, Java, Scripting, ArcGIS Enterprise, Kubernetes, AWS, Azure
2mo
Save
Mark Applied
Hide
Senior Software Engineer - Internal Observability
Menlo Park, California, United States
$200k-$288k/yr OnsiteFull Time
Snowflake
SnowflakeNYSE: SNOW: Cloud-based platform for data storage, processing, and analytics.
7+ YOE7+ years software engineering experience with distributed systems, cloud services, strong programming in Java/Scala/C++/Python, observability and telemetry expertise, and technical leadership.
Java, Scala, C++, Python, OpenTelemetry, AWS, Azure, GCP
2mo
Save
Mark Applied
Hide
Sr. Staff Software Engineer — Observability, Insights & Governance
Bellevue or Seattle or Mountain View
$217k-$299k/yr HybridFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years building large-scale distributed systems; technical leadership; CS fundamentals; observability/telemetry, query engines, data governance; BS in CS; MS/PhD a plus.
Observability tooling, Telemetry pipelines, Query engines, Distributed tracing, Metrics, Time-series storage, Data governance, Logging systems, LLMs (agentic experiences)
4w
Save
Mark Applied
Hide
Sr. Software Engineer, Observability - Slack
San Francisco or Atlanta or Seattle
$173k-$260k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Provides cloud-based customer relationship management and enterprise software.
Experience building and maintaining high-volume log pipelines and distributed observability services; strong communication, mentoring, testing, debugging, and security knowledge; familiarity with observability tooling.
Astra, Prometheus, Go, Python, Java, Elasticsearch, Logstash, Kibana, AWS
6d
Save
Mark Applied
Hide
Senior Software Engineer, AIOps and Observability
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years product development experience, 5+ years building/operating observability platforms, Bachelor's in CS/Engineering or equivalent, experience with Prometheus/Grafana/OpenTelemetry, Kubernetes, cloud/on-prem observability, and Go/Python/Java/C#.
Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, BigPanda, PagerDuty, Datadog, Kubernetes, Nomad, Docker, NATS, Kafka, Go, Python, Java, C#