350 observability engineer jobs at 160 companies in Aptos, CA

1mo
Save
Mark Applied
Hide
Senior Engineer, Network Observability
Warsaw or Bellevue or Ireland or Livingston or London or New York or Sunnyvale
OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Cloud computing infrastructure specialized for large-scale AI workloads.
Experience building and operating network observability: Prometheus/Grafana/Alertmanager, gNMI/SNMP, Python/Go/Bash, Kubernetes, Linux networking; telemetry collectors/exporters, on-call support, and mentoring junior engineers.
Prometheus, Grafana, Alertmanager, gNMI, SNMP, Python, Golang, Go, Bash, Kubernetes, Ansible, Jinja2, TensorFlow, scikit-learn, OpenTelemetry, Jaeger, Zipkin, Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, Linux
4d
Save
Mark Applied
Hide
Sr. Staff Observability Software Engineer
Mountain View or Seattle
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
8+ YOEBachelor's in CS/EE/Math,8+ years building large-scale distributed systems,deep observability experience (metrics,logs,tracing),SLO/KPI definition,cloud and container familiarity,programming in Go/Java/Python/Ruby.
OpenTelemetry, Prometheus, Grafana, Elastic stack, Docker, Kubernetes, Dynatrace, AppDynamics, Jaeger, Zipkin, AWS, Azure, Google Cloud Platform, Go, Java, Python, Ruby
1mo
Save
Mark Applied
Hide
Engineering Manager Observability
Mountain View or Austin or Sunnyvale or Warren or United States
$219k-$335k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
7+ Mgmt7+ years managing software or SRE teams; deep observability knowledge (logs, metrics, traces) with Prometheus, Grafana, OpenTelemetry; distributed systems architecture; programming in Go, Python, Typescript; cloud (GCP/AWS/Azure), Kubernetes, Docker; strong communication.
Prometheus, Grafana, OpenTelemetry, Go, Python, Typescript, GCP, AWS, Azure, Kubernetes, Docker, Istio, Terraform, TSDBs, CI/CD pipelines
1w
Save
Mark Applied
Hide
Senior Software Engineer, AIOps and Observability
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years product development experience with 5+ years building and operating observability platforms; proficiency with Prometheus, OpenTelemetry, EM systems; bachelor’s in CS/engineering or equivalent.
Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, BigPanda, PagerDuty, Datadog, Kubernetes, Nomad, Docker, NATS, Kafka, Go, Python, Java, C#
2mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability (193698)
Fremont or San Francisco or Oakland
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Manufacturing equipment used to fabricate advanced semiconductor microchips.
12+ YOE6+ MgmtBS/MS/PhD or equivalent, 12+ years in infrastructure/SRE/DevOps/network engineering, 6+ years leading SRE/observability teams; multi-cloud networking, DR/BCP, observability platforms, IaC, automation, Python/Go experience.
Azure, AWS, GCP, Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Go, Kubernetes, AKS, EKS, GKE, ServiceNow
1w
Save
Mark Applied
Hide
Staff Software Engineer, Observability
Menlo Park or Toronto
HybridFull Time
Robinhood
RobinhoodNASDAQ: HOOD: Provides a commission-free platform for investing and financial services.
8+ YOE8+ years software engineering experience, deep Kubernetes and public cloud (AWS) expertise, strong coding in Go/Python, experience owning observability/control-plane and telemetry pipelines.
Kubernetes, AWS, Go, Python, Vector, Prometheus, Grafana, Honeycomb, Humio, Sentry
2mo
Save
Mark Applied
Hide
Sr. Staff Software Engineer — Observability, Insights & Governance
Bellevue or Seattle or Mountain View
$217k-$299k/yr HybridFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years building large-scale distributed systems; technical leadership; CS fundamentals; observability/telemetry, query engines, data governance; BS in CS; MS/PhD a plus.
Observability tooling, Telemetry pipelines, Query engines, Distributed tracing, Metrics, Time-series storage, Data governance, Logging systems, LLMs (agentic experiences)
1w
Save
Mark Applied
Hide
Senior Software Engineer, AIOps and Observability
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years product development experience, 5+ years building/operating observability platforms, Bachelor's in CS/Engineering or equivalent, experience with Prometheus/Grafana/OpenTelemetry, Kubernetes, cloud/on-prem observability, and Go/Python/Java/C#.
Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, BigPanda, PagerDuty, Datadog, Kubernetes, Nomad, Docker, NATS, Kafka, Go, Python, Java, C#
2w
Save
Mark Applied
Hide
Engineering Manager, Observability
Sunnyvale, California, United States
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE2+ Mgmt5+ years software engineering, 2+ years engineering management, experience with observability platforms, reliability engineering, scaling telemetry, and hiring/managing teams.
OpenTelemetry, Grafana, Prometheus, Kubernetes
1mo
Save
Mark Applied
Hide
Principal Software Development Engineer - Observability
San Jose or Seattle
$249k-$349k/yr OnsiteFull Time
Expedia Group
Expedia GroupNASDAQ: EXPE: Operates a global platform for travel bookings and services.
10+ YOE10+ years building and operating large-scale distributed systems; deep observability expertise (logs, metrics, traces); proficiency in Go/Java/Python, cloud-native tech, OpenTelemetry, and observability tooling; leadership and architecture experience.
Go, Java, Python, AWS, Kubernetes, Docker, OpenTelemetry, Prometheus, Grafana, Datadog, Splunk, Clickhouse, Terraform, Crossplane, Electronic Medical Records (EMR)
2w
Save
Mark Applied
Hide
Staff Software Engineer - Observability
Mountain View or San Diego
OnsiteFull Time
Intuit
IntuitNASDAQ: INTU: Provides financial software for accounting, tax, and personal finance.
8+ YOE8+ years building and operating large-scale distributed systems; architect-level Splunk experience; public cloud (AWS/GCP); Java/Go/Python; CI/CD and infrastructure-as-code; bachelor's degree in CS/engineering.
Splunk, S3, Kinesis, CloudWatch, Fluent Bit, GCP Logs Processor, MCP Server, Java, Go, Python, Git, CI/CD, AWS, GCP
2w
Save
Mark Applied
Hide
Principal Software Engineer – Sensor, Telemetry & Observability (Hybrid)
Sunnyvale or Redmond
$195k-$290k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
Architect and deliver at-scale telemetry and observability across endpoint sensors and cloud ingestion; strong systems engineering and cross-platform experience; mentoring and technical leadership expected.
C/C++, ETW, PMU, KQL, SQL, Snowflake, Go
1w
Save
Mark Applied
Hide
Senior/Lead Site Reliability Engineer Observability
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOE7+ years SRE/Platform/DevOps experience; hands-on Splunk, ELK/Elasticsearch, Prometheus, Grafana, Tempo, OpenTelemetry, Kafka; Terraform, Kubernetes, cloud, Python/Go/Ruby/Bash; Splunk certification; ability to obtain US security clearance.
Splunk Enterprise, Splunk Cloud, Splunk SPL, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Azure, GCP, Ansible, Consul, CI/CD pipelines, service mesh
2mo
Save
Mark Applied
Hide
Engineering Manager, Observability Infrastructure
San Mateo, California, United States
$295k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
3+ Mgmt3+ years engineering management experience; strong background in building and operating large-scale distributed systems and data infrastructure; experience with observability or AI/ML infrastructure preferred; excellent communication and cross-functional collaboration skills.
2mo
Save
Mark Applied
Hide
Senior Software Engineer (Observability Solutions) - Enterprise Technology Services
Sunnyvale or United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
8+ YOEBachelor's in CS/CE or related field; strong Go backend experience; distributed systems; 8+ years; petabyte-scale data workloads.
Go, Distributed Systems, Prometheus, Elasticsearch, APM
1mo
Save
Mark Applied
Hide
Principal Engineer - Isovalent Secure Workload Observability (remote)
Milpitas or San Jose or Sunnyvale or Mountain View or San Francisco or Palo Alto
$231k-$332k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
10+ YOE10+ years building software (Go, C++, Java), 5+ years technical leadership, degree in computer science/engineering or related field, experience with scalable distributed systems, Kubernetes, cloud APIs, algorithms and performance engineering.
Go, Golang, C++, Java, Kubernetes, AWS, Azure, GCP, eBPF, Cilium
3mo
Save
Mark Applied
Hide
Senior Software Engineer, Platform Engineering
San Mateo, California, United States
$130k-$200k/yr OnsiteFull Time
IXL Learning
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOE6+ years in software/platform engineering; proficient in Java/C++/Python/Go; cloud (AWS/GCP); Docker/Kubernetes; backend systems; observability; strong collaboration; quick learner.
Java, C++, Python, Go, AWS, GCP, Docker, Kubernetes, Observability
2mo
Save
Mark Applied
Hide
Frontend Engineer, Media Infra Systems & Observability - L4
Los Gatos, California, United States
$250k-$413k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
Strong frontend skills in React, TypeScript, GraphQL; experience with data-heavy UIs; ability to define requirements with engineering partners; familiar with backend reading/debugging.
React, TypeScript, GraphQL
3mo
Save
Mark Applied
Hide
Tech Lead – Network Observability
Palo Alto, California, United States
$180k-$260k/yr OnsiteFull Time
Clockwork Systems
Clockwork Systems: Software-driven network fabrics for GPU cluster performance optimization.
Lead architecture and development of a high-performance network observability platform; strong distributed systems, Linux networking, and observability tooling; mentor engineers.
Prometheus, Grafana, Datadog, OpenTelemetry, NCCL, rdma-core, libibverbs, tcpdump, Wireshark, ethtool, iproute2, Linux kernel networking tools
2w
Save
Mark Applied
Hide
Staff Cloud Engineer
Milpitas, California, United States
$131k-$223k/yr OnsiteFull Time
Sandisk
SandiskNasdaq: SNDK: Designs and manufactures flash memory and data storage products.
10+ YOEBachelor's in CS/Engineering, 10+ years experience with 7+ in cloud engineering/migrations, deep AWS expertise, Terraform, Kubernetes/EKS, GitHub/GitHub Actions, Python/Bash, observability, and large-scale migration experience.
AWS, Kubernetes, EKS, Terraform, Python, Bash, GitHub, GitHub Actions, JFrog Artifactory, CloudWatch, ArgoCD, AWS Application Migration Service, DMS