316 observability engineer jobs at 149 companies in Soquel, CA
1mo
Save
Mark Applied
Hide
1mo
Senior Engineer, Network Observability
Warsaw or Bellevue or Ireland or Livingston or London or New York or Sunnyvale
OnsiteFull Time
CoreWeaveNasdaq: CRWV: Cloud computing infrastructure specialized for large-scale AI workloads.
Experience building and operating network observability: Prometheus/Grafana/Alertmanager, gNMI/SNMP, Python/Go/Bash, Kubernetes, Linux networking; telemetry collectors/exporters, on-call support, and mentoring junior engineers.
Prometheus, Grafana, Alertmanager, gNMI, SNMP, Python, Golang, Go, Bash, Kubernetes, Ansible, Jinja2, TensorFlow, scikit-learn, OpenTelemetry, Jaeger, Zipkin, Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, Linux
Livingston or New York City or Sunnyvale or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
Deep experience building network observability platforms, proficiency with Python/Go/Bash, Kubernetes, networking (routing/switching), collector and persistence technologies, ability to lead cross-team initiatives and participate on on-call rotation.
gNMI, SNMP, Prometheus, OTEL, Loki, Clickhouse, Grafana, Alertmanager, Python, Go, Bash, Ansible, Jinja2, Kubernetes, Jaeger, Zipkin, OpenTelemetry, SONiC, HPE Junos, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux
SynopsysNasdaq: SNPS: Provides software and IP for semiconductor design and manufacturing.
8+ YOE8+ years in software/platform/SRE or infrastructure engineering with experience building observability capabilities; hands-on with Elastic, Grafana, Kafka, Logstash, OpenTelemetry, Prometheus; scripting in Python/Ruby/Bash; Linux, Kubernetes, Ansible; Bachelor's degree required.
Lisbon or San Francisco or Sunnyvale or Raleigh or Seattle or Boston or London or Bengaluru or Dublin or Kyiv or United States or Portugal
RemoteFull Time
SingleStore: Provides a distributed SQL database for real-time analytics.
2+ YOE2+ years building distributed systems or backend services; strong Go (Golang) proficiency; familiarity with Kubernetes, cloud providers (AWS/GCP/Azure), observability concepts, and debugging production systems.
RobinhoodNASDAQ: HOOD: Provides a commission-free platform for investing and financial services.
8+ YOE8+ years software engineering experience, deep Kubernetes and public cloud (AWS) expertise, strong coding in Go/Python, experience owning observability/control-plane and telemetry pipelines.
SnowflakeNYSE: SNOW: Cloud-based platform for data storage, processing, and analytics.
7+ YOE7+ years software engineering experience with distributed systems, cloud services, strong programming in Java/Scala/C++/Python, observability and telemetry expertise, and technical leadership.
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years building large-scale distributed systems; technical leadership; CS fundamentals; observability/telemetry, query engines, data governance; BS in CS; MS/PhD a plus.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years product development experience, 5+ years building/operating observability platforms, Bachelor's in CS/Engineering or equivalent, experience with Prometheus/Grafana/OpenTelemetry, Kubernetes, cloud/on-prem observability, and Go/Python/Java/C#.
Principal Software Development Engineer - Observability
San Jose or Seattle
$249k-$349k/yrOnsiteFull Time
Expedia GroupNASDAQ: EXPE: Operates a global platform for travel bookings and services.
10+ YOE10+ years building and operating large-scale distributed systems; deep observability expertise (logs, metrics, traces); proficiency in Go/Java/Python, cloud-native tech, OpenTelemetry, and observability tooling; leadership and architecture experience.
Go, Java, Python, AWS, Kubernetes, Docker, OpenTelemetry, Prometheus, Grafana, Datadog, Splunk, Clickhouse, Terraform, Crossplane, Electronic Medical Records (EMR)
Blackhawk Network: Provider of global branded payment and gift card solutions.
6+ YOE6+ years in platform engineering/SRE/DevOps/architecture; experience designing large-scale observability for cloud/AWS; expertise with New Relic, Splunk, Datadog, Coralogix; OpenTelemetry and observability pipelines; strong communication.
IntuitNASDAQ: INTU: Provides financial software for accounting, tax, and personal finance.
8+ YOE8+ years building and operating large-scale distributed systems; architect-level Splunk experience; public cloud (AWS/GCP); Java/Go/Python; CI/CD and infrastructure-as-code; bachelor's degree in CS/engineering.
Principal Software Engineer – Sensor, Telemetry & Observability (Hybrid)
Sunnyvale or Redmond
$195k-$290k/yrHybridFull Time
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
Architect and deliver at-scale telemetry and observability across endpoint sensors and cloud ingestion; strong systems engineering and cross-platform experience; mentoring and technical leadership expected.
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
3+ Mgmt3+ years engineering management experience; strong background in building and operating large-scale distributed systems and data infrastructure; experience with observability or AI/ML infrastructure preferred; excellent communication and cross-functional collaboration skills.
Clockwork Systems: Software-driven network fabrics for GPU cluster performance optimization.
Lead architecture and development of a high-performance network observability platform; strong distributed systems, Linux networking, and observability tooling; mentor engineers.
SandiskNasdaq: SNDK: Designs and manufactures flash memory and data storage products.
10+ YOEBachelor's in CS/Engineering, 10+ years experience with 7+ in cloud engineering/migrations, deep AWS expertise, Terraform, Kubernetes/EKS, GitHub/GitHub Actions, Python/Bash, observability, and large-scale migration experience.
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOE5+ years software engineering experience, deep CI/CD and automation expertise, proficiency in Python/Go/Rust, strong Linux/Unix fundamentals, containerization (Kubernetes preferred), observability and on-call experience.