216 observability engineer jobs at 67 companies in Marina, CA
1mo
Save
Mark Applied
Hide
1mo
Senior Engineer, Network Observability
Warsaw or Bellevue or Ireland or Livingston or London or New York or Sunnyvale
OnsiteFull Time
CoreWeaveNasdaq: CRWV: Cloud computing infrastructure specialized for large-scale AI workloads.
Experience building and operating network observability: Prometheus/Grafana/Alertmanager, gNMI/SNMP, Python/Go/Bash, Kubernetes, Linux networking; telemetry collectors/exporters, on-call support, and mentoring junior engineers.
Prometheus, Grafana, Alertmanager, gNMI, SNMP, Python, Golang, Go, Bash, Kubernetes, Ansible, Jinja2, TensorFlow, scikit-learn, OpenTelemetry, Jaeger, Zipkin, Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, Linux
Livingston or New York City or Sunnyvale or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
Deep experience building network observability platforms, proficiency with Python/Go/Bash, Kubernetes, networking (routing/switching), collector and persistence technologies, ability to lead cross-team initiatives and participate on on-call rotation.
gNMI, SNMP, Prometheus, OTEL, Loki, Clickhouse, Grafana, Alertmanager, Python, Go, Bash, Ansible, Jinja2, Kubernetes, Jaeger, Zipkin, OpenTelemetry, SONiC, HPE Junos, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
7+ YOE7+ years of production experience with distributed applications; large-scale observability platform experience; BS in Computer Science or related field; strong collaboration in distributed teams.
SynopsysNasdaq: SNPS: Provides software and IP for semiconductor design and manufacturing.
8+ YOE8+ years in software/platform/SRE or infrastructure engineering with experience building observability capabilities; hands-on with Elastic, Grafana, Kafka, Logstash, OpenTelemetry, Prometheus; scripting in Python/Ruby/Bash; Linux, Kubernetes, Ansible; Bachelor's degree required.
Lisbon or San Francisco or Sunnyvale or Raleigh or Seattle or Boston or London or Bengaluru or Dublin or Kyiv or United States or Portugal
RemoteFull Time
SingleStore: Provides a distributed SQL database for real-time analytics.
2+ YOE2+ years building distributed systems or backend services; strong Go (Golang) proficiency; familiarity with Kubernetes, cloud providers (AWS/GCP/Azure), observability concepts, and debugging production systems.
Senior Systems Software Engineer, Observability and Telemetry Platform
Santa Clara or United States
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years in infrastructure automation and distributed systems; BS in CS or related (or equivalent); experience with observability, capacity, Linux, containers, and infrastructure tooling.
Senior Systems Software Engineer, Observability and Telemetry Platform
Santa Clara or United States
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEBS in CS or equivalent; 8+ years infrastructure automation/distributed systems experience; 5+ years building observability platforms; proficiency with Linux, networking, containers and coding in Python/Go/Perl/Ruby.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
10+ YOE10+ years software engineering/architecture experience; deep Java expertise; cloud native and distributed systems design; experience with telemetry, TSDBs, real-time streaming (Kafka), SQL, Kubernetes, and Unix/Linux.
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
Experienced engineering manager to lead cross-domain squads in release delivery and real-time observability; preferred 10+ years engineering and 5+ years people management; experience with distributed systems, data pipelines, and CI/CD.
Frontend Engineer, Media Infra Systems & Observability - L4
Los Gatos, California, United States
$250k-$413k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global video streaming and media production service.
Strong frontend skills in React, TypeScript, GraphQL; experience with data-heavy UIs; ability to define requirements with engineering partners; familiar with backend reading/debugging.
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOE5+ years software engineering experience, deep CI/CD and automation expertise, proficiency in Python/Go/Rust, strong Linux/Unix fundamentals, containerization (Kubernetes preferred), observability and on-call experience.
Mountain View or Menlo Park or San Jose or Sacramento
$120k-$140k/yrHybridFull Time
Aro Homes: Designs and builds sustainable, precision-engineered modular homes.
8+ YOECalifornia-licensed Civil PE, 8+ years designing light-framed wood systems, architectural steel, and foundations; Revit/BIM experience; ability to create engineering tools/processes and perform site observations.
McAfee: Provides cybersecurity and privacy protection software for consumers.
8+ YOE8+ years in DevOps/platform engineering with production Kubernetes (EKS/GKE), Terraform, CI/CD (GitHub Actions/Jenkins/Harness), Backstage, Golang or Python, DevSecOps, service meshes (Istio/Envoy), Kafka, and platform observability.
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
7+ YOE7+ years building production software, hands-on experience shipping generative AI products, deep knowledge of LLM behavior, prompt engineering, eval engineering, distributed systems, reliability, and production observability.
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
2+ YOEBachelor's degree in engineering, 2+ years test-infrastructure or HIL experience, proficiency with observability tools, Python/Bash scripting, Linux, CI/CD, strong problem solving and communication skills.
Datadog, Grafana, Prometheus, Python, Bash, Linux, Unix, Ansible, Chef, Terraform, Kubernetes, Git, AUTOSAR, Vector, dSPACE, ASPICE, ISO 26262, ISO 21448
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Vistance NetworksNASDAQ: VISN: Vistance Networks provides infrastructure solutions for communications and data networks.
5+ YOE5+ years SRE/DevOps or production engineering experience; strong Python and Linux skills; experience with GCP, Kubernetes, observability tools (Prometheus, Grafana, OpenTelemetry, ELK), ClickHouse; incident response and distributed systems troubleshooting.