OpenAI: Develops artificial intelligence models and generative AI software services.
Strong software engineering experience; Python or Go; IaC (Terraform); data pipelines for security; collaborate with security and engineering teams; proactive problem solving.
StackAI: Build and deploy custom AI agents without writing code.
4+ YOE4+ years building backend services and public APIs; strong REST/OpenAPI skills, observability and distributed tracing, Python and FastAPI, familiarity with TypeScript/Node.js, and analytics-driven metrics and reporting.
Railway: Infrastructure platform for automated application deployment and cloud hosting.
Experience building distributed systems, observability tooling, backend services in Golang/Rust, GRPC; familiarity with Terraform and Ansible; strong communication and ownership.
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
7+ YOE7+ years in distributed systems and data engineering; expert in Java, Python, Go, or Scala; strong observability (metrics/logs/traces) with OpenTelemetry/Prometheus/Grafana; experience building scalable observability platforms; cloud-native (Kubernetes); product mindset and collaboration skills.
Lisbon or San Francisco or Sunnyvale or Raleigh or Seattle or Boston or London or Bengaluru or Dublin or Kyiv or United States or Portugal
RemoteFull Time
SingleStore: Provides a distributed SQL database for real-time analytics.
2+ YOE2+ years building distributed systems or backend services; strong Go (Golang) proficiency; familiarity with Kubernetes, cloud providers (AWS/GCP/Azure), observability concepts, and debugging production systems.
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman: Platform for building, testing, and managing software APIs.
10+ YOE10+ years software engineering experience with distributed systems, observability expertise, production operations, strong programming in Go/Java/Python/Node.js, and experience with monitoring/logging/tracing tools.
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years in distributed systems, observability or governance; strong CS fundamentals; cross-functional communication; BS in CS (MS/PhD a plus).
Blackhawk Network: Provider of global branded payment and gift card solutions.
6+ YOE6+ years in platform engineering/SRE/DevOps/architecture; experience designing large-scale observability for cloud/AWS; expertise with New Relic, Splunk, Datadog, Coralogix; OpenTelemetry and observability pipelines; strong communication.
LaunchDarkly: Provides a feature management platform for software development teams.
5+ YOE5+ years of professional software engineering; TypeScript/React frontend and Go backend; experience with integrations, data pipelines or APIs; IaC tooling; RBAC and GitOps; strong communication.
Blackhawk Network: Provider of branded payment solutions and gift cards.
6+ YOE6+ years in platform engineering/SRE/DevOps or solution architecture; bachelor’s or equivalent; expertise with observability platforms (New Relic, Splunk, Datadog, Coralogix), OpenTelemetry and telemetry pipelines; strong distributed systems and AWS experience.
Senior Software Engineer - Observability and Reliability
New York City or San Francisco or London or Sydney
$170k-$240k/yrOnsiteFull Time
Sigma Computing: Cloud-native analytics platform featuring a spreadsheet-style interface.
5+ YOE5+ years building high-quality software, strong CS fundamentals, experience building observability tools, proficiency with Go, OpenTelemetry, Kubernetes, participation in on-call rotations, cloud service administration (GCP/AWS/Azure) preferred.
HUD: Platform for reinforcement learning environments and AI agent evaluations.
Production infra and backend experience owning uptime, performance, deployment safety, and cost; strong AWS, Kubernetes/EKS, Terraform, CI/CD, observability, and backend engineering skills.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Staff-level experience building production infrastructure software, strong distributed systems and telemetry pipeline background, networking and high-performance network knowledge, experience processing high-volume operational data.
Triumph: Skill-based mobile gaming platform for real money tournaments.
Experience operating and scaling large production systems; deep Postgres knowledge; CI/CD; observability tooling; able to lead a function independently.
Felt Technologies: Provides embedded telehealth and provider network infrastructure for applications.
Proficiency with NextJS, React, Postgres; experience with serverless environments, observability, and healthtech/PII/PHI best practices; mobile SDK experience (Swift/Kotlin) is a plus. Must work US timezones.