Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Ivo: AI-powered contract review and intelligence platform for legal teams.
5+ YOEMinimum 5 years experience; own uptime and reliability, define SLIs/SLOs/SLAs, design failover and disaster recovery, implement security controls, lead incident response; familiarity with cloud and LLM-driven systems.
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yrRemoteFull Time
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Coupa: Cloud-based platform for managing and optimizing business expenditures.
8+ YOE8+ years hands-on DBA experience, deep MySQL expertise, scripting (Bash/Python/Ruby), cloud (AWS/RDS/Aurora) and automation experience, monitoring and HA/DR skills, ability to lead architecture and mentor engineers.
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
Observability Lead - Cloud SRE & Network Reliability (193698)
Fremont or San Francisco or Oakland
$114k-$253k/yrHybridFull Time
Lam ResearchNASDAQ: LRCX: Manufacturing equipment used to fabricate advanced semiconductor microchips.
12+ YOE6+ MgmtBS/MS/PhD or equivalent, 12+ years in infrastructure/SRE/DevOps/network engineering, 6+ years leading SRE/observability teams; multi-cloud networking, DR/BCP, observability platforms, IaC, automation, Python/Go experience.
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
6+ YOE6+ years SRE/DevOps experience preferred; strong Kubernetes, IaC (Terraform/Pulumi), cloud networking, Go or Python, CI/CD, observability, and distributed systems experience; blockchain node operation experience is a plus.
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
7+ YOEU.S. citizen required. 7+ years SRE/platform/cloud experience; B.S. in CS/Engineering or equivalent; experience with large-scale cloud production systems, SLOs/SLIs, observability, incident management, automation, and IaC. Programming in Python/Go/Java/PowerShell/Bash.
Charlotte or Chandler or San Francisco or Columbus
$119k-$224k/yrHybridFull Time
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
5+ YOERequires 5+ years in systems engineering or architecture, 5+ years SRE, cloud observability experience, hosting platforms, and familiarity with DevOps, Agile, and IT service management.