Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
San Mateo or Arizona or California or Colorado or Florida or Georgia or Illinois or Nevada or North Carolina or Oregon or Texas or Utah or Washington
$140k-$150k/yrRemoteFull Time
VyncaCare: Offers palliative care services and advance care planning technology.
3+ YOE3+ years SRE/DevOps experience, strong AWS and Terraform skills, Kubernetes and Helm experience, observability and incident response knowledge, bachelor's or equivalent, on-call participation, East Coast hours.
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years SRE or software engineering experience; Bachelor’s in Computer Science or equivalent; fluency in Go, Java, or C#; experience with Kubernetes, Nomad, Vault, and Consul; strong reliability and observability practices.
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.