OpenMind: Private robotics software building AI infrastructure and robot orchestration for human-facing service and industrial fleets.
Experience in reliability engineering, test infrastructure, SRE, or autonomous-system simulation; strong Python, Go, or C++ fundamentals; robotics simulation experience; and expertise in measurement and failure analysis.
Seattle or San Francisco or Detroit or United States
$180k-$279k/yrHybridFull Time
Rocket Homes: Technology-driven real estate service provider connecting home buyers and sellers with listings and real estate agents.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
Austin or New York City or San Francisco or Seattle
$203k-$232k/yrOnsiteFull Time
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
United States or Austin or New York City or Boston or San Francisco
$125k-$130k/yrRemoteFull Time
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
5+ YOERequires 5 years of experience with complex cloud infrastructure, 3 years with Kubernetes, production distributed systems, Linux, monitoring, troubleshooting, customer support, DevOps or CI/CD, and Python scripting.
TikTok: Short-form mobile video and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Console: AI-native IT service-management software platform automating employee support requests for companies.
5+ YOE5+ years infrastructure/platform/backend experience; hands-on AWS and Kubernetes; experience with Pulumi or Terraform; comfortable in TypeScript/Node/React codebases; production reliability, observability, and enterprise deployment experience.
AWS, Kubernetes, Pulumi, Terraform, TypeScript, Node, React, Slack, Microsoft Teams
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yrRemoteFull Time
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
2+ YOEAssociate's degree or equivalent experience, 2 years in semiconductor lab or manufacturing, 1 year troubleshooting electromechanical systems, and precision soldering and mechanical assembly experience.
Manufacturing Execution System (MES), Work-in-Progress (WIP), Python, Linux, Unix, Google Workspace, Google Docs, Google Sheets, Google Slides
Ayar Labs: -packaged optics providing connectivity for hyperscale AI infrastructure.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yrHybridFull Time
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
CrewAI: Private software providing multi-agent AI orchestration and enterprise workflow automation for businesses.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
LangChain: Private AI software providing agent-engineering platforms and open-source frameworks for developers and enterprises.
5+ YOERequires 5+ years in infrastructure, platform engineering, or SRE; Kubernetes and cloud infrastructure expertise; scripting or systems programming; infrastructure-as-code, CI/CD, reliability engineering, and production stateful workload experience.
Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Sunnyvale, California, United States
$104k-$171k/yrOnsiteFull Time
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
2+ YOE2+ years in distributed systems reliability engineering; high-availability architecture, Kafka/RocketMQ, Kubernetes, automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed.
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
1+ YOEBachelor's degree or equivalent experience, 1+ year in SRE, DevOps, or systems engineering, Linux and networking knowledge, distributed systems experience, programming, scripting, CI/CD, and automation skills.
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance: Global technology specializing in AI-powered content platforms.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale: AI infrastructure software helping developers and AI teams build, deploy, and scale machine-learning workloads with Ray.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.