SandiskNasdaq: SNDK: Designs and manufactures flash memory and data storage products.
6+ YOEBachelor's degree in engineering/CS, 6+ years engineering experience in test infrastructure or reliability, hands-on with environmental chambers, failure analysis, vendor management, and global lab alignment.
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Seattle or San Francisco or Detroit or United States
$180k-$279k/yrHybridFull Time
Rocket CompaniesNYSE: RKT: Provides digital mortgage, real estate, and personal finance services.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
Austin or New York City or San Francisco or Seattle
$203k-$232k/yrOnsiteFull Time
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
Console: Automates IT support and internal operations using AI agents.
5+ YOE5+ years infrastructure/platform/backend experience; hands-on AWS and Kubernetes; experience with Pulumi or Terraform; comfortable in TypeScript/Node/React codebases; production reliability, observability, and enterprise deployment experience.
AWS, Kubernetes, Pulumi, Terraform, TypeScript, Node, React, Slack, Microsoft Teams
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yrRemoteFull Time
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure; proficient in Python; strong cloud, Kubernetes, networking, storage, and automation; focus on reliability.
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.