1,182 site reliability engineer jobs at 578 companies in United States
1w
Save
Mark Applied
Hide
1w
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
San Jose, California, United States
OnsiteInternship
ByteDance: Global technology specializing in AI-powered content platforms.
Currently pursuing a bachelor's degree in computer science or related technical discipline; programming experience in C, C++, Java, Python, Go, or Rust; knowledge of Unix/Linux internals, networking, and distributed systems.
Electrolux GroupNasdaq Stockholm: ELUX B: Global home appliance manufacturer reinventing taste, care, and wellbeing.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
Seekr Technologies: Private American enterprise AI providing explainable, secure AI software and hardware to government and critical-infrastructure customers.
5+ YOERequires 5+ years in site reliability engineering and Linux systems, monitoring and logging expertise, programming or scripting skills, Docker, Kubernetes, configuration automation, and incident response experience.
Enzo Health: AI-native home health EHR and operations platform for U.S. home health agencies.
5+ YOEAt least 5 years in site reliability, platform, infrastructure, or production engineering; hands-on AWS, Kubernetes, Terraform, Postgres, CI/CD, observability, automation, and incident response experience.
PRADCO Outdoor Brands: Family-owned U.S. manufacturer of hunting, fishing, and pet products for hunters, anglers, and pet owners.
3+ YOERequires 3–5 years in site reliability or similar engineering, Azure expertise, distributed-systems support, CI/CD, performance diagnostics, automation, and a bachelor's degree or equivalent experience.
Microsoft Azure, App Service Plans, Azure Functions, Event Grid, Event Hub, Service Bus, Azure SQL Service, Azure DevOps, CI/CD, SQL
Oracle CorporationNYSE: ORCL: Cloud infrastructure and enterprise software solutions provider.
3+ YOEBachelor’s degree in Computer Science or equivalent experience; 3+ years in Site Reliability Engineering, DevOps, or Systems Engineering; cloud operations, incident management, automation, programming, and infrastructure tooling experience.
CognizantNASDAQ: CTSH: Global professional services providing technology and consulting services.
3+ YOERequires 3+ years in software engineering focused on reliability, infrastructure, or platforms; Java, microservices, MVC, JDBC, REST, frameworks, Azure, SRE, troubleshooting, and on-call support.
Exiger: Private supply-chain AI software serving corporations, government agencies, and banks with risk and compliance technology.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
ShopifyNasdaq: SHOP: Provides internet infrastructure and tools for commerce.
Experienced SRE/engineer with on-call experience, ability to build resilient production tooling, improve observability, respond to alerts, and collaborate across engineering teams.
Supabase: Developer platform providing Postgres databases, authentication, storage, realtime, REST APIs, and edge functions for application developers.
7+ YOE7+ years in SRE/production engineering, experience shaping SRE practices, defining and operationalizing SLOs/SLIs, incident response and postmortems, software engineering mindset, cloud infra (AWS) and IaC (Pulumi/Terraform/CDK).
Akamai TechnologiesNASDAQ: AKAM: Cloud and edge computing platform for secure digital experiences.
Bachelor's in Computer Science/Engineering or equivalent; experience in SRE/Software Engineering for large-scale distributed systems; Terraform and IAC experience; familiarity with SaltStack/Ansible/Chef/Puppet; Linux, CI/CD, observability, and participation in on-call rotation.
QualityAI: AI-first quality engineering and software testing firm.
6+ YOE6+ years SRE experience with AWS, monitoring/observability, Linux/Unix, scripting, CI/CD, and building automated reliability and performance solutions.
Runpod: AI cloud computing platform providing on-demand GPUs and serverless compute to developers, researchers, and AI companies.
5+ YOE5+ years SRE or production engineering experience; strong Linux, networking, container, distributed systems, SLI/SLO, incident response, and scripting skills.
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.