62 infrastructure reliability engineer jobs at 32 companies in Aromas, CA
1mo
Save
Mark Applied
Hide
1mo
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
6+ YOEBS in CS/CE or equivalent experience,6+ years in Linux-based infrastructure,4+ years hardware troubleshooting,experience with hardware qualification/integration,firmware/BMC knowledge,and automation for infrastructure at scale.
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
8+ YOEBachelor's in CS or equivalent,8+ years building infrastructure/distributed systems,5+ years programming in C++ or Go,5+ years reliability engineering,EMR not mentioned,experience with distributed systems and stakeholder collaboration.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOE5+ years PostgreSQL and production data-system experience; 3+ years Linux engineering and infrastructure automation; 2+ years Python, Bash, Go, Ruby, or Perl; cloud, Kubernetes, and distributed data-system experience.
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Model AI: Building high-performance infrastructure for agentic AI systems.
Strong ML systems and distributed systems experience; optimizing large-model inference, accelerator utilization, batching, scheduling, KV cache and runtime efficiency; ability to ship reliable production systems.
Software Engineer - Backend Infrastructure, Standalone Apps Team
Menlo Park, California, United States
$219k-$301k/yrOnsiteFull Time
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
12+ YOEBachelor's degree or equivalent, 12+ years software engineering experience with infrastructure/platform/distributed systems, strong distributed systems/storage/reliability knowledge, track record operating infrastructure at scale.
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years in infrastructure/platform/backend engineering; fluent in Go or similar; AWS, containerization, and IaC experience (Terraform or Pulumi); SRE concepts (SLOs, error budgets); platform engineering mindset and familiarity with AI tools.