19 sre manager jobs at 14 companies in Pacific Grove, CA

6d
Save
Mark Applied
Hide
Senior/Lead SRE Platform Services Engineer Technical Leader
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOERequires 8–12+ years in SRE, platform engineering, infrastructure, distributed systems, or cloud operations; software engineering, infrastructure-as-code, cloud, Kubernetes, AWS, CI/CD, security, and technical leadership experience.
Python, Go, Ruby, Kubernetes, AWS, CI/CD, FedRAMP High, DoD IL5, GovCloud, FIPS
2mo
Save
Mark Applied
Hide
Technical Program Manager
San Jose, California, United States
$125k-$187k/yr HybridFull Time
F5
F5NASDAQ: FFIV: Provides application delivery networking and multi-cloud security solutions.
Lead cross-functional engineering programs end-to-end; manage scope, milestones, risks; collaborate with engineering, product, QA, security, and SRE; excellent communication.
Jira, Confluence, Smartsheet, Excel
2w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
2mo
Save
Mark Applied
Hide
Senior Incident Manager
United States or San Jose
$125k-$195k/yr RemoteFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
8+ YOE8+ years in incident management/SRE/infrastructure operations; experience with large-scale distributed infrastructure, data center operations, GPU clusters, networking, cloud platforms; incident frameworks (ITIL/SRE); strong leadership, communication, and stakeholder management.
PagerDuty, ServiceNow, Jira, Datadog, Prometheus, Grafana, Incident command system (ICS)
1w
Save
Mark Applied
Hide
IT Systems Engineer - Internal Platforms & SRE
San Francisco or San Jose
$206k-$275k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
1mo
Save
Mark Applied
Hide
SRE/Devops Engineer- San Jose, the US
San Jose, California, United States
HybridFull Time
Kody
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident management ownership, and absolute fluency in Mandarin and English.
AWS, GitHub
1mo
Save
Mark Applied
Hide
Engineering Manager - Cloud Infrastructure and Devops
San Jose or Chicago or Scottsdale or Austin
$177k-$262k/yr HybridFull Time
PayPal
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
7+ YOE2+ Mgmt7+ years software engineering experience, 2+ years engineering management leading 4+ engineers, strong cloud/DevOps/SRE background, experience with AWS, Kubernetes, Terraform, hiring and stakeholder management.
AWS, Kubernetes, Terraform, CI/CD
2w
Save
Mark Applied
Hide
Manager of Platform DevOps
San Jose or Seattle
$124k-$271k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
5+ YOE3+ Mgmt5+ years SRE/DevOps experience,3+ years managing technical teams,BS/MS or equivalent,expertise with Kubernetes,Terraform,cloud providers,CI/CD and incident response,programming and scripting skills.
Kubernetes, Terraform, AWS, OCI, Git, Jenkins, Argo CD, JFrog artifactory, ELK stack, Prometheus, Grafana, BrightHire
1w
Save
Mark Applied
Hide
Principal Cloud Engineer
San Jose, California, United States
$156k-$224k/yr OnsiteFull Time
Bloom Energy
Bloom EnergyNYSE: BE: Manufactures solid oxide fuel cell systems for onsite power.
10+ YOE3+ MgmtBachelor’s degree and 10+ years in cloud, infrastructure, platform engineering, or DevOps, including 3+ years in senior technical leadership. Requires AWS, networking, IaC, CI/CD, Kubernetes, observability, security, and SRE expertise.
AWS, Terraform, CloudFormation, Prometheus, Grafana, Datadog, ELK, CloudWatch, Jenkins, GitLab, GitHub Actions, ArgoCD, Kubernetes, Python, Bash, PowerShell, Azure, Google Cloud, Oracle Cloud, VPC, Transit Gateway, Cloud WAN, Direct Connect, BGP, IAM, SRE, GitOps, CI/CD, AI/ML, GPU
4d
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
2w
Save
Mark Applied
Hide
Head of Infrastructure and Employee Services
San Jose, California, United States
$195k-$426k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
15+ YOE5+ Mgmt15+ years leading IT infrastructure and operations with 5+ years in senior management; experience managing large global teams, IAM, SRE/automation, AI/ML in support environments, XaaS and on‑prem architectures, and compliance frameworks (NIST, SOX, SOC).
Okta, SailPoint, macOS, Windows, Linux, Workspace One, Jamf, InTune, VDI, Qualys, Tenable, Zoom, Google Workspace, Proofpoint, SendGrid, ServiceNow, Jira, PagerDuty, CrowdStrike, Digital Guardian, AWS, GCP, BrightHire
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - System Service Global
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in related field and strong experience with large-scale Linux host management, core data-center services (DNS, NTP, DHCP, NAT, APT, Kerberos), DevOps tooling, SRE practices, and troubleshooting.
BIND, PowerDNS, NTP, DHCP, NAT, APT, Kerberos, Ansible, Salt, Puppet, CI/CD, Python, Go, Bash, Linux
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Managed Kubernetes
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
6+ YOERequires 6+ years in SRE or operations, deep Linux and production Kubernetes expertise, strong Go and Python skills, GitOps, Helm, observability, CI/CD, and Kubernetes provisioning experience.
Kubernetes, Python, Golang, GitOps, ArgoCD, Helm, Linux, EKS, GKE, Prometheus, Grafana, FluentBit, CI/CD, kubeadm, Cluster API, CRDs, CSI, CNI
1mo
Save
Mark Applied
Hide
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOE5+ years SRE/DevOps/Linux operations experience, bachelor’s in CS or related field, proficiency in Go/Python/C++, strong troubleshooting, monitoring, and reliability practices.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG, AMI
1mo
Save
Mark Applied
Hide
Cloud & Engineering Solution Capability Leader
San Francisco or Oakland or San Jose
$240k-$290k/yr OnsiteFull Time
Slalom
Slalom: Provides business and technology consulting and software engineering services.
Senior leader with deep technology delivery and leadership experience in product engineering, cloud modernization, AI-accelerated engineering, SRE/operations, and go-to-market capability development.
AWS, Microsoft, Google
3mo
Save
Mark Applied
Hide
Staff Software Engineer - Managed Kubernetes
Bellevue or San Francisco or San Jose
$314k-$465k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
10+ YOE10+ years in software/platform engineering or SRE; 5+ years Kubernetes at scale; strong Go and Python; deep Kubernetes internals; GPU orchestration; multi-tenant infrastructure; distributed systems; observability at scale; Linux networking; IaC and GitOps.
Go, Python, Kubernetes, NVIDIA GPU Operator, NCCL, DCGM, GPUDirect, CNI, InfiniBand, RDMA, GaP? note: ignore invalid, GitOps
1d
Save
Mark Applied
Hide
Platform Services Technical Leader IRC302309
San Jose, California, United States
$130k-$140k/yr RemoteFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
8+ YOERequires 8–12+ years in SRE, platform engineering, infrastructure, distributed systems, or cloud operations; technical leadership, software engineering, cloud systems, CI/CD, Kubernetes, AWS, security, and incident response expertise.
Ansible, Argo CD, Amazon Web Services (AWS), Helm, Terraform, Python, Go, Ruby, Kubernetes
1mo
Save
Mark Applied
Hide
Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$245k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's in CS or related, strong CS foundation, Linux and storage/network knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; 5+ years SRE/cloud experience preferred.
Linux, Python, Go, Java, PHP, C, C++
2mo
Save
Mark Applied
Hide
Director, Enterprise IT Infrastructure (CA, US, 95110)
San Jose, California, United States
$191k-$280k/yr OnsiteFull Time
QuantumScape
QuantumScapeNYSE: QS: Develops next-generation solid-state batteries for electric vehicles.
15+ YOE5+ Mgmt15+ years IT experience with 5+ years leading enterprise infrastructure; Bachelor's in CS/IT/Engineering; hands-on expertise with GCP/Azure, Kubernetes (GKE), networking, identity, endpoints, SRE, automation, and OT/IT integration; strong leadership and cross-functional communication.
GCP, Azure, Compute Engine, GKE, Cloud Storage, IAM, VPC, LAN/WAN, Wi-Fi, SD-WAN, Palo Alto firewalls, Kubernetes, SSO, MFA, MDM, EDR, ThreatLocker, Google Workspace, Microsoft 365, JIRA, ManageEngine, Cisco, OT/ICS, NIST CSF, ISO 27001, SOC 2, AIOps

Explore Jobs

Expand Your Job Search