39 principal site reliability engineer jobs at 26 companies in United States

2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Nashville, Tennessee, United States
$85k-$210k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
3+ YOEExperience in site reliability, Windows/Linux administration, scripting/automation, cloud infrastructure, patching and incident response; 3+ years relevant experience; strong communication and documentation skills.
PowerShell, Bash, Python, Ansible, Chef, Oracle Cloud Infrastructure, Citrix, Security Technical Implementation Guides (STIG)
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Orlando or Glendale
$176k-$235k/yr HybridFull Time
The Walt Disney Company
The Walt Disney CompanyNYSE: DIS: Produces media content and operates global theme parks.
10+ YOE10+ years experience; expertise in observability, multi-cloud (AWS/Azure/GCP), CI/CD, Terraform/Cloud Formation/Ansible/Chef, containers/Kubernetes, Git, and leadership/mentoring skills.
Gitlab, AWS CodeBuild, CodeDeploy, CodePipeline, Azure DevOps, Terraform, Cloud Formation, Ansible, Chef, Harness, Kubernetes, AWS, Azure, GCP, Datadog, New Relic, Dynatrace, Git
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Orlando or Glendale
$176k-$247k/yr HybridFull Time
The Walt Disney Company
The Walt Disney CompanyNYSE: DIS: Produces movies, operates theme parks, and provides streaming services.
10+ YOE10+ years experience; expertise in multi-cloud (AWS, Azure, GCP), observability, CI/CD, infrastructure as code (Terraform/CloudFormation), Linux systems administration; bachelor's degree or equivalent; strong leadership and communication.
AWS, Azure, GCP, Git, Gitlab, AWS CodeBuild, CodeDeploy, CodePipeline, Azure DevOps, Terraform, Cloud Formation, Ansible, Chef, Harness, Kubernetes, Datadog, New Relic, Dynatrace
1mo
Save
Mark Applied
Hide
Consulting/Principal Site Reliability Engineer
Mumbai or United States
OnsiteFull Time
LexisNexis Risk Solutions
LexisNexis Risk SolutionsNYSE: RELX: Provides data and analytics for risk management and compliance.
Design, build, and operate reliable scalable cloud infrastructure and observability; proficiency with Python, PowerShell, Shell, AWS, Microsoft Azure, Terraform, Docker, Kubernetes, Git/GitHub, Prometheus, Grafana; strong systems administration and incident response skills.
Python, PowerShell, Shell, AWS, Microsoft Azure, Terraform, GitHub, Git, Docker, Kubernetes, Prometheus, Grafana, CI/CD, Windows, Linux, Unix
1mo
Save
Mark Applied
Hide
Principal Software Engineer, Site Reliability
Bellevue, Washington, United States
$218k-$250k/yr OnsiteFull Time
UiPath
UiPathNYSE: PATH: Provides robotic process automation software for enterprise workflow automation.
10+ YOE10+ years building large-scale distributed systems; proficiency in OOP (C#, C++, Java, Python), cloud (Azure/AWS/GCP), Kubernetes, databases, CI/CD, SRE and incident management; strong mentoring and product-focused engineering experience.
C#, C++, Java, Python, Kubernetes, Azure, AWS, GCP, AKS, GKE, Azure SQL, CosmosDB, Azure Data Lake, Power BI, MongoDB, MySQL, DynamoDB, CI/CD, DevOps
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Durham, North Carolina, United States
OnsiteFull Time
Fidelity Investments
Fidelity Investments: Provides investment management, retirement planning, and brokerage services.
3+ YOEBachelor's in CS/IT/Engineering + 5 years SRE experience, or Master's + 3 years; experience with CI/CD, cloud (AWS/Azure), observability (Datadog, Splunk), performance testing, Python/Shell scripting.
Datadog, Splunk, Grafana, Java, JMeter, Cloud-test, Rush-hour, Python, Shell, Kubernetes, uDeploy, Jenkins Core, Ansible AWX, Terraform, Azure, AWS, AWS Route53, Azure Load Balancer, F5, AVI
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - Remote
Minnetonka or Minneapolis or Washington or United States
RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
10+ YOE10+ years SRE/Software/Cloud engineering experience, Bachelor’s or equivalent, strong reliability and automation background, Azure and observability experience, ability to influence cross-functional teams.
Azure, OpenTelemetry, AIOps, AI/ML
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - Austin, Texas
Austin, Texas, United States
HybridFull Time
ShipperHQ
ShipperHQ: Provides shipping rate management and checkout optimization software for e-commerce.
10+ YOE10+ years SRE/platform/devops experience; proven AWS cloud experience; expert Terraform and GitLab CI/CD; Kubernetes and container expertise; strong software engineering, observability, SLO/SLI, networking, Linux, and mentoring skills.
AWS, Terraform, GitLab, Kubernetes, Linux
1mo
Save
Mark Applied
Hide
Principal Reliability Engineer
Burlington or Indianapolis or Durham
OnsiteFull Time
Labcorp
LabcorpNYSE: LH: Provides clinical laboratory testing and drug development services globally.
10+ YOE10+ years progressive experience in reliability/project/facilities engineering, 5+ years leading complex multi‑site projects, bachelor’s in engineering (or equivalent experience), 5+ years with CMMS and reliability analytics, up to 25% travel.
CMMS, predictive maintenance technologies, reliability analytics
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Colorado or United States or North America or Europe
$117k-$180k/yr RemoteFull Time
Zayo Group
Zayo Group: Provides high-capacity fiber networks and communications infrastructure.
10+ YOE10+ years SRE experience, strong Linux and system administration, Python and shell scripting, Kubernetes and Docker, monitoring/alerting and automation expertise, cloud experience (AWS/Google), demonstrated leadership.
Python, Kubernetes, Docker, SevOne, Assure1, Nagios, Prometheus, Grafana, Cacti, AWS, Google, Ansible, Terrafor, Puppet, Ciena Blue Planet MDSO, Cisco NSO, Nokia NSP, netconf
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - CTJ - Secret
Redmond, Washington, United States
$143k-$275k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
2+ YOEDegree in CS/IT (or equivalent experience) with minimum 2+ years technical experience (Doctorate path) and ability to obtain required background investigations (T3/CJIS) for government cloud environments; SRE, incident response, and cloud systems experience.
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - ARINCDirect (Remote)
Arlington or United States
$108k-$205k/yr RemoteFull Time
RTX
RTXNYSE: RTX: RTX provides advanced aerospace and defense systems and services.
8+ YOESTEM degree with 8+ years relevant experience (or advanced degree with 5+ years, or 12+ years without degree). Must be authorized to work in the U.S. without sponsorship. Experience in Linux, Docker, Kubernetes, infrastructure automation (Saltstack, Ansible, Terraform), hardware (servers, switches, cabling), monitoring, incident response, and capac...
Docker, Kubernetes, Saltstack, Ansible, Terraform, GitOps, Python, Linux Shell (bash, awk, sed), Linux, PostgreSQL, SQL, AWS, DNS, DHCP, LDAP, NFS
1w
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Machine Learning
Cambridge, Massachusetts, United States
$142k-$178k/yr OnsiteFull Time
Cambridge Mobile Telematics
Cambridge Mobile Telematics: Providing telematics and behavioral analytics for safer driving and insurance.
7+ YOEBachelor's or equivalent, 7+ years SRE/IT experience, AWS (EC2,EKS, S3,RDS), Databricks, Ray, Terraform, Python, Linux, Datadog/CloudWatch, strong incident response and system design skills.
Ray, AWS EKS, Databricks, CloudWatch, Datadog, EC2, S3, RDS/Aurora, Dynamo, SQS, Lambda, IAM, Terraform, Python, Docker, Kubernetes, Unity Catalog, CI/CD
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer - Volcano
United States
$150k-$210k/yr RemoteFull Time
Kong
Kong: Provides cloud-native API management and service mesh platforms.
B.S. in Computer Science or equivalent; substantial Staff/Principal-level SRE/platform engineering experience; deep Kubernetes expertise; experience defining SLOs, incident response, and operating multi-tenant data services; familiarity with ArgoCD, Helm, Terraform/Terragrunt, Datadog, Prometheus, and Grafana.
Kubernetes, ArgoCD, Helm, Terraform, Terragrunt, PostgreSQL, Redis, Datadog, Prometheus, Grafana, GitOps, CI/CD, CNI
2mo
Save
Mark Applied
Hide
Principal Architect, Site Reliability Engineering
Southlake or Austin
$221k-$252k/yr OnsiteFull Time
Charles Schwab
Charles SchwabNYSE: SCHW: Financial services, brokerage, and investment management provider.
5+ YOE3+ Mgmt5+ years in SRE with 3+ years in architect/leadership; design scalable, fault-tolerant systems; strong observability; CI/CD; postmortems; SRE leadership.
Prometheus, Grafana, Datadog, Splunk
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Chicago, Illinois, United States
$132k-$220k/yr HybridFull Time
CME Group
CME GroupNASDAQ: CME: Operates global derivatives marketplaces for trading futures and options.
10+ YOE3+ Mgmt10+ years in SRE/Systems/Software engineering; 3+ years in a staff/principal/tech lead role; experience leading cloud migrations or re-architecting platforms; GCP/Kubernetes; Python; Go; financial/regulatory environment familiarity.
Python, Go, Kafka, Kubernetes, GCP, Terraform, ArgoCD, Node.js
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - CTJ - Secret (200038967)
United States
OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Global provider of software, cloud, and AI technology solutions.
No qualifications specified in the posting; title indicates a position requiring Secret security clearance.