66 site reliability devops platform engineer jobs at 54 companies in United States

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
2w
Save
Mark Applied
Hide
DevOps & Site Reliability Lead-Retail Devops
Deerfield, Illinois, United States
$120k-$160k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
10+ YOEExpertise in Java full‑stack and Spring, REST, Oracle/DB2/MySQL, Azure platform (AKS, containers, networking), Azure DevOps CI/CD, PowerShell/Bash scripting, Dynatrace, SRE leadership and reliability engineering.
Java, Spring Boot, Spring MVC, Spring Security, REST, Oracle, DB2, MySQL, BASE24 EPS, C++, AS400, Python, Microsoft Azure, Azure Container Apps, AKS, Docker, VNETs, Private Endpoints, Application Gateway, Azure Key Vault, Azure Monitor, Log Analytics, Azure DevOps, PowerShell, Bash, GitOps, Dynatrace
2w
Save
Mark Applied
Hide
Staff Site Reliability Engineer- Developer Platform
Palo Alto, California, United States
$186k-$233k/yr OnsiteFull Time
Rivian and Volkswagen Group Technologies
Rivian and Volkswagen Group Technologies: A joint venture creating cloud, connectivity and software-defined vehicle solutions for electric vehicles.
5+ YOE5+ years in Platform/DevOps/SRE; Terraform, Kubernetes, GitOps (ArgoCD or Flux), cloud (AWS/Azure/GCP), scripting (Python, Bash, Go); strong communication and mentoring skills.
Terraform, Kubernetes, ArgoCD, Flux, Python, Bash, GoLang, AWS, Azure, GCP
1mo
Save
Mark Applied
Hide
Senior DevOps Engineer/Site Reliability Engineer-East Coast
United States
$165k-$215k/yr RemoteFull Time
Stellar Cyber
Stellar Cyber: Unified platform for automated cyber threat detection and response.
5+ YOE5+ years in DevOps/SRE/Platform Engineering; Kubernetes, Docker, cloud (OCI/AWS/GCP/Azure), Terraform/Helm, CI/CD, observability (Prometheus/Grafana/Loki/Alertmanager/Elastic Stack), Python/Bash/Go, Linux, on-call experience; East Coast residency.
Kubernetes, Docker, OCI, AWS, GCP, Azure, Terraform, Helm, Prometheus, Grafana, Loki, Alertmanager, Elastic Stack, Elasticsearch, Python, Bash, Go, Kafka, Spark, Redis, MongoDB, Linux, ArgoCD, GitHub Actions
2w
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS/Engineering, strong Linux, networking, databases, Kubernetes, SRE/DevOps toolset knowledge, experience with ClickHouse/Spark/Presto/Doris/Hadoop, coding in Python/Shell/Java/Go, strong problem-solving and communication.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Python, Shell, Java, Go
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
Dimensional Fund Advisors: Provides systematic investment solutions based on financial science.
5+ YOE5+ years in SRE/DevOps/platform engineering with Python and .NET tooling experience.
Python, .NET, Poetry, Virtual Environments, IDE Tools, PowerShell, Bash, VS Code, JetBrains IDEs, Visual Studio
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Kubernetes Platform (Starshield)
Hawthorne or Redmond
$125k-$175k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
1+ YOEBachelor's in CS/IT/engineering + 1+ year SRE/DevOps experience (or 3+ years experience), Linux, Terraform/Ansible, Kubernetes and OCI containers, Bash/Python scripting, development in Python/C++/Go, willingness to obtain Top Secret clearance.
Terraform, Ansible, OCI containers, Kubernetes, Bash, Python, C++, Go, Bazel, Makefiles, TCP/IP, Linux
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
United States
$165k-$230k/yr RemoteFull Time
SimSpace
SimSpace: Provides high-fidelity cyber ranges for cybersecurity training and testing.
8+ YOE8+ years in SRE/Platform/DevOps with expertise in distributed systems, Kubernetes, observability, security architecture, CI/CD, and strong software engineering skills.
Jsonnet, Grafana Tanka, Kustomize, GitHub Actions, ArgoCD, Grafana, Go, Python, Kubernetes
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Houston, Texas, United States
HybridFull Time
NOV
NOVNYSE: NOV: Manufactures and provides equipment for the energy industry.
5+ YOE5+ years SRE/DevOps experience; expertise in Kubernetes, AKKA.NET, cloud platforms (AWS/Azure/GCP), scripting (Bash/PowerShell/Python), observability stacks, and PostgreSQL tuning; proven incident management skills.
Prometheus, Grafana, Datadog, OpenTelemetry, ELK, Phobos, AKKA.NET, PostgreSQL, GitHub Actions, Azure Pipelines, GitLab CI, Bash, PowerShell, Python, AWS, Azure, GCP, Kubernetes, Docker, Terraform, GitHub, GitLab, Azure DevOps, C#
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or Canada
$180k-$250k/yr RemoteFull Time
Orkes
Orkes: Cloud platform for building distributed, event-driven applications.
5+ YOE5+ years in SRE/DevOps/Platform Engineering; cloud platforms (AWS/GCP/Azure); Kubernetes; observability; automation; CI/CD; strong incident management; cross-functional collaboration.
Kubernetes, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Terraform, Python, Bash, CI/CD, Git, Linux
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
United States
$150k-$200k/yr RemoteFull Time
Novellia
Novellia: A platform for patient-authorized medical records and biopharma research insights.
5+ YOE5+ years in SRE/platform/DevOps with production ownership, comfortable coding in Python/Go/TypeScript, incident management, SLOs, observability, and strong collaboration and process skills.
Python, Go, TypeScript
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Greenville, South Carolina, United States
$100k-$120k/yr RemoteFull Time
CURO Financial Technologies
CURO Financial TechnologiesNYSE: CURO: Provider of consumer credit products and financial services.
5+ YOE5+ years SRE/platform/DevOps experience with production ownership; deep hands-on Kubernetes, AWS, Terraform, Python, CI/CD, observability, incident response, and compliance awareness.
Kubernetes, ArgoCD, Helm, Terraform, Python, AWS, EKS, ECS, EC2, ECR, IAM/IRSA, VPC, ALB, NLB, CloudWatch, Secrets Manager, KMS, GitHub Actions, Azure DevOps, Grafana, Prometheus, Mimir, Loki, Tempo, OpenTelemetry, Bash, Argo Rollouts, Sigstore, cosign, SBOMs, SLSA, .NET, C#, Cilium, eBPF, Karpenter
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
United States
$160k-$200k/yr RemoteFull Time
Epic
EpicNYSE: TAL: Subscription-based digital reading and learning platform for children.
5+ YOE5+ years in infrastructure/platform/DevOps, experience with GCP, Kubernetes, CI/CD, observability, Terraform, scripting (Python/Bash), and SRE practices including SLO/SLI definition.
Google Cloud Platform (GCP), GCE, GCS, VPC, IAM, Cloud Monitoring, Docker, Kubernetes, GKE, Helm, GitHub Actions, ArgoCD, Jenkins, New Relic, Terraform, Python, Bash, Dagster, Airflow, PromRelay
1mo
Save
Mark Applied
Hide
GOV Site Reliability Engineer
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yr RemoteFull Time
Veeam
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.
VDC, Prometheus, Grafana, OpenTelemetry, ELK stack, Terraform, Terragrunt, Pulumi, Kubernetes, GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, TypeScript, JS, Go, Java, C#, Azure Government, AWS GovCloud
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Foundation Team
Dublin or United States or Europe or Asia
OnsiteFull Time
Global-e
Global-eNasdaq: GLBE: Platform enabling global direct-to-consumer e-commerce for brands.
2+ YOEMinimum 2 years in DevOps/SRE or platform engineering with hands-on AWS, Terraform, Kubernetes, CI/CD, observability, and scripting (Python/Bash); familiarity with SRE principles.
Terraform, AWS, EKS, RDS, Kinesis, S3, IAM, Kubernetes, GitHub Actions, ArgoCD, Jenkins, Datadog, Prometheus, Grafana, Python, Bash, AIOps
4d
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Berkeley, Missouri, United States
$198k-$268k/yr OnsiteFull Time
Boeing
BoeingNYSE: BA: Designing and manufacturing commercial aircraft and defense systems.
14+ YOEActive U.S. Secret clearance, 14+ years SRE/DevOps experience, Bachelor's degree, leadership in GitLab/Jira/PostgreSQL platforms, cloud and automation expertise, ability to obtain Security+.
GitLab, GitLab CI/CD, Jira, Confluence, PostgreSQL, Artifactory, SonarQube, Azure DevOps, AWS, Microsoft Azure, Infrastructure as Code, Ansible, Docker, Kubernetes, Jenkins
1mo
Save
Mark Applied
Hide
Platform Operations Engineer (Site Reliability Engineer)
Westerville or Columbus
OnsiteFull Time
Vertiv
VertivNYSE: VRT: Provides critical power and cooling for data centers.
5+ YOEBachelor's in CS/IS/Engineering or equivalent, 5+ years in platform ops/SRE/DevOps, 3+ years with enterprise observability tools, incident/SLA/SLO ownership, cloud (AWS), containers, IaC, and programming for automation.
Compass AI, Writer AI, Site Scope, UiPath, Workato, Cursor, Power Automate, Datadog, Grafana, Prometheus, Azure Monitor, Splunk, AWS, Docker, Kubernetes, Terraform, Ansible, Python, Ruby, Powershell, Java, Javascript, C#, GitLab, Jenkins, GitHub Actions, Azure DevOps, SAST/DAST
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Denver or Dallas
HybridFull Time
Analytic Partners
Analytic Partners: Provides commercial analytics software and marketing measurement solutions.
4+ YOE4+ years in Platform Engineering/DevOps or related; strong Linux/Windows; automation with Python, Bash, or PowerShell; deep AWS and Azure experience; CI/CD; Infrastructure as Code; containers.
Linux, Windows, Python, Bash, PowerShell, AWS, Azure, Jenkins, GitHub Actions, Terraform, CloudFormation, Arm, Docker, Kubernetes, Nomad, Consul, Vault, Splunk, Sumo Logic
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
United States
$180k-$200k/yr RemoteFull Time
PayNearMe
PayNearMe: A payments technology delivering secure, scalable payments solutions for businesses and consumers.
3+ YOE3+ years in SRE/DevOps; cloud platforms (AWS/GCP/Azure); Kubernetes & Docker; Terraform; monitoring (Datadog); CI/CD with GitLab; scripting in Python/Bash/Go; production support; strong collaboration.
Terraform, Kubernetes, Docker, Datadog, Prometheus, Grafana, ELK, Splunk, Python, Bash, Go, GitLab CI, Ansible, Puppet, Chef
1w
Save
Mark Applied
Hide
Site Reliability Engineer (Senior+)
United States
$103k-$287k/yr RemoteFull Time
Miris
Miris: Provides infrastructure for streaming high-fidelity 3D spatial content.
7+ YOE7+ years SRE/DevOps experience, expertise in cloud platforms (AWS Fargate, CoreWeave), Terraform, Kubernetes, observability (Prometheus, Grafana), SLI/SLO and incident response, security and compliance (SOC 2, GDPR, ISO 27001).
Terraform, AWS Fargate, CoreWeave, Kubernetes, Prometheus, Grafana
1mo
Save
Mark Applied
Hide
Site Reliability Engineer II
McLean, Virginia, United States
$104k-$150k/yr HybridFull Time
Medallia
Medallia: Provides enterprise experience management and customer feedback software.
2+ YOE2+ years SRE/DevOps experience with Kubernetes, cloud platforms (AWS/OCI/GCP), Linux administration, scripting (Python/Bash/Go), CI/CD/Git workflows, networking fundamentals, and ability to participate in on-call rotation.
Kubernetes, AWS, OCI, GCP, Linux, Python, Bash, Go, Git, GitOps, ArgoCD, Terraform, Prometheus, Grafana, Loki, OpenTelemetry, CI/CD