36 cloud reliability engineer jobs at 25 companies in Londonderry, NH

3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Hybrid)
Merrimack or San Diego
$118k-$201k/yr HybridFull Time
BAE Systems
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Juniper, Ansible, Helm Charts, NFS, JDFS, Ceph, S3, VMware, Open Stack, Azure Stack, Kubernetes, Terraform
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer- Eng
Lowell, Massachusetts, United States
$130k-$186k/yr HybridFull Time
UKG
UKG: Workforce management and human capital management software provider.
5+ YOE5+ years software/systems/cloud engineering; public cloud experience (GCP/AWS/Azure); observability, SLOs, incident response, Linux, coding in Python/Java/C++; GitHub Actions and dashboarding (Splunk/Grafana).
GCP, AWS, Azure, Python, Java, C++, Linux, GitHub Actions, Splunk, Grafana, Kubernetes, Terraform, Ansible
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer I
Boston or Seattle or Atlanta
$134k-$215k/yr HybridFull Time
Axon
AxonNASDAQ: AXON: Develops public safety technologies, devices, and cloud software.
7+ YOEBachelor's in CS/Engineering, 7+ years software engineering experience, expertise in distributed systems, Kubernetes, cloud (Azure/AWS/GCP), observability, Kafka, Terraform/Pulumi, and experience with agentic AI/LLM tooling preferred.
Kubernetes, Terraform, Pulumi, Kafka, Grafana, Datadog, New Relic, MySQL, Cassandra, PostgreSQL, Azure, AWS, GCP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Newton, Massachusetts, United States
$160k-$205k/yr OnsiteFull Time
Manifold
Manifold: The Enterprise Agent Platform for life sciences that helps biopharma and research teams analyze governed biomedical data.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
AWS, GCP, Azure, Terraform, Github Action, Okta, Auth0, Docker, ECS, Packer, Tailscale, WireGuard, Snowflake, Airflow, dbt, PostgreSQL, LLM, CI/CD
1w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Boston, Massachusetts, United States
$148k-$185k/yr OnsiteFull Time
DraftKings
DraftKingsNasdaq: DKNG: Publicly traded U.S. sports entertainment and gaming serving fans with fantasy sports, betting, lottery, and events.
7+ YOEBachelor's degree or equivalent experience; 7+ years in site reliability engineering with SLOs, SLIs, and error budgets; Datadog expertise; distributed systems, cloud infrastructure, communication, and cross-team influence skills.
Datadog, Amazon Web Services, Kubernetes
3mo
Save
Mark Applied
Hide
Sr. Manager, Site Reliability Engineer (SRE)
Boston or Hampton
$160k-$185k/yr HybridFull Time
Planet Fitness
Planet FitnessNYSE: PLNT: Leading fitness center franchisor and operator.
7+ YOE7+ years leading SRE/DevOps with cloud (AWS/Azure/GCP), incident management, SLO/SLI experience, CI/CD and observability expertise; bachelor\u0002s degree or equivalent experience.
AWS, Azure, GCP, CI/CD, Infrastructure as Code (IaC)
2w
Save
Mark Applied
Hide
Customer Reliability Engineer - Airflow
United States or New York City or Boston or San Francisco
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
4+ YOERequires data engineering background, 4 years with Python, 1 year administering Airflow and creating DAGs, Kubernetes, Docker, containers, cloud provider experience, troubleshooting, communication, autonomy, and mentoring experience.
Apache Airflow, Python, Kubernetes, Docker, AWS, GCP, Azure, Zoom, SQL, PostgreSQL, Databricks, Snowflake, Redshift, dbt
2w
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNYSE: IBM: Global technology and consulting focusing on cloud and AI.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
London or Singapore or Tokyo or Houston or Boston
HybridFull Time
Veson Nautical
Veson Nautical: Private maritime software serving shipowners, charterers, traders, and operators with commercial freight management solutions.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Waltham, Massachusetts, United States
$130k-$140k/yr HybridFull Time
SS&C Technologies
SS&C TechnologiesNASDAQ: SSNC: Global provider of financial services and healthcare technology solutions.
7+ YOEExperience in Unix/Linux, Java, Python or Bash, Kubernetes, Docker, cloud (AWS/Azure), IaC (Terraform/Ansible), monitoring tools, networking protocols, and on-call incident response.
Unix, Linux, Java, Python, Bash, Kubernetes, Docker, Azure, AWS, CloudWatch, EKS, EFS, S3, RedShift, Terraform, Ansible, HTTP(s), JMS, TCP, UDP, Splunk, Datadog, Dynatrace, Zabbix, Prometheus, Akamai, Cloudflare, DNS, CDN, DataStream, WAF, Oracle, PostgreSQL, MongoDB, RabbitMQ, Interconnect, AMQ
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (Senior or Staff)
Boston or Miami or New Jersey or New York City or Princeton or Raleigh or Washington or Toronto or North America
$127k-$249k/yr HybridFull Time
MongoDB
MongoDBNASDAQ: MDB: Unified data platform for building modern applications.
6+ YOE6+ years software development experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud (AWS, Google Cloud Platform, Microsoft Azure) expertise; Linux and networking knowledge.
Argo Workflows, ArgoCD, Kubernetes, Python, Go, AWS, Google Cloud Platform (GCP), Microsoft Azure, Linux
2mo
Save
Mark Applied
Hide
Senior DevOps Site Reliability Engineer
North Andover, Massachusetts, United States
HybridFull Time
TSD Mobility Solutions
TSD Mobility Solutions: Provider of automotive dealership software and business solutions.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Jenkins, AWS, EC2, S3, RDS, Lambda, VPC, IAM, Windows IIS, Linux, Azure DevOps, Bash, Python, PowerShell, Terraform, Ansible, Web Deploy (MSDeploy), PowerShell DSC, Docker, Kubernetes, Datadog, Grafana, CloudWatch
3mo
Save
Mark Applied
Hide
Site Reliability Engineer - Disaster Recovery & Business Continuity
Boston or Chicago
$130k-$150k/yr HybridFull Time
Charles River Associates
Charles River AssociatesNasdaq Global Select Market: CRAI: Public global economic, financial, and management consulting firm serving law firms, corporations, accounting firms, and governments.
Experience with IT service continuity, disaster recovery, cross-functional coordination, and documentation.
Windows, Microsoft 365, Cloud, SaaS, Backups, Virtualization, Identity, Networking
1d
Save
Mark Applied
Hide
SRE - Enterprise & Cloud Security - AI Driven Security - Manager
New York City or Atlanta or Chicago or Washington or Boston or Dallas or San Francisco or Seattle or Houston
$99k-$232k/yr OnsiteFull Time
PwC
PwC: Global professional services network providing audit, tax, and consulting services.
5+ YOEBachelor's degree and 5+ years of experience required. Preferred fields include data science, AI, computer science, information systems, or engineering; cloud, data engineering, or machine learning credentials preferred.
Python, C++, AWS, Google Cloud, Microsoft Azure, Databricks, Snowflake
2w
Save
Mark Applied
Hide
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yr HybridFull Time
Zelis
Zelis: Private healthcare financial technology serving payers, providers, and healthcare consumers with claims, payment, and member-engagement solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEBS, MS, or PhD in computer science or electrical engineering, or equivalent experience; 6+ years in systems programming; expertise in distributed systems, C++ and Python, HPC or cloud RCA pipelines, CPU metrics, and cluster managers.
C++, Python, Slurm, Kubernetes, LSF, Linux kernel, /dev/mcelog, dmesg, journald, NVIDIA DCGM (Data Center GPU Manager), NVIDIA Management Library (NVML), CRIU
1mo
Save
Mark Applied
Hide
Member of Technical Staff – Senior Engineer, Data Infrastructure & Data Operations
San Francisco or Cambridge
$255k-$340k/yr OnsiteFull Time
Walden Robotics
Walden Robotics: Private full-stack physical AI building and deploying general-purpose robots for manufacturing and logistics.
Experience building production data infrastructure and high-throughput pipelines, cloud-based data platform development, platform reliability and cost ownership, and collaboration with ML teams.
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
C++, Python, Slurm, LSF, Kubernetes, CUDA, DCGM, NVML, CRIU, /dev/mcelog, dmesg, journald, Linux kernel
1w
Save
Mark Applied
Hide
DevOps Engineer
Epalinges or Lausanne or Menlo Park or Boston or Europe
HybridFull Time
Atinary Technologies
Atinary Technologies: Private AI deeptech providing machine-learning and robotics software for scientific research and development.
3+ YOERequires 3+ years in DevOps, site reliability, or cloud infrastructure, or 2+ software engineering and 1+ DevOps years; AWS, CI/CD, containers, Python, Bash, and Infrastructure as Code experience required.
Python, AWS, GitHub Actions, Bash, Terraform, OpenTofu, Infrastructure as Code (IaC)