45 site reliability engineer jobs at 27 companies in Westford, MA

2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Cambridge or United States
$76k-$136k/yr HybridFull Time
Akamai Technologies
Akamai TechnologiesNASDAQ: AKAM: Cloud and edge computing platform for secure digital experiences.
Bachelor's in Computer Science/Engineering or equivalent; experience in SRE/Software Engineering for large-scale distributed systems; Terraform and IAC experience; familiarity with SaltStack/Ansible/Chef/Puppet; Linux, CI/CD, observability, and participation in on-call rotation.
Terraform, SaltStack, Ansible, Chef, Puppet, Linux, CI/CD, Infrastructure as Code (IAC)
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Hybrid)
Merrimack or San Diego
$118k-$201k/yr HybridFull Time
BAE Systems
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Juniper, Ansible, Helm Charts, NFS, JDFS, Ceph, S3, VMware, Open Stack, Azure Stack, Kubernetes, Terraform
1w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Boston, Massachusetts, United States
$148k-$185k/yr OnsiteFull Time
DraftKings
DraftKingsNasdaq: DKNG: Publicly traded U.S. sports entertainment and gaming serving fans with fantasy sports, betting, lottery, and events.
7+ YOEBachelor's degree or equivalent experience; 7+ years in site reliability engineering with SLOs, SLIs, and error budgets; Datadog expertise; distributed systems, cloud infrastructure, communication, and cross-team influence skills.
Datadog, Amazon Web Services, Kubernetes
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer I
Boston or Seattle or Atlanta
$134k-$215k/yr HybridFull Time
Axon
AxonNASDAQ: AXON: Develops public safety technologies, devices, and cloud software.
7+ YOEBachelor's in CS/Engineering, 7+ years software engineering experience, expertise in distributed systems, Kubernetes, cloud (Azure/AWS/GCP), observability, Kafka, Terraform/Pulumi, and experience with agentic AI/LLM tooling preferred.
Kubernetes, Terraform, Pulumi, Kafka, Grafana, Datadog, New Relic, MySQL, Cassandra, PostgreSQL, Azure, AWS, GCP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Newton, Massachusetts, United States
$160k-$205k/yr OnsiteFull Time
Manifold
Manifold: The Enterprise Agent Platform for life sciences that helps biopharma and research teams analyze governed biomedical data.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
AWS, GCP, Azure, Terraform, Github Action, Okta, Auth0, Docker, ECS, Packer, Tailscale, WireGuard, Snowflake, Airflow, dbt, PostgreSQL, LLM, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Somerville, Massachusetts, United States
$160k-$200k/yr HybridFull Time
Tulip Interfaces
Tulip Interfaces: Private software providing AI-enabled, no-code frontline operations tools for manufacturers.
5+ YOE5+ years experience with observability tools, OpenTelemetry instrumentation, Prometheus metrics, experience with time-series data and producing SLIs/SLOs, strong systems reasoning and communication.
Grafana, Loki, Tempo, Mimir, OpenTelemetry, Prometheus, promQL, TypeScript, Go, Kubernetes, MongoDB, PostGres, Alloy, Claude Skills, Gemini Gems
2d
Save
Mark Applied
Hide
Site Reliability Engineer
Waltham, Massachusetts, United States
$166k-$220k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology developing AI-powered autonomous military systems.
3+ YOERequires 3+ years in SRE, DevOps, field/systems engineering, or production support; Linux and networking expertise; communication skills; on-call availability; and eligibility for U.S. Secret clearance.
Linux, IP, VPN, Python, Bash, Nix, NixOS, systemd, PagerDuty
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Woonsocket, Rhode Island, United States
$79k-$159k/yr HybridFull Time
CVS Health
CVS HealthNYSE: CVS: Diversified healthcare integrating retail, pharmacy, and insurance services.
Experience in SRE/platform engineering, edge/distributed architecture, Java/Python/Go, cloud (AWS/Azure/GCP), monitoring/observability tools, CI/CD and infrastructure as code; bachelor’s degree or equivalent experience.
Java, Python, Go, AWS, Azure, GCP, Prometheus, Grafana, Splunk, AppInsights, OpenTelemetry, Dynatrace, Datadog Watchdog, Splunk ITSI, Azure Monitor, GitHub Copilot, ChatGPT
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer- Eng
Lowell, Massachusetts, United States
$130k-$186k/yr HybridFull Time
UKG
UKG: Workforce management and human capital management software provider.
5+ YOE5+ years software/systems/cloud engineering; public cloud experience (GCP/AWS/Azure); observability, SLOs, incident response, Linux, coding in Python/Java/C++; GitHub Actions and dashboarding (Splunk/Grafana).
GCP, AWS, Azure, Python, Java, C++, Linux, GitHub Actions, Splunk, Grafana, Kubernetes, Terraform, Ansible
3mo
Save
Mark Applied
Hide
Sr. Manager, Site Reliability Engineer (SRE)
Boston or Hampton
$160k-$185k/yr HybridFull Time
Planet Fitness
Planet FitnessNYSE: PLNT: Leading fitness center franchisor and operator.
7+ YOE7+ years leading SRE/DevOps with cloud (AWS/Azure/GCP), incident management, SLO/SLI experience, CI/CD and observability expertise; bachelor\u0002s degree or equivalent experience.
AWS, Azure, GCP, CI/CD, Infrastructure as Code (IaC)
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Machine Learning
Cambridge, Massachusetts, United States
$142k-$178k/yr OnsiteFull Time
Cambridge Mobile Telematics
Cambridge Mobile Telematics: Telematics and AI helping insurers, automakers, mobility firms, and public agencies reduce driving risk and crashes.
7+ YOEBachelor's or equivalent, 7+ years SRE/IT experience, AWS (EC2,EKS, S3,RDS), Databricks, Ray, Terraform, Python, Linux, Datadog/CloudWatch, strong incident response and system design skills.
Ray, AWS EKS, Databricks, CloudWatch, Datadog, EC2, S3, RDS/Aurora, Dynamo, SQS, Lambda, IAM, Terraform, Python, Docker, Kubernetes, Unity Catalog, CI/CD
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
London or Singapore or Tokyo or Houston or Boston
HybridFull Time
Veson Nautical
Veson Nautical: Private maritime software serving shipowners, charterers, traders, and operators with commercial freight management solutions.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Waltham, Massachusetts, United States
$130k-$140k/yr HybridFull Time
SS&C Technologies
SS&C TechnologiesNASDAQ: SSNC: Global provider of financial services and healthcare technology solutions.
7+ YOEExperience in Unix/Linux, Java, Python or Bash, Kubernetes, Docker, cloud (AWS/Azure), IaC (Terraform/Ansible), monitoring tools, networking protocols, and on-call incident response.
Unix, Linux, Java, Python, Bash, Kubernetes, Docker, Azure, AWS, CloudWatch, EKS, EFS, S3, RedShift, Terraform, Ansible, HTTP(s), JMS, TCP, UDP, Splunk, Datadog, Dynatrace, Zabbix, Prometheus, Akamai, Cloudflare, DNS, CDN, DataStream, WAF, Oracle, PostgreSQL, MongoDB, RabbitMQ, Interconnect, AMQ
2w
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNYSE: IBM: Global technology and consulting focusing on cloud and AI.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Boston, Massachusetts, United States
$140k-$211k/yr OnsiteFull Time
Federal Reserve Bank of Boston
Federal Reserve Bank of Boston: Government-chartered regional reserve bank serving New England through monetary policy, financial supervision, payments, and community development.
Senior SRE with AWS, Terraform, Docker, Linux, CI/CD, IaC, observability, and automation experience; able to operate large-scale distributed systems in a production environment.
AWS, EC2, EKS, RDS, Aurora, S3, Route 53, ELB, IAM, Terraform, Consul, Vault, Ansible, Python, Java, Go, Docker, ECR, OpenSearch, Dynatrace, Grafana, Prometheus, CloudWatch, ChaosToolkit, Gremlin, Chaos Monkey
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (Senior or Staff)
Boston or Miami or New Jersey or New York City or Princeton or Raleigh or Washington or Toronto or North America
$127k-$249k/yr HybridFull Time
MongoDB
MongoDBNASDAQ: MDB: Unified data platform for building modern applications.
6+ YOE6+ years software development experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud (AWS, Google Cloud Platform, Microsoft Azure) expertise; Linux and networking knowledge.
Argo Workflows, ArgoCD, Kubernetes, Python, Go, AWS, Google Cloud Platform (GCP), Microsoft Azure, Linux
2mo
Save
Mark Applied
Hide
Senior DevOps Site Reliability Engineer
North Andover, Massachusetts, United States
HybridFull Time
TSD Mobility Solutions
TSD Mobility Solutions: Provider of automotive dealership software and business solutions.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Jenkins, AWS, EC2, S3, RDS, Lambda, VPC, IAM, Windows IIS, Linux, Azure DevOps, Bash, Python, PowerShell, Terraform, Ansible, Web Deploy (MSDeploy), PowerShell DSC, Docker, Kubernetes, Datadog, Grafana, CloudWatch
2mo
Save
Mark Applied
Hide
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yr OnsiteFull Time
QuEra Computing
QuEra Computing: Neutral-atom quantum computing building quantum computers for researchers, businesses, governments, and high-performance computing centers.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Hardware-in-the-loop (HIL), Kubernetes, Docker, Git, Python, Bash, Go, GitLab CI, Jenkins, Ansible, Terraform, Grafana, Prometheus, ELK stack, CI/CD, Ubuntu, Debian, Redhat, Linux, Windows, VLAN, DNS, DHCP, TCP/IP
3mo
Save
Mark Applied
Hide
Site Reliability Engineer - Disaster Recovery & Business Continuity
Boston or Chicago
$130k-$150k/yr HybridFull Time
Charles River Associates
Charles River AssociatesNasdaq Global Select Market: CRAI: Public global economic, financial, and management consulting firm serving law firms, corporations, accounting firms, and governments.
Experience with IT service continuity, disaster recovery, cross-functional coordination, and documentation.
Windows, Microsoft 365, Cloud, SaaS, Backups, Virtualization, Identity, Networking
2w
Save
Mark Applied
Hide
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yr HybridFull Time
Zelis
Zelis: Private healthcare financial technology serving payers, providers, and healthcare consumers with claims, payment, and member-engagement solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management