Bachelor's in Computer Science/Engineering or equivalent; experience in SRE/Software Engineering for large-scale distributed systems; Terraform and IAC experience; familiarity with SaltStack/Ansible/Chef/Puppet; Linux, CI/CD, observability, and participation in on-call rotation.
Anduril Industries: Defense technology building autonomous military hardware and software.
3+ YOE3+ years SRE/DevOps/field or production support experience; strong Linux and networking fundamentals; on-call rotation experience; ability to diagnose cross-stack issues; eligibility for U.S. Secret clearance.
Manifold: AI platform for life sciences data and research collaboration.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
Tulip: Provides a no-code platform for industrial frontline operations.
5+ YOE5+ years experience with observability tools, OpenTelemetry instrumentation, Prometheus metrics, experience with time-series data and producing SLIs/SLOs, strong systems reasoning and communication.
Senior Site Reliability Engineer - Government Cloud
Boston or Dublin or United States
$210k-$220k/yrRemoteFull Time
Tines: No-code workflow automation for security and IT teams.
5+ YOE5+ years in infrastructure/DevOps/cloud engineering with strong AWS experience; hands-on IaC (CDK or Terraform), container image pipelines and hardening, observability, FedRAMP/CMMC/FISMA familiarity, documentation and assessment experience; U.S. citizenship required.
Flock Safety: Sells AI-powered cameras and software for public safety surveillance.
Experience writing production Go or TypeScript, proficiency with Kubernetes, Helm, Terraform, GitHub Actions, and AWS; observability and CI/CD expertise; participates in on-call rotations.
CVS HealthNYSE: CVS: Provides retail pharmacy, health insurance, and pharmacy benefit management services.
5+ YOE5+ years SRE/DevOps experience, observability and monitoring expertise, cloud and containerization knowledge, 2+ years Java/Python and AI/AIOps experience, CI/CD and source control experience.
Splunk, Dynatrace, Datadog, Prometheus, Grafana, Java, Python, AWS, Microsoft Azure, Google Cloud, Rancher, Docker, Kubernetes, OpenShift, GitHub, BitBucket, Jenkins, Apigee, Data power
Planet FitnessNYSE: PLNT: Operates a global network of low-cost fitness centers.
7+ YOE7+ years leading SRE/DevOps with cloud (AWS/Azure/GCP), incident management, SLO/SLI experience, CI/CD and observability expertise; bachelor\u0002s degree or equivalent experience.
AWS, Azure, GCP, CI/CD, Infrastructure as Code (IaC)
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yrOnsiteFull Time
QuEra Computing: Develops and operates neutral-atom quantum computing systems.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Veson Nautical: Develops enterprise software for global maritime freight management.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
Federal Reserve System: The central bank of the United States.
Senior SRE with AWS, Terraform, Docker, Linux, CI/CD, IaC, observability, and automation experience; able to operate large-scale distributed systems in a production environment.
Boston or Miami or New Jersey or New York City or Princeton or Raleigh or Washington or Toronto or North America
$127k-$249k/yrHybridFull Time
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud (AWS, Google Cloud Platform, Microsoft Azure) expertise; Linux and networking knowledge.
Argo Workflows, ArgoCD, Kubernetes, Python, Go, AWS, Google Cloud Platform (GCP), Microsoft Azure, Linux
Epalinges or Lausanne or Menlo Park or Boston or Europe
HybridFull Time
Atinary Technologies: AI platform for autonomous materials discovery and R&D optimization.
3+ YOERequires 3+ years in DevOps, site reliability, or cloud infrastructure, or 2+ software engineering and 1+ DevOps years; AWS, CI/CD, containers, Python, Bash, and Infrastructure as Code experience required.
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yrHybridFull Time
Zelis: Providing healthcare payment and claims cost management solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management
Harvey or Philadelphia or New Bedford or New Bedford or Cartersville or Dallas
$126k-$174k/yrFieldFull Time
AtkoreNYSE: ATKR: Manufactures electrical, mechanical, and safety infrastructure products.
8+ YOEBachelor’s degree in engineering, 8+ years in engineering or maintenance in steel manufacturing, multi-site responsibility, maintenance reliability, project management, lean methods, TPM, CMMS, and leadership skills.
Total Productive Maintenance (TPM), Computerized Maintenance Management Systems (CMMS), Overall Equipment Effectiveness (OEE), Mean Time Between Failure (MTBF), Mean Time to Repair (MTTR), Failure Modes Effects Analysis (FMEA)