5+ YOE5+ years relevant experience, Bachelor's in Computer Engineering/CS or equivalent, proficiency in Python, observability stacks, networking (BGP, IPv4/IPv6), incident response, and building automation and runbooks.
Federal Reserve System: The central bank of the United States.
Senior SRE with AWS, Terraform, Docker, Linux, CI/CD, IaC, observability, and automation experience; able to operate large-scale distributed systems in a production environment.
Tulip: Provides a no-code platform for industrial frontline operations.
5+ YOE5+ years experience with observability tools, OpenTelemetry instrumentation, Prometheus metrics, experience with time-series data and producing SLIs/SLOs, strong systems reasoning and communication.
CVS HealthNYSE: CVS: Provides retail pharmacy, health insurance, and pharmacy benefit management services.
5+ YOE5+ years SRE/DevOps experience, observability and monitoring expertise, cloud and containerization knowledge, 2+ years Java/Python and AI/AIOps experience, CI/CD and source control experience.
Splunk, Dynatrace, Datadog, Prometheus, Grafana, Java, Python, AWS, Microsoft Azure, Google Cloud, Rancher, Docker, Kubernetes, OpenShift, GitHub, BitBucket, Jenkins, Apigee, Data power
Senior Site Reliability Engineer, Fleet Infrastructure
Boston or Washington
$166k-$220k/yrOnsiteFull Time
Anduril Industries: Defense technology building autonomous military hardware and software.
8+ YOEBuild and operate high-availability observability and telemetry systems, collaborate with engineering teams, participate in on-call rotations, and meet U.S. Person access requirements.
Senior Site Reliability Engineer - Government Cloud
Boston or Dublin or United States
$210k-$220k/yrRemoteFull Time
Tines: No-code workflow automation for security and IT teams.
5+ YOE5+ years in infrastructure/DevOps/cloud engineering with strong AWS experience; hands-on IaC (CDK or Terraform), container image pipelines and hardening, observability, FedRAMP/CMMC/FISMA familiarity, documentation and assessment experience; U.S. citizenship required.
Manifold: AI platform for life sciences data and research collaboration.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
Senior/Staff Site Reliability Engineer - Data Center
Boston, Massachusetts, United States
$166k-$224k/yrHybridFull Time
PathAI: AI-powered platform for pathology research and clinical diagnostics.
8+ YOERequires 8+ years of relevant experience, infrastructure operations expertise, a bachelor's degree in Computer Science or equivalent experience, and knowledge of datacenter, virtualization, storage, automation, monitoring, and incident response.
DraftKingsNASDAQ: DKNG: Provide online sports betting, fantasy sports, and casino gaming.
4+ YOE4+ years managing distributed cloud and on‑prem environments, strong AWS and Kubernetes experience, proficiency in Go or Python, networking and Linux knowledge, bachelor's in CS or equivalent experience.
Flock Safety: Sells AI-powered cameras and software for public safety surveillance.
Experience writing production Go or TypeScript, proficiency with Kubernetes, Helm, Terraform, GitHub Actions, and AWS; observability and CI/CD expertise; participates in on-call rotations.
Planet FitnessNYSE: PLNT: Operates a global network of low-cost fitness centers.
7+ YOE7+ years leading SRE/DevOps with cloud (AWS/Azure/GCP), incident management, SLO/SLI experience, CI/CD and observability expertise; bachelor\u0002s degree or equivalent experience.
AWS, Azure, GCP, CI/CD, Infrastructure as Code (IaC)
Boston or Miami or New Jersey or New York City or Princeton or Raleigh or Washington or Toronto or North America
$127k-$249k/yrHybridFull Time
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud (AWS, Google Cloud Platform, Microsoft Azure) expertise; Linux and networking knowledge.
Argo Workflows, ArgoCD, Kubernetes, Python, Go, AWS, Google Cloud Platform (GCP), Microsoft Azure, Linux
Veson Nautical: Develops enterprise software for global maritime freight management.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yrOnsiteFull Time
QuEra Computing: Develops and operates neutral-atom quantum computing systems.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yrHybridFull Time
Zelis: Providing healthcare payment and claims cost management solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management