30 infrastructure reliability engineer jobs at 23 companies in Milton, MA
19h
Save
Mark Applied
Hide
19h
Customer Reliability Engineer, Infrastructure
United States or Austin or New York City or Boston or San Francisco
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOERequires 5 years of experience with complex cloud infrastructure, 3 years with Kubernetes, production distributed systems, Linux, monitoring, troubleshooting, customer support, DevOps or CI/CD, and Python scripting.
Senior Site Reliability Engineer, Fleet Infrastructure
Boston or Washington
$166k-$220k/yrOnsiteFull Time
Anduril Industries: Defense technology building autonomous military hardware and software.
8+ YOEBuild and operate high-availability observability and telemetry systems, collaborate with engineering teams, participate in on-call rotations, and meet U.S. Person access requirements.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
DraftKingsNASDAQ: DKNG: Provide online sports betting, fantasy sports, and casino gaming.
2+ YOE2+ years in DRE/SRE or related role, experience with relational and NoSQL databases, Kubernetes stateful workloads, infrastructure as code, Go or Python development, observability and reliability practices.
Manifold: AI platform for life sciences data and research collaboration.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
Senior Site Reliability Engineer - Government Cloud
Boston or Dublin or United States
$210k-$220k/yrRemoteFull Time
Tines: No-code workflow automation for security and IT teams.
5+ YOE5+ years in infrastructure/DevOps/cloud engineering with strong AWS experience; hands-on IaC (CDK or Terraform), container image pipelines and hardening, observability, FedRAMP/CMMC/FISMA familiarity, documentation and assessment experience; U.S. citizenship required.
Member of Technical Staff – Senior Engineer, Data Infrastructure & Data Operations
San Francisco or Cambridge
$255k-$340k/yrOnsiteFull Time
Walden Robotics: Builds general-purpose robots and develops the teams and infrastructure to scale robot applications and improve quality of life.
Experience building production data infrastructure and high-throughput pipelines, cloud-based data platform development, platform reliability and cost ownership, and collaboration with ML teams.
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yrOnsiteFull Time
QuEra Computing: Develops and operates neutral-atom quantum computing systems.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Principal Infrastructure & Systems Software Engineer
Newton, Massachusetts, United States
OnsiteFull Time
Magnendo: Develops robotic magnetic navigation technology for endovascular interventions.
9+ YOE9+ years in systems/infrastructure software engineering; expertise in OS selection, real-time orchestration, IPC, hardware support; proficiency with Python, C++, Linux, GIT; experience with high-reliability and regulatory environments.
Reynolds and Reynolds: Provides software and services for automotive retailers.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Transdev: Operator of public transportation systems and mobility solutions.
10+ YOEMinimum 10 years railroad engineering and infrastructure management experience; unionized environment experience; knowledge of FRA regulations, PTC coordination, asset reliability, and facilities management; PE preferred; bachelor\u0002s degree required.
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
United States or Canada or New York City or San Francisco or Boston or Toronto or Chicago or Los Angeles or Washington
$160k-$190k/yrRemoteFull Time
Voltus: Connecting distributed energy resources to electricity markets for revenue
6+ YOE6+ years of software engineering experience, including DevOps/SRE production operations. Requires Go and/or Python, deep AWS and Kubernetes expertise, Terraform, observability, stateful systems, reliability ownership, and on-call experience.
CloudZero: Platform for cloud cost intelligence and FinOps optimization.
5+ YOE5+ years building and operating distributed systems in AWS; strong production Python; SLO and reliability experience; infrastructure as code; observability and debugging experience.
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yrHybridFull Time
Zelis: Providing healthcare payment and claims cost management solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management
OnRamp: SaaS platform automating B2B customer onboarding and implementation processes.
Deep hands-on AWS experience, infrastructure-as-code (Terraform or CDK), containers and CI/CD, building observability/reliability, security/compliance (SOC 2/HIPAA), and using AI/LLM tooling and coding agents to automate operations.
Epalinges or Lausanne or Menlo Park or Boston or Europe
HybridFull Time
Atinary Technologies: AI platform for autonomous materials discovery and R&D optimization.
3+ YOERequires 3+ years in DevOps, site reliability, or cloud infrastructure, or 2+ software engineering and 1+ DevOps years; AWS, CI/CD, containers, Python, Bash, and Infrastructure as Code experience required.
Ahold Delhaize USAEuronext Amsterdam: AD: Operates a portfolio of omnichannel grocery brands.
8+ YOE8+ years in platform or infrastructure engineering, automation, CI/CD, container/runtime services, incident response, reliability engineering, and effective communication. Bachelor's degree or equivalent experience required.