25 infrastructure reliability engineer jobs at 19 companies in Acton, MA
1mo
Save
Mark Applied
Hide
1mo
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Senior Site Reliability Engineer, Fleet Infrastructure
Boston or Washington
$166k-$220k/yrOnsiteFull Time
Anduril Industries: Defense technology building autonomous military hardware and software.
8+ YOEBuild and operate high-availability observability and telemetry systems, collaborate with engineering teams, participate in on-call rotations, and meet U.S. Person access requirements.
DraftKingsNASDAQ: DKNG: Provide online sports betting, fantasy sports, and casino gaming.
2+ YOE2+ years in DRE/SRE or related role, experience with relational and NoSQL databases, Kubernetes stateful workloads, infrastructure as code, Go or Python development, observability and reliability practices.
Manifold: AI platform for life sciences data and research collaboration.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
Senior Site Reliability Engineer - Government Cloud
Boston or Dublin or United States
$210k-$220k/yrRemoteFull Time
Tines: No-code workflow automation for security and IT teams.
5+ YOE5+ years in infrastructure/DevOps/cloud engineering with strong AWS experience; hands-on IaC (CDK or Terraform), container image pipelines and hardening, observability, FedRAMP/CMMC/FISMA familiarity, documentation and assessment experience; U.S. citizenship required.
Member of Technical Staff – Senior Engineer, Data Infrastructure & Data Operations
San Francisco or Cambridge
$255k-$340k/yrOnsiteFull Time
Walden Robotics: Builds general-purpose robots and develops the teams and infrastructure to scale robot applications and improve quality of life.
Experience building production data infrastructure and high-throughput pipelines, cloud-based data platform development, platform reliability and cost ownership, and collaboration with ML teams.
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yrOnsiteFull Time
QuEra Computing: Develops and operates neutral-atom quantum computing systems.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Principal Infrastructure & Systems Software Engineer
Newton, Massachusetts, United States
OnsiteFull Time
Magnendo: Develops robotic magnetic navigation technology for endovascular interventions.
9+ YOE9+ years in systems/infrastructure software engineering; expertise in OS selection, real-time orchestration, IPC, hardware support; proficiency with Python, C++, Linux, GIT; experience with high-reliability and regulatory environments.
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman: Platform for building, testing, and managing software APIs.
Experience leading engineering teams building GenAI or AI infrastructure and distributed systems; strong cloud, accelerator, and performance optimization knowledge; proficiency in Python or Go; architecture and reliability experience.
Reynolds and Reynolds: Provides software and services for automotive retailers.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Senior Director, Reliability and Security Engineering
Boston or United States
OnsiteFull Time
Beacon Biosignals: Provides AI-powered EEG monitoring and analytics for brain health.
10+ YOE5+ Mgmt10+ years in SRE/DevOps/infrastructure/security engineering with 5+ years managing engineering teams; experience building security practices, infra-as-code, incident response, and compliance in regulated environments.
Kubernetes, SOC 2, ISO 27001, HITRUST r2, infrastructure-as-code, policy-as-code
Transdev: Operator of public transportation systems and mobility solutions.
10+ YOEMinimum 10 years railroad engineering and infrastructure management experience; unionized environment experience; knowledge of FRA regulations, PTC coordination, asset reliability, and facilities management; PE preferred; bachelor\u0002s degree required.
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
CloudZero: Platform for cloud cost intelligence and FinOps optimization.
5+ YOE5+ years building and operating distributed systems in AWS; strong production Python; SLO and reliability experience; infrastructure as code; observability and debugging experience.
OnRamp: SaaS platform automating B2B customer onboarding and implementation processes.
Deep hands-on AWS experience, infrastructure-as-code (Terraform or CDK), containers and CI/CD, building observability/reliability, security/compliance (SOC 2/HIPAA), and using AI/LLM tooling and coding agents to automate operations.
AxonNASDAQ: AXON: Develops weapons and software for law enforcement and safety.
Deep VoIP/UC and SIP experience, cloud platform expertise ideally with AWS, familiarity with FreeSWITCH or Kamailio, and the ability to build reliable distributed voice infrastructure.