30 infrastructure reliability engineer jobs at 22 companies in Salem, NH

2d
Save
Mark Applied
Hide
Customer Reliability Engineer, Infrastructure
United States or Austin or New York City or Boston or San Francisco
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
5+ YOERequires 5 years of experience with complex cloud infrastructure, 3 years with Kubernetes, production distributed systems, Linux, monitoring, troubleshooting, customer support, DevOps or CI/CD, and Python scripting.
Astro, Apache Airflow, Kubernetes, AWS, GCP, Azure, Linux, Python
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Fleet Infrastructure
Boston or Washington
$166k-$220k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology developing AI-powered autonomous military systems.
8+ YOEBuild and operate high-availability observability and telemetry systems, collaborate with engineering teams, participate in on-call rotations, and meet U.S. Person access requirements.
Lattice OS, Docker, Kubernetes, AWS, GCP, Azure, ClickHouse, ClickStack, Victoria Metrics, Prometheus, Grafana, ELK
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
3w
Save
Mark Applied
Hide
Database Reliability Engineer
Boston or United States
$112k-$140k/yr HybridFull Time
DraftKings
DraftKingsNASDAQ: DKNG: Provide online sports betting, fantasy sports, and casino gaming.
2+ YOE2+ years in DRE/SRE or related role, experience with relational and NoSQL databases, Kubernetes stateful workloads, infrastructure as code, Go or Python development, observability and reliability practices.
PostgreSQL, MySQL, Aurora MySQL, MongoDB, Redis, ScyllaDB, Aerospike, Kubernetes, StatefulSets, Persistent Volumes, database operators, Go, Python, Terraform, Pulumi, FluxCD, ArgoCD, GKE, EKS, GitOps, GitHub Copilot, Claude, Cursor, MCP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Newton, Massachusetts, United States
$160k-$205k/yr OnsiteFull Time
Manifold
Manifold: The Enterprise Agent Platform for life sciences that helps biopharma and research teams analyze governed biomedical data.
7+ YOE7+ years in infrastructure/DevOps/SRE with deep cloud (AWS/GCP/Azure), Terraform, CI/CD (Github Action), container tooling, identity systems, data platform services, and experience operating secure multi-account environments.
AWS, GCP, Azure, Terraform, Github Action, Okta, Auth0, Docker, ECS, Packer, Tailscale, WireGuard, Snowflake, Airflow, dbt, PostgreSQL, LLM, CI/CD
1mo
Save
Mark Applied
Hide
Member of Technical Staff – Senior Engineer, Data Infrastructure & Data Operations
San Francisco or Cambridge
$255k-$340k/yr OnsiteFull Time
Walden Robotics
Walden Robotics: Private full-stack physical AI building and deploying general-purpose robots for manufacturing and logistics.
Experience building production data infrastructure and high-throughput pipelines, cloud-based data platform development, platform reliability and cost ownership, and collaboration with ML teams.
2mo
Save
Mark Applied
Hide
Sr. Control System Engineer/Site Reliability Engineer (SRE)
Boston, Massachusetts, United States
$160k-$225k/yr OnsiteFull Time
QuEra Computing
QuEra Computing: Neutral-atom quantum computing building quantum computers for researchers, businesses, governments, and high-performance computing centers.
10+ YOEDesign, implement, and maintain hardware and software control systems for quantum computers; strong Linux/Windows administration, networking (LAN/WAN/VLAN/DNS/DHCP/TCP/IP), scripting (Python/Bash/Go), containerization, CI/CD, infrastructure-as-code, observability, and rack server experience; 10+ years experience.
Hardware-in-the-loop (HIL), Kubernetes, Docker, Git, Python, Bash, Go, GitLab CI, Jenkins, Ansible, Terraform, Grafana, Prometheus, ELK stack, CI/CD, Ubuntu, Debian, Redhat, Linux, Windows, VLAN, DNS, DHCP, TCP/IP
1mo
Save
Mark Applied
Hide
Principal Infrastructure & Systems Software Engineer
Newton, Massachusetts, United States
OnsiteFull Time
Magnendo
Magnendo: Medical device startup developing robotic navigation systems for neurovascular stroke and aneurysm procedures.
9+ YOE9+ years in systems/infrastructure software engineering; expertise in OS selection, real-time orchestration, IPC, hardware support; proficiency with Python, C++, Linux, GIT; experience with high-reliability and regulatory environments.
Python, C++, Linux, GIT, Yocto, Gentoo
2mo
Save
Mark Applied
Hide
Senior DevOps Site Reliability Engineer
North Andover, Massachusetts, United States
HybridFull Time
TSD Mobility Solutions
TSD Mobility Solutions: Provider of automotive dealership software and business solutions.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Jenkins, AWS, EC2, S3, RDS, Lambda, VPC, IAM, Windows IIS, Linux, Azure DevOps, Bash, Python, PowerShell, Terraform, Ansible, Web Deploy (MSDeploy), PowerShell DSC, Docker, Kubernetes, Datadog, Grafana, CloudWatch
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEBS, MS, or PhD in computer science or electrical engineering, or equivalent experience; 6+ years in systems programming; expertise in distributed systems, C++ and Python, HPC or cloud RCA pipelines, CPU metrics, and cluster managers.
C++, Python, Slurm, Kubernetes, LSF, Linux kernel, /dev/mcelog, dmesg, journald, NVIDIA DCGM (Data Center GPU Manager), NVIDIA Management Library (NVML), CRIU
3w
Save
Mark Applied
Hide
Head of Infrastructure
Boston, Massachusetts, United States
$200k-$280k/yr OnsiteFull Time
Transdev North America
Transdev North America: Provides public transportation, mobility solutions, and fleet maintenance services.
10+ YOEMinimum 10 years railroad engineering and infrastructure management experience; unionized environment experience; knowledge of FRA regulations, PTC coordination, asset reliability, and facilities management; PE preferred; bachelor\u0002s degree required.
PTC
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald
2mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
C++, Python, Slurm, LSF, Kubernetes, CUDA, DCGM, NVML, CRIU, /dev/mcelog, dmesg, journald, Linux kernel
1d
Save
Mark Applied
Hide
Senior Software Engineer, Infrastructure
United States or Canada or New York City or San Francisco or Boston or Toronto or Chicago or Los Angeles or Washington
$160k-$190k/yr RemoteFull Time
Voltus
Voltus: Privately held virtual power plant operator that connects commercial, industrial, residential, and transportation energy resources to electricity markets.
6+ YOE6+ years of software engineering experience, including DevOps/SRE production operations. Requires Go and/or Python, deep AWS and Kubernetes expertise, Terraform, observability, stateful systems, reliability ownership, and on-call experience.
GitHub, Kubernetes, HashiCorp Nomad, HashiCorp Consul, HashiCorp Vault, AWS, Python, Postgres, Go, FastAPI, Temporal, Delta Lake, ClickHouse, TypeScript, React, Docker, Terraform, Buildkite, Prometheus, Grafana, Elasticsearch, OpenSearch, Argo CD, Flux, Helm, Jenkins, OpenTelemetry, Amazon MSK, AWS Organizations, AWS Control Tower, Auth0, Okta, Amazon Cognito, Keycloak, Java, C++, Claude Code, MCP, AWS Bedrock, GitOps
3w
Save
Mark Applied
Hide
Senior CloudOps Engineer
Boston or San Francisco
$130k-$190k/yr HybridFull Time
CloudZero
CloudZero: Private SaaS platform helping finance, IT, and engineering teams optimize cloud and AI costs.
5+ YOE5+ years building and operating distributed systems in AWS; strong production Python; SLO and reliability experience; infrastructure as code; observability and debugging experience.
Python, Kafka, Kinesis, SQS, Pulsar, Step Functions, MSK, CloudFormation, SAM, Terraform, Pulumi, Sumo Logic, Datadog, Prometheus, Splunk, Cortex, Backstage, GitHub Actions, Claude, Codex, Gemini, AWS, Azure, GCP
2w
Save
Mark Applied
Hide
Senior Manager, Site Reliability & Operational Resilience
Morristown or Boston or St. Petersburg or St. Louis or Atlanta or Hyderabad
$139k-$177k/yr HybridFull Time
Zelis
Zelis: Providing healthcare payment and claims cost management solutions.
8+ YOE3+ MgmtRequires 8+ years in SRE, production, platform, DevOps, cloud, or infrastructure engineering; 3+ years leading people; enterprise resilience experience; bachelor's degree or equivalent; no visa sponsorship.
LogicMonitor, New Relic, Splunk, Datadog, Python, PowerShell, Go, Terraform, Azure, AWS, Kubernetes, OpenTelemetry, Jira Service Management
1mo
Save
Mark Applied
Hide
Devops / Platform Engineer
Boston, Massachusetts, United States
HybridFull Time
OnRamp
OnRamp: AI-enabled B2B customer onboarding and engagement platform helping enterprises accelerate adoption, relationships, and revenue.
Deep hands-on AWS experience, infrastructure-as-code (Terraform or CDK), containers and CI/CD, building observability/reliability, security/compliance (SOC 2/HIPAA), and using AI/LLM tooling and coding agents to automate operations.
AWS, Terraform, CDK, CI/CD, LLM
1w
Save
Mark Applied
Hide
DevOps Engineer
Epalinges or Lausanne or Menlo Park or Boston or Europe
HybridFull Time
Atinary Technologies
Atinary Technologies: AI platform for autonomous materials discovery and R&D optimization.
3+ YOERequires 3+ years in DevOps, site reliability, or cloud infrastructure, or 2+ software engineering and 1+ DevOps years; AWS, CI/CD, containers, Python, Bash, and Infrastructure as Code experience required.
Python, AWS, GitHub Actions, Bash, Terraform, OpenTofu, Infrastructure as Code (IaC)
1mo
Save
Mark Applied
Hide
ServiceNow Platform Engineer IV
Salisbury or Chicago or Quincy
$147k-$220k/yr HybridFull Time
Ahold Delhaize USA
Ahold Delhaize USAEuronext Amsterdam: AD: The U.S. service division for a global grocery retailer.
8+ YOE8+ years in platform or infrastructure engineering, automation, CI/CD, container/runtime services, incident response, reliability engineering, and effective communication. Bachelor's degree or equivalent experience required.
ServiceNow, CI/CD
1mo
Save
Mark Applied
Hide
Service Now Platform Engineer IV
Salisbury or Chicago or Quincy
$147k-$220k/yr HybridFull Time
Ahold Delhaize USA
Ahold Delhaize USAEuronext Amsterdam: AD: The U.S. service division for a global grocery retailer.
8+ YOE8+ years platform or infrastructure engineering experience; bachelor's degree or equivalent; strong automation, CI/CD, container/runtime experience; incident response and reliability engineering skills; effective communication.
ServiceNow, CI/CD