137 site reliability engineer jobs at 87 companies in New York
2mo
Save
Mark Applied
Hide
2mo
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Mistral AI: Developer of open-weight and frontier AI models.
7+ YOE7+ years SRE/DevOps experience, Master’s in CS/Engineering or related, strong cloud and distributed systems skills, Kubernetes/CI-CD/infra-as-code proficiency, scripting experience, observability and on-call experience.
M&T BankNYSE: MTB: A diversified financial services providing banking and wealth management.
7+ YOEExpert in reliability engineering, SLO/SLI frameworks, incident and problem management, observability, automation, cloud platforms, and production operations; 7+ years systems analysis/application development or equivalent.
Greenhouse: Private hiring software helping organizations source, interview, and onboard candidates with structured, AI-powered recruiting tools.
3+ YOE3+ years in site reliability or infrastructure, production AWS and Kubernetes experience, software delivery experience, and strong fluency with AI coding tools such as Claude. Must be eligible to work in Canada.
Sunset: Private startup wind-down service helping founders close companies through legal, tax, and operational work.
Production cloud infrastructure and reliability experience across multiple services, strong software engineering skills, infrastructure and application coding, incident leadership, recovery expertise, and AI engineering tool proficiency.
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yrHybridFull Time
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
MarketAxessNASDAQ: MKTX: Public financial-technology operating an electronic fixed-income trading platform for institutional investors and broker-dealers.
Experience with Java/Python, shell scripting, Linux, SQL, CI/CD (Jenkins), FIX, containers, AWS, monitoring, SRE practices, automation, and strong communication; Bachelor\u0002s/Master\u0002s in CS/Engineering or related field.
6+ YOE6+ years in SRE/DevOps with production Kubernetes, CI/CD pipeline expertise, IaC (Terraform/Helm/Pulumi), Python/Bash/Go proficiency, observability tooling, and experience with Kafka/Flink/ClickHouse and AI/LLM tooling integration.
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
Requires a computer science bachelor's degree or equivalent experience, Python, production ML inference, AWS cloud infrastructure, Kubernetes, vulnerability management, distributed-systems debugging, and on-call participation.
Retool: Private software providing an internal-tools development platform for business and enterprise teams.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Ripple: Enables institutions to move, manage, and tokenize value.
7+ YOE7+ years SRE/Platform experience focused on observability, New Relic, Terraform, PowerShell, Azure/AWS, incident management (Incident.IO/PagerDuty/OpsGenie), and coaching engineering teams.
New Relic, NRQL, Terraform, PowerShell, Azure, AWS, Azure DevOps, Octopus Deploy, Incident.IO, PagerDuty, OpsGenie, Slack, Python, Bash, Jira, SQL Server
Optimum CommunicationsNYSE: OPTU: US broadband communications and video services provider.
6+ YOEBachelor's in related field, 6+ years IP networking experience (4+ in reliability), expert routing protocol knowledge, automation (Python, Ansible, Terraform), observability (Grafana, Prometheus, ELK), and leadership/documentation skills.
Longbridge Group: Singapore-headquartered fintech group operating online brokerage, institutional trading technology, and AI financial infrastructure for investors and institutions.
5+ YOERequires 5+ years in SRE, DevOps, or production engineering; AWS or equivalent cloud expertise; Docker, Kubernetes, Linux, CI/CD, programming, incident management, and distributed-systems troubleshooting.