106 site reliability manager jobs at 68 companies in New York

4d
Save
Mark Applied
Hide
Manager, Site Reliability Engineer
San Francisco or New York City
$150k-$220k/yr OnsiteFull Time
Forge Global
Forge GlobalNYSE: FRGE: Financial technology operating a private-market marketplace and data, custody, and investment solutions for companies and investors.
10+ YOE5+ MgmtRequires 5+ years leading SRE, DevOps, cloud operations, or reliability functions; 10+ years in engineering or operations; bachelor's degree or equivalent; cloud infrastructure, distributed systems, observability, CI/CD, and automation experience.
AWS, Azure, Kubernetes, Terraform, Ansible, Datadog, CloudWatch
2d
Save
Mark Applied
Hide
Site Reliability Engineer - Enterprise Technology
New York City, New York, United States
$200k-$250k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: Private quantitative trading firm providing liquidity across global markets and directly to financial-market clients.
5+ YOERequires 5+ years in site reliability or related disciplines, Python, Linux, Kubernetes, observability, containerized infrastructure, CI/CD, IaC, configuration management, and cloud platform experience.
Linux, Kubernetes, Python, Jenkins, GitHub Actions, ArgoCD, Terraform, SaltStack, Chef, Puppet, Ansible, AWS, Azure, GCP
2d
Save
Mark Applied
Hide
Site Reliability Engineer - Enterprise Technology
New York City, New York, United States
$200k-$250k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: Private quantitative trading firm providing liquidity across global markets and directly to financial-market clients.
5+ YOERequires 5+ years in site reliability or related disciplines, Python, containerized infrastructure, CI/CD, IaC, configuration management, cloud platforms, and Linux, Kubernetes, and observability expertise.
Linux, Kubernetes, Python, Jenkins, GitHub Actions, ArgoCD, Terraform, SaltStack, Chef, Puppet, Ansible, AWS, Azure, GCP
1w
Save
Mark Applied
Hide
Site Reliability Engineer
New York, United States
HybridFull Time
Longbridge Group
Longbridge Group: Singapore-headquartered fintech group operating online brokerage, institutional trading technology, and AI financial infrastructure for investors and institutions.
5+ YOERequires 5+ years in SRE, DevOps, or production engineering; AWS or equivalent cloud expertise; Docker, Kubernetes, Linux, CI/CD, programming, incident management, and distributed-systems troubleshooting.
Terraform, Ansible, Helm, Kubernetes, Prometheus, AWS, GCP, Docker, Python, Go, Linux, CI/CD
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Buffalo, New York, United States
$140k-$233k/yr OnsiteFull Time
M&T Bank
M&T BankNYSE: MTB: A diversified financial services providing banking and wealth management.
7+ YOEExpert in reliability engineering, SLO/SLI frameworks, incident and problem management, observability, automation, cloud platforms, and production operations; 7+ years systems analysis/application development or equivalent.
AWS, Azure, CI/CD, SDLC, SLO/SLI
1w
Save
Mark Applied
Hide
Site Reliability Engineer
New York City, New York, United States
HybridFull Time
Chariot
Chariot: US fintech helping nonprofits receive and process donor-advised fund gifts and grant payments.
4+ YOERequires 4+ years software development, 2+ years backend application development, bachelor's degree preferred, and proficiency with Go, Node, Docker, Terraform, Kubernetes, Postgres, REST APIs, gRPC, and AWS.
Go, Node, Docker, Terraform, Kubernetes, Postgres, REST APIs, gRPC, AWS
3d
Save
Mark Applied
Hide
Manager, Site Reliability Engineering (Auth0)
New York City or Washington or California or Colorado or Illinois or Washington
$182k-$251k/yr HybridFull Time
Auth0
Auth0Nasdaq: OKTA: Public American identity-security providing cloud-based authentication, authorization, and access-management services to organizations.
8+ YOE3+ MgmtRequires 8+ years industry experience, 3+ years SRE or software engineering team leadership, AWS/Azure, Terraform, containers, Kubernetes, microservices, databases, Go or Python, and U.S. Person status.
AWS, Azure, Terraform, Kubernetes, Go, Python
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York, New York, United States
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services and investment banking firm.
5+ YOEFormal SRE training or certification, 5+ years applied SRE/production management experience, leadership in production support, experience with Front Office sales platforms, observability and incident management expertise.
Dynatrace, Splunk, Geneos, Grafana, AWS, Microsoft Azure, GCP, Python, Shell, PowerShell, Ansible, Terraform, Kubernetes, OpenShift
2d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Scottsdale or San Francisco or Chicago or New York City or Phoenix or Chicago or San Francisco
$118k-$183k/yr HybridFull Time
Early Warning Services
Early Warning Services: U.S. bank-owned fintech and consumer reporting agency providing identity, fraud-risk, and real-time payment solutions to financial institutions.
5+ YOEBachelor's degree in business, computer science, or related field; 5+ years of technical experience; incident management, Linux, scripting, observability, Git, security protocols, and enterprise production experience.
Observability, Linux, Git, Java, Ruby, Python, JavaScript, Go, CI/CD, TCP/UDP/IP, AWS, Docker, Kubernetes, Swarm
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York, New York, United States
$178k-$258k/yr RemoteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
Requires a computer science bachelor's degree or equivalent experience, Python, production ML inference, AWS cloud infrastructure, Kubernetes, vulnerability management, distributed-systems debugging, and on-call participation.
Python, PHP, Node.js, Ruby, SageMaker, OpenAI, Bedrock, EC2 Auto Scaling Groups, Kubernetes, AWS, Azure, GCP, AMI, LangGraph, LLM gateway, MCP, Aurora PostgreSQL, Memcached, Terraform, Terragrunt, Atlantis, Chef, Ansible, SSM, Docker, bash, Jenkins, Argo CD, New Relic, Splunk, Grafana, Prometheus, Fastly, Datadome, WAF, vLLM, LangSmith, Lambda
3d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Santa Barbara or California or Illinois or Texas or Minnesota or Colorado or Georgia or New York or Massachusetts or Connecticut
$175k-$230k/yr HybridFull Time
PayJunction
PayJunction: Privately held U.S. payment processor serving businesses with in-store, online, and mobile payments.
10+ years relevant experience; 5+ years Linux administration; 3+ years AWS, containers, infrastructure as code, configuration management, and automation scripting; physical server and data center experience required.
AWS, Terraform, Puppet, OpenVox, Ansible, Linux, containers, scripting, BI
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
2mo
Save
Mark Applied
Hide
Staff Engineer, Site Reliability
New York City or Los Angeles or United States or Canada
RemoteFull Time
Babylist
Babylist: Private baby-registry and e-commerce platform helping expecting parents plan, shop, and prepare.
Hands-on Terraform expertise, proven AWS experience (EKS, RDS, networking, CDNs), production Kubernetes operation, CI/CD design, observability and alerting, on-call/incident management, cross-functional developer support, familiarity with AI tooling.
Ruby on Rails, AWS, Sidekiq, MySQL, Redis, Terraform, EKS, RDS, Kubernetes, CircleCI, GitHub Actions, Datadog, Sentry, PagerDuty, Cronitor, Claude, ChatGPT
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Observability
Chicago or New York City
$160k-$200k/yr HybridFull Time
Ripple
Ripple: Enables institutions to move, manage, and tokenize value.
7+ YOE7+ years SRE/Platform experience focused on observability, New Relic, Terraform, PowerShell, Azure/AWS, incident management (Incident.IO/PagerDuty/OpsGenie), and coaching engineering teams.
New Relic, NRQL, Terraform, PowerShell, Azure, AWS, Azure DevOps, Octopus Deploy, Incident.IO, PagerDuty, OpsGenie, Slack, Python, Bash, Jira, SQL Server
1mo
Save
Mark Applied
Hide
Site Reliability Engineer/L3 Support
New York or Kansas or Pennsylvania or New Hampshire
$110k-$120k/yr RemoteFull Time
SS&C Technologies
SS&C TechnologiesNASDAQ: SSNC: Global provider of financial services and healthcare technology solutions.
3+ YOEU.S. citizen with 3+ years SRE/DevOps experience supporting cloud production systems; strong Linux, Kubernetes, AWS, scripting, monitoring, incident management and communication skills.
Kubernetes, AWS, EKS, RDS, IAM, CloudWatch, Route 53, VPC, AWS Backup, Python, Bash, PowerShell, Go, Prometheus, Grafana, Datadog, Splunk, OpenTelemetry, PagerDuty, Jira Service Management, Istio, CI/CD
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer (SRE)
Chicago or New York City
$175k-$220k/yr HybridFull Time
Optimal Market Technologies
Optimal Market Technologies: Private FINRA-registered broker-dealer providing options execution, ATS, routing, and algorithms to retail brokers and institutional trading firms.
Hands-on Linux and network administration, strong scripting/automation, Infrastructure-as-Code experience, production support and incident response ownership, staff management/mentoring, familiarity with C++, Python, SQL, Azure, and real-time trading systems.
C++, Python, SQL, Linux, CentOS7, RHEL9, Microsoft Azure, Microsoft Azure Virtual Desktop (AVD), Claude Code, FIX Protocol, PostgreSQL, AERON
1w
Save
Mark Applied
Hide
Maintenance and Reliability Site Leader
Tonawanda or Buffalo
OnsiteFull Time
DuPont
DuPontNYSE: DD: Global innovation providing technology-based materials and solutions.
7+ YOEBachelor's degree in engineering or equivalent experience, 7–10 years of manufacturing experience, large-site maintenance leadership, and Process Safety Management experience required; union experience preferred.
Microsoft Office, Microsoft Word, Microsoft Excel, Microsoft PowerPoint, Computerized Maintenance Management Software (CMMS), SAP, ultrasound, infrared, vibration
1w
Save
Mark Applied
Hide
Site Reliability Engineer- Team Lead
New York City or Europe or United States or Asia-Pacific
$170k-$190k/yr HybridFull Time
Pico
Pico: Private financial-markets technology providing trading infrastructure, connectivity, market data, software, and analytics to institutions.
Bachelor's degree or higher in engineering or related discipline, operational team leadership experience, Linux performance expertise, networking knowledge, financial technology experience, and programming or scripting skills.
Linux, Python, C, C++, Java
1w
Save
Mark Applied
Hide
Site Reliability Team Leader
Tel Aviv-Yafo or Israel or New York City or London or Edinburgh or Brazil or Estonia or Ukraine
OnsiteFull Time
Optimove
Optimove: Private SaaS providing AI-powered customer engagement and personalized marketing software to consumer brands.
5+ YOE5+ years in SRE, platform, DevOps, or infrastructure engineering; Kubernetes, GCP/AWS, programming, automation, CI/CD, observability, Linux, networking, distributed systems, and strong communication skills.
Kubernetes, GCP, AWS, Python, Go, Bash, CI/CD, Datadog, Prometheus, Grafana, Linux, Terraform, Ansible, Kafka, Pub/Sub, Redis, OpenTelemetry, Canary, Blue/Green, Feature Flags
2mo
Save
Mark Applied
Hide
Site Reliability Engineer, Tech Infrastructure - USDS
New York, New York, United States
$137k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture LLC
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
3+ YOE3+ years SRE/systems engineering experience, bachelor\u0002s in CS or related, proficiency in Python/Go/Java/Shell, Linux, cloud and distributed systems, monitoring and incident management.
Python, Go, Java, Shell, Linux, Docker, Kubernetes, Prometheus, Grafana