148 cloud reliability engineer jobs at 88 companies in Lake Success, NY

1mo
Save
Mark Applied
Hide
Cloud Reliability Engineer
Englewood Cliffs or New York City
$135k-$165k/yr HybridFull Time
Versant Media Group
Versant Media GroupNasdaq: VSNT: Media and entertainment managing diverse content and digital brands.
3+ YOEBachelor's degree or equivalent experience; 3–7 years in SRE/Cloud/DevOps roles; strong AWS experience (enterprise scale), Terraform and CloudFormation, scripting (Python/PowerShell/Bash), CI/CD, monitoring/observability, incident management.
AWS, AWS Organizations, Control Tower, Identity Center, Terraform, CloudFormation, Python, PowerShell, Bash, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
GCP, Azure, Terraform, Terraform Enterprise, Log Analytics, Dynatrace, Resource Graph, CI/CD, IAM
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York, New York, United States
HybridFull Time
Mistral AI
Mistral AI: Developer of open-weight and frontier AI models.
7+ YOE7+ years SRE/DevOps experience, Master’s in CS/Engineering or related, strong cloud and distributed systems skills, Kubernetes/CI-CD/infra-as-code proficiency, scripting experience, observability and on-call experience.
Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast
4w
Save
Mark Applied
Hide
Staff , Site Reliability Engineer - Cloud Platform
New York City, New York, United States
$190k-$210k/yr HybridFull Time
Butterfly Network
Butterfly NetworkNew York Stock Exchange: BFLY: Public U.S. medical technology making handheld point-of-care ultrasound hardware and AI-powered clinical software for healthcare professionals.
8+ YOE8+ years managing production systems; deep hands-on AWS and Kubernetes experience; strong programming/scripting and automation skills; observability platform ownership; incident response leadership and mentoring ability.
Kubernetes, EKS, AWS, NewRelic, Datadog, DICOM, HL7, FHIR, PACS, VNA, EMR
2w
Save
Mark Applied
Hide
Platform / Site Reliability Engineer
New York City, New York, United States
OnsiteFull Time
Sunset
Sunset: Private startup wind-down service helping founders close companies through legal, tax, and operational work.
Production cloud infrastructure and reliability experience across multiple services, strong software engineering skills, infrastructure and application coding, incident leadership, recovery expertise, and AI engineering tool proficiency.
AWS, Terraform, CI/CD, SOC 2
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Terraform, Terraform Enterprise, Google Cloud Platform (GCP), Azure, Log Analytics, Dynatrace, Resource Graph, CI/CD
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: AI-powered marketing software helping marketers personalize customer experiences across email, mobile, and web.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Greenwich, Connecticut, United States
HybridFull Time
Interactive Brokers
Interactive BrokersNASDAQ: IBKR: Global electronic brokerage firm providing automated trading technology.
5+ YOE5+ years experience in Linux/Unix, networking and coding; experience with cloud (AWS or Azure), Terraform or CloudFormation, Docker and Kubernetes; bachelor's or master's in CS/STEM; CI/CD, on-call rotation, mentoring skills.
CI/CD, Terraform, CloudFormation, AWS, Azure, Docker, Kubernetes, Linux, Unix
1mo
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered warehouse intelligence serving logistics operators with AI agents for monitoring, forecasting, and automation.
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York, New York, United States
$178k-$258k/yr RemoteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
Requires a computer science bachelor's degree or equivalent experience, Python, production ML inference, AWS cloud infrastructure, Kubernetes, vulnerability management, distributed-systems debugging, and on-call participation.
Python, PHP, Node.js, Ruby, SageMaker, OpenAI, Bedrock, EC2 Auto Scaling Groups, Kubernetes, AWS, Azure, GCP, AMI, LangGraph, LLM gateway, MCP, Aurora PostgreSQL, Memcached, Terraform, Terragrunt, Atlantis, Chef, Ansible, SSM, Docker, bash, Jenkins, Argo CD, New Relic, Splunk, Grafana, Prometheus, Fastly, Datadome, WAF, vLLM, LangSmith, Lambda
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yr HybridFull Time
ThousandEyes
ThousandEyesNASDAQ: CSCO: Global leader in networking, cybersecurity, and cloud-native technology solutions.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
2mo
Save
Mark Applied
Hide
Site Reliability Engineer, Tech Infrastructure - USDS
New York, New York, United States
$137k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture LLC
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
3+ YOE3+ years SRE/systems engineering experience, bachelor\u0002s in CS or related, proficiency in Python/Go/Java/Shell, Linux, cloud and distributed systems, monitoring and incident management.
Python, Go, Java, Shell, Linux, Docker, Kubernetes, Prometheus, Grafana
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Cloud, Observability & Automation)
Jersey City, New Jersey, United States
HybridFull Time
DTCC
DTCC: Global post-trade market infrastructure for the financial services industry.
8+ YOEBachelor's degree in computer science, engineering, or equivalent experience; 8+ years in SRE, production engineering, DevOps, or application support; AWS, Python, Java, Go, Linux, monitoring, incident management, and distributed systems expertise.
AWS, Splunk, Grafana, Dynatrace, ITSI, Python, Java, Amazon Q, Kiro, Go, Linux/Unix
3w
Save
Mark Applied
Hide
Service Reliability Engineer
London or Manchester or New York City
HybridFull Time
Fitch Group
Fitch Group: Global provider of financial information, credit ratings, and analytics.
Deep SRE, DevOps, or platform engineering experience with AWS, Azure, Docker, Kubernetes, Linux, Windows, CI/CD, cloud security, networking, and Python, PowerShell, or Bash.
AWS, Azure, Docker, Kubernetes, Linux, Windows, IIS, .NET, Java Spring Boot, GitHub Actions, Bamboo, Python, PowerShell, Bash, Datadog, Microsoft Teams, AWS Bedrock, SageMaker, Model Context Protocol (MCP), IAM, OPA, AWS Config, AWS CloudTrail, AWS Security Hub, Wiz, DNS, CIS, NIST, ISO 27001
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
New York City, New York, United States
OnsiteContract
Bahwan CyberTek
Bahwan CyberTek: Global digital transformation providing predictive analytics, AI, and digital supply chain solutions to enterprise clients.
8+ YOERequires 8–12 years of experience with SRE/DevOps, observability, automation, cloud platforms, ETL/ELT, CI/CD, reliability engineering, and large datasets; Python, PowerShell, and Bash skills required.
Python, PowerShell, Bash, Azure DataBricks, Unity Catalog, AWS S3, AWS RDS, ETL, ELT, CI/CD
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer (FedRAMP)
Bellevue or Chicago or New York City or San Francisco or Washington
$174k-$267k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
8+ YOE8+ years operations experience in cloud and Linux, strong networking and web server knowledge, proficiency with Terraform/Chef and scripting (Bash, Python, Go), experience with automation tools and on-call duty.
AWS, Terraform, Chef, Ansible, Puppet, Apache httpd, nginx, Apache Tomcat, Bash, Python, Golang, git, gdb, strace, ltrace, tcpdump, Wireshark, Docker, Kubernetes
3w
Save
Mark Applied
Hide
Site Reliability Engineer, Global Banking & Markets, Vice President
New York City, New York, United States
$150k-$250k/yr OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities and investment management firm.
8+ YOE8+ years reliability/software engineering experience, strong Java (Java 17+) skills, cloud (GCP/AWS), Kubernetes/Docker, SLO/SLI incident experience, AI-assisted tooling familiarity, excellent communication and risk acumen.
Java 17, Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist, Apache Kafka, GCP, AWS, Kubernetes, Docker, Terraform, Helm, Prometheus, Grafana, OpenTelemetry, Spring Boot, gRPC, Protocol Buffers, Apache Camel, Spring Integration, Vert.x, Netty
1mo
Save
Mark Applied
Hide
Cloud AWS Support Reliability Engineer (SRE)
New York, New York, United States
$160k-$230k/yr OnsiteFull Time
Federal Reserve System
Federal Reserve System: The central bank of the United States.
Proven Terraform and GitLab CI/CD proficiency for AWS provisioning and deployment, strong AWS infrastructure knowledge, SRE/DevOps experience, incident remediation and on-call support, programming experience (Java/Python/Scala).
Terraform, GitLab, AWS Lambda, S3, SNS, EKS, Helm, EC2, RDS, Oracle-RDS, VPC, transit gateway, Direct Connect, ALB, ELB, Grafana, CloudWatch, Route53, DNS, EFS, Glue, ECS, CloudBeaver, Java, Python, Scala, NGINX, JBoss, Tomcat, JIRA, Okta
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer, Vice President
New York City, New York, United States
$150k-$190k/yr OnsiteFull Time
Morgan Stanley
Morgan StanleyNYSE: MS: Global financial services firm providing investment and banking solutions.
10+ YOE10+ years production experience, strong scripting (Python/Perl/Shell), DB and batch scheduler experience, CI/CD, cloud (Azure/AWS), monitoring and production support expertise.
Python, Perl, Shell, Ruby, Java, C#, DB2, Sybase, Oracle, Autosys, Jenkins, Train, Windeploy, Splunk, IP Soft, Sockeye, Azure, AWS