35 cloud reliability engineer jobs at 24 companies in Newfane, NY
1mo
Save
Mark Applied
Hide
1mo
Cloud Performance Engineering - Site Reliability Engineer
Toronto, Ontario, Canada
$110k-$125k/yrRemoteFull Time
Smile Digital Health: Provides software for healthcare data management and interoperability.
Expertise with cloud providers (Azure), performance testing, observability, autoscaling, Kafka tuning, IaC (Terraform/Ansible/Chef), and production Linux operations; strong troubleshooting and security/compliance experience.
iManage: Intelligent document and email management software for professionals.
Experience in reliability engineering with automation, cloud platforms, observability, and on-call responsibility; strong collaboration and architectural skills.
7+ YOEExpert in reliability engineering, SLO/SLI frameworks, incident and problem management, observability, automation, cloud platforms, and production operations; 7+ years systems analysis/application development or equivalent.
5+ YOE5+ years IT experience (or equivalent), proficiency in Python/Golang, Linux administration, infrastructure-as-code, networking and identity/auth systems, Agile experience, and strong problem-solving and collaboration skills.
Python, Golang, Linux, Kubernetes, ECS, CI/CD, Docker, Istio, Envoy, Consul, Mesos/Marathon, Amazon Web Services, EC2, RDS, Dynamo DB, Route53, Elastic Load Balancers, AMIs, IAM Roles, Ops Works, Cloud Formation
SimCorp: Provides integrated software solutions for investment and asset managers.
3+ YOE3+ years in Site Reliability, DevOps, or Cloud Engineering; Azure expertise; IaC with Bicep/ARM/Terraform; monitoring/logging tools; IdP onboarding and security; Kubernetes/Docker; ITIL familiarity.
Domino Data Lab: Enterprise MLOps platform for developing and managing AI models.
Deep SRE/platform engineering experience with Kubernetes, Linux, cloud platforms, observability, Python or Go, incident response, SLO/SLI definition, and mentoring/technical leadership.
Electronic ArtsNASDAQ: EA: Develops and publishes video games and interactive entertainment software.
7+ YOE7+ years experience with cloud, containers, virtualization, Linux, automation and distributed systems; strong scripting/programming in Python, Golang or Java; experience with Terraform, Helm, Chef, Puppet, Packer and Kubernetes.
Royal Bank of CanadaTSX: RY: Provides personal, commercial, and investment banking services worldwide.
5+ YOERequires 5–7 years in site reliability engineering or cloud development, cross-functional leadership, Kubernetes and cloud experience, CI/CD and DevOps knowledge, production support, and proficiency with listed SRE technologies.
Kubernetes, CI/CD, DevOps, Agile Methodology, Python, YAML, Shell, OpenShift, Linux, MongoDB, Dynatrace, Prometheus, PagerDuty, Moog, Splunk, Elastic, Ansible, Grafana, Chaos Engineering, MQ, Kafka, GitHub, Elastic Stack (ELK), Red Hat Ansible, Red Hat OpenShift
Staff Site Reliability Engineer - Confluent Incident Management & Reliability
Markham or Toronto
$134k-$248k/yrRemoteFull Time
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
10+ YOE10+ years SRE/incident management experience, cloud experience (AWS, GCP, or Azure), deep incident tooling knowledge (Rootly, PagerDuty), Kubernetes and observability expertise, strong communication and coaching skills.
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development and distributed systems experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud platform expertise (AWS, GCP, Azure); Linux and networking knowledge; participate in 24/7 on-call.
Questrade Financial Group: Online platform for brokerage, wealth management, and banking services.
8+ YOEBachelor's or master's degree or equivalent; 8+ years of software or site reliability engineering experience; multi-language coding, cloud scalability, hybrid architecture, observability, CI/CD, infrastructure-as-code, incident management, and financial systems knowledge.
DoceboNASDAQ: DCBO: Cloud platform for enterprise learning management and training delivery.
4+ YOERequires 4–8 years in SRE, DevOps, or production engineering in SaaS, with Linux, containers, cloud infrastructure and networking, infrastructure-as-code, CI/CD, observability, version control, and incident response experience.
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Toronto, Ontario, Canada
$113k-$210k/yrHybridFull Time
ManulifeTSX: MFC: Provides insurance, wealth management, and investment services globally.
5+ YOE5+ years platform/DevOps experience operating cloud-native distributed systems; experience with Azure, Kubernetes, Terraform/Ansible; on-call and incident response; knowledge of LLM/AI infrastructure and backend services.
Senior Staff Cloud Security Engineer – Financial Services - Office of the CISO
Toronto, Ontario, Canada
RemoteFull Time
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
12+ YOE12+ years IT experience with 5+ years in cloud security, expertise in cloud operations and enterprise cyber defense, excellent communication and presentation skills, eligible for Canadian Reliability Status.
Los Angeles or San Francisco or Toronto or Raleigh or United States
$172k-$229k/yrHybridFull Time
BuildOps: SaaS platform for managing commercial contracting businesses.
Extensive experience solving cross-cutting reliability and quality problems, leading multi-team initiatives, systems thinking, cloud (AWS) experience, strong programming in TypeScript or Java, observability and CI/CD familiarity, and strong communication.
CIBCToronto Stock Exchange: CM: Provides personal, commercial, and investment banking and wealth management.
10+ YOE10+ years in application support/incident/problem management with SRE, cloud, and enterprise Tier 1 critical applications; strong leadership, automation, CI/CD, and vendor management experience.
Kong: Provides cloud-native API management and service mesh platforms.
Experienced SRE with deep Kubernetes and multi-cloud expertise, strong Golang skills, CI/CD and IaC (Terraform/Ansible) experience, monitoring/alerting knowledge, and ability to lead enterprise implementations.