35 cloud reliability engineer jobs at 24 companies in Newfane, NY

1mo
Save
Mark Applied
Hide
Cloud Performance Engineering - Site Reliability Engineer
Toronto, Ontario, Canada
$110k-$125k/yr RemoteFull Time
Smile Digital Health
Smile Digital Health: Provides software for healthcare data management and interoperability.
Expertise with cloud providers (Azure), performance testing, observability, autoscaling, Kafka tuning, IaC (Terraform/Ansible/Chef), and production Linux operations; strong troubleshooting and security/compliance experience.
FHIR, Otel, Grafana, Prometheus, JMeter, Gatling, Azure Load Testing, Terraform, Ansible, Chef, Kubernetes, OpenShift, Azure Monitor, Application Insights, Log Analytics, Kafka, Java
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
HybridFull Time
iManage
iManage: Intelligent document and email management software for professionals.
Experience in reliability engineering with automation, cloud platforms, observability, and on-call responsibility; strong collaboration and architectural skills.
Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK, EFK, CI/CD, Bash, Python, Java
1w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Buffalo, New York, United States
$140k-$233k/yr OnsiteFull Time
M&T Bank
M&T BankNYSE: MTB: Provides retail, commercial, and wealth management banking services.
7+ YOEExpert in reliability engineering, SLO/SLI frameworks, incident and problem management, observability, automation, cloud platforms, and production operations; 7+ years systems analysis/application development or equivalent.
AWS, Azure, CI/CD, SDLC, SLO/SLI
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer
Toronto or Canada
$100k-$150k/yr RemoteFull Time
SPS Commerce
SPS CommerceNASDAQ: SPSC: Cloud-based supply chain management and retail analytics software provider.
5+ YOE5+ years IT experience (or equivalent), proficiency in Python/Golang, Linux administration, infrastructure-as-code, networking and identity/auth systems, Agile experience, and strong problem-solving and collaboration skills.
Python, Golang, Linux, Kubernetes, ECS, CI/CD, Docker, Istio, Envoy, Consul, Mesos/Marathon, Amazon Web Services, EC2, RDS, Dynamo DB, Route53, Elastic Load Balancers, AMIs, IAM Roles, Ops Works, Cloud Formation
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Warsaw or Toronto
HybridFull Time
SimCorp
SimCorp: Provides integrated software solutions for investment and asset managers.
3+ YOE3+ years in Site Reliability, DevOps, or Cloud Engineering; Azure expertise; IaC with Bicep/ARM/Terraform; monitoring/logging tools; IdP onboarding and security; Kubernetes/Docker; ITIL familiarity.
Microsoft Azure, Terraform, Bicep, ARM, Kubernetes, Docker, DataDog, Log Analytics, Application Insights, OpenTelemetry, Playwright, SAML, OAuth, OIDC, KQL, Defender for Cloud
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Argentina or Toronto or United States
$200k-$230k/yr RemoteFull Time
Domino Data Lab
Domino Data Lab: Enterprise MLOps platform for developing and managing AI models.
Deep SRE/platform engineering experience with Kubernetes, Linux, cloud platforms, observability, Python or Go, incident response, SLO/SLI definition, and mentoring/technical leadership.
Kubernetes, Linux, Python, Go, LLM
4w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
1w
Save
Mark Applied
Hide
Site Reliability Engineer III
Vancouver or Toronto or Edmonton or Victoria
$122k-$171k/yr HybridFull Time
Electronic Arts
Electronic ArtsNASDAQ: EA: Develops and publishes video games and interactive entertainment software.
7+ YOE7+ years experience with cloud, containers, virtualization, Linux, automation and distributed systems; strong scripting/programming in Python, Golang or Java; experience with Terraform, Helm, Chef, Puppet, Packer and Kubernetes.
AWS, Kubernetes, Docker, Terraform, Helm, Chef, Puppet, Packer, Python, Golang, Java, Linux
5d
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Toronto, Ontario, Canada
OnsiteFull Time
Royal Bank of Canada
Royal Bank of CanadaTSX: RY: Provides personal, commercial, and investment banking services worldwide.
5+ YOERequires 5–7 years in site reliability engineering or cloud development, cross-functional leadership, Kubernetes and cloud experience, CI/CD and DevOps knowledge, production support, and proficiency with listed SRE technologies.
Kubernetes, CI/CD, DevOps, Agile Methodology, Python, YAML, Shell, OpenShift, Linux, MongoDB, Dynatrace, Prometheus, PagerDuty, Moog, Splunk, Elastic, Ansible, Grafana, Chaos Engineering, MQ, Kafka, GitHub, Elastic Stack (ELK), Red Hat Ansible, Red Hat OpenShift
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer - Confluent Incident Management & Reliability
Markham or Toronto
$134k-$248k/yr RemoteFull Time
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
10+ YOE10+ years SRE/incident management experience, cloud experience (AWS, GCP, or Azure), deep incident tooling knowledge (Rootly, PagerDuty), Kubernetes and observability expertise, strong communication and coaching skills.
Rootly, PagerDuty, Jira, Confluence, Slack, Kubernetes, AWS, GCP, Azure
2w
Save
Mark Applied
Hide
Advanced Site Reliability / DevOps Engineer
Toronto, Ontario, Canada
$100k-$120k/yr RemoteFull Time
Tech Mahindra
Tech MahindraNational Stock Exchange of India: TECHM: Global provider of information technology and business process services.
7+ YOEMinimum 7 years infrastructure/software engineering experience; strong Azure (AKS, PaaS/IaaS), Terraform, Ansible, Kubernetes, GitHub Actions, CI/CD, scripting (Python), and cloud architecture skills; effective communicator.
Terraform, Ansible, GitHub Actions, Docker, Kubernetes, Azure AKS, Azure Data Factory (ADF), MS Entra Id, Cosmos, Azure SQL MI, Azure SQL DB, Azure Synapse, Databricks, Azure AI/ML, Azure CLI, Python, Git, Rancher, Prometheus, Grafana, Datadog, Dynatrace, Azure Monitor, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (Senior or Staff)
Toronto or New York City or North America
$144k-$200k/yr HybridFull Time
MongoDB
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development and distributed systems experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud platform expertise (AWS, GCP, Azure); Linux and networking knowledge; participate in 24/7 on-call.
Python, Go, Argo Workflows, ArgoCD, Kubernetes, AWS, Google Cloud Platform (GCP), Azure, Linux, CI/CD
6d
Save
Mark Applied
Hide
Senior Principal Site Reliability Engineer
Toronto, Ontario, Canada
$150k-$190k/yr HybridFull Time
Questrade Financial Group
Questrade Financial Group: Online platform for brokerage, wealth management, and banking services.
8+ YOEBachelor's or master's degree or equivalent; 8+ years of software or site reliability engineering experience; multi-language coding, cloud scalability, hybrid architecture, observability, CI/CD, infrastructure-as-code, incident management, and financial systems knowledge.
Java, .NET, Node.js, TypeScript, Python, AWS, Azure, GCP, Prometheus, Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace, Terraform, Ansible, CloudFormation, PagerDuty, Opsgenie, CI/CD
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
$103k-$137k/yr HybridFull Time
Docebo
DoceboNASDAQ: DCBO: Cloud platform for enterprise learning management and training delivery.
4+ YOERequires 4–8 years in SRE, DevOps, or production engineering in SaaS, with Linux, containers, cloud infrastructure and networking, infrastructure-as-code, CI/CD, observability, version control, and incident response experience.
Linux, CI/CD, version control
6d
Save
Mark Applied
Hide
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Toronto, Ontario, Canada
$113k-$210k/yr HybridFull Time
Manulife
ManulifeTSX: MFC: Provides insurance, wealth management, and investment services globally.
5+ YOE5+ years platform/DevOps experience operating cloud-native distributed systems; experience with Azure, Kubernetes, Terraform/Ansible; on-call and incident response; knowledge of LLM/AI infrastructure and backend services.
Azure, Kubernetes, Terraform, Ansible, Python, Java, Scala, TypeScript, CI/CD, GitOps
3w
Save
Mark Applied
Hide
Senior Staff Cloud Security Engineer – Financial Services - Office of the CISO
Toronto, Ontario, Canada
RemoteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
12+ YOE12+ years IT experience with 5+ years in cloud security, expertise in cloud operations and enterprise cyber defense, excellent communication and presentation skills, eligible for Canadian Reliability Status.
1mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability (Scenario Testing) (262286)
Toronto, Ontario, Canada
OnsiteFull Time
Scotiabank
ScotiabankToronto Stock Exchange: BNS: Provides global personal, commercial, and investment banking services.
7+ YOE7+ years in SRE/platform/production engineering, strong incident analysis, proficiency with SQL, Python, Jupyter, PowerBI, cloud (GCP/Azure), Kubernetes, observability and resilience engineering practices; strong leadership and communication.
SQL, Python, Jupyter, PowerBI, ServiceNow, GCP, Azure, Kubernetes
1w
Save
Mark Applied
Hide
Staff Software Engineer, Quality & Reliability
Los Angeles or San Francisco or Toronto or Raleigh or United States
$172k-$229k/yr HybridFull Time
BuildOps
BuildOps: SaaS platform for managing commercial contracting businesses.
Extensive experience solving cross-cutting reliability and quality problems, leading multi-team initiatives, systems thinking, cloud (AWS) experience, strong programming in TypeScript or Java, observability and CI/CD familiarity, and strong communication.
TypeScript, Java, AWS, CI/CD
2w
Save
Mark Applied
Hide
Director, Application Reliability Engineering - Operations
Toronto, Ontario, Canada
OnsiteFull Time
CIBC
CIBCToronto Stock Exchange: CM: Provides personal, commercial, and investment banking and wealth management.
10+ YOE10+ years in application support/incident/problem management with SRE, cloud, and enterprise Tier 1 critical applications; strong leadership, automation, CI/CD, and vendor management experience.
2w
Save
Mark Applied
Hide
Senior SRE, Managed Gateways
Toronto or Canada
$118k-$167k/yr RemoteFull Time
Kong
Kong: Provides cloud-native API management and service mesh platforms.
Experienced SRE with deep Kubernetes and multi-cloud expertise, strong Golang skills, CI/CD and IaC (Terraform/Ansible) experience, monitoring/alerting knowledge, and ability to lead enterprise implementations.
Kubernetes, Golang, CI/CD, Terraform, Ansible, Prometheus, Grafana, ELK stack, Datadog, Istio, Linkerd, PostgreSQL, Cassandra, Kong Konnect