32 cloud reliability engineer jobs at 24 companies in Lockport, NY

3w
Save
Mark Applied
Hide
Cloud Performance Engineering - Site Reliability Engineer
Toronto, Ontario, Canada
$110k-$125k/yr RemoteFull Time
Smile Digital Health
Smile Digital Health: Provides software for healthcare data management and interoperability.
Expertise with cloud providers (Azure), performance testing, observability, autoscaling, Kafka tuning, IaC (Terraform/Ansible/Chef), and production Linux operations; strong troubleshooting and security/compliance experience.
FHIR, Otel, Grafana, Prometheus, JMeter, Gatling, Azure Load Testing, Terraform, Ansible, Chef, Kubernetes, OpenShift, Azure Monitor, Application Insights, Log Analytics, Kafka, Java
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
HybridFull Time
iManage
iManage: Intelligent document and email management software for professionals.
Experience in reliability engineering with automation, cloud platforms, observability, and on-call responsibility; strong collaboration and architectural skills.
Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK, EFK, CI/CD, Bash, Python, Java
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer
Toronto or Canada
$100k-$150k/yr RemoteFull Time
SPS Commerce
SPS CommerceNASDAQ: SPSC: Cloud-based supply chain management and retail analytics software provider.
5+ YOE5+ years IT experience (or equivalent), proficiency in Python/Golang, Linux administration, infrastructure-as-code, networking and identity/auth systems, Agile experience, and strong problem-solving and collaboration skills.
Python, Golang, Linux, Kubernetes, ECS, CI/CD, Docker, Istio, Envoy, Consul, Mesos/Marathon, Amazon Web Services, EC2, RDS, Dynamo DB, Route53, Elastic Load Balancers, AMIs, IAM Roles, Ops Works, Cloud Formation
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Warsaw or Toronto
HybridFull Time
SimCorp
SimCorp: Provides integrated software solutions for investment and asset managers.
3+ YOE3+ years in Site Reliability, DevOps, or Cloud Engineering; Azure expertise; IaC with Bicep/ARM/Terraform; monitoring/logging tools; IdP onboarding and security; Kubernetes/Docker; ITIL familiarity.
Microsoft Azure, Terraform, Bicep, ARM, Kubernetes, Docker, DataDog, Log Analytics, Application Insights, OpenTelemetry, Playwright, SAML, OAuth, OIDC, KQL, Defender for Cloud
2w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Argentina or Toronto or United States
$200k-$230k/yr RemoteFull Time
Domino Data Lab
Domino Data Lab: Enterprise MLOps platform for developing and managing AI models.
Deep SRE/platform engineering experience with Kubernetes, Linux, cloud platforms, observability, Python or Go, incident response, SLO/SLI definition, and mentoring/technical leadership.
Kubernetes, Linux, Python, Go, LLM
3w
Save
Mark Applied
Hide
Site Reliability Engineer IV
Buffalo, New York, United States
$140k-$233k/yr HybridFull Time
M&T Bank
M&T BankNYSE: MTB: Provides retail, commercial, and wealth management banking services.
7+ YOEExpert-level SRE experience, 7+ years systems analysis/application development (or 9+ with associate), advanced scripting, observability, IaC (Terraform), and cloud (Azure) experience.
OpenTelemetry (OTel), Dynatrace, Terraform, Microsoft Azure, Azure Monitor, Application Insights, PowerShell, Python, Bash
5d
Save
Mark Applied
Hide
Site Reliability Engineer III
Vancouver or Toronto or Edmonton or Victoria
$122k-$171k/yr HybridFull Time
Electronic Arts
Electronic ArtsNASDAQ: EA: Develops and publishes video games and interactive entertainment software.
7+ YOE7+ years experience with cloud, containers, virtualization, Linux, automation and distributed systems; strong scripting/programming in Python, Golang or Java; experience with Terraform, Helm, Chef, Puppet, Packer and Kubernetes.
AWS, Kubernetes, Docker, Terraform, Helm, Chef, Puppet, Packer, Python, Golang, Java, Linux
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer - Confluent Incident Management & Reliability
Markham or Toronto
$134k-$248k/yr RemoteFull Time
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
10+ YOE10+ years SRE/incident management experience, cloud experience (AWS, GCP, or Azure), deep incident tooling knowledge (Rootly, PagerDuty), Kubernetes and observability expertise, strong communication and coaching skills.
Rootly, PagerDuty, Jira, Confluence, Slack, Kubernetes, AWS, GCP, Azure
1w
Save
Mark Applied
Hide
Advanced Site Reliability / DevOps Engineer
Toronto, Ontario, Canada
$100k-$120k/yr RemoteFull Time
Tech Mahindra
Tech MahindraNational Stock Exchange of India: TECHM: Global provider of information technology and business process services.
7+ YOEMinimum 7 years infrastructure/software engineering experience; strong Azure (AKS, PaaS/IaaS), Terraform, Ansible, Kubernetes, GitHub Actions, CI/CD, scripting (Python), and cloud architecture skills; effective communicator.
Terraform, Ansible, GitHub Actions, Docker, Kubernetes, Azure AKS, Azure Data Factory (ADF), MS Entra Id, Cosmos, Azure SQL MI, Azure SQL DB, Azure Synapse, Databricks, Azure AI/ML, Azure CLI, Python, Git, Rancher, Prometheus, Grafana, Datadog, Dynatrace, Azure Monitor, Linux
3d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE), Cloud Operations
Toronto, Ontario, Canada
OnsiteFull Time
Royal Bank of Canada
Royal Bank of CanadaTSX: RY: Provides personal, commercial, and investment banking services worldwide.
5+ YOE5+ years SRE/DevOps experience, Kubernetes/OpenShift administration, Ansible and Terraform proficiency, Python scripting, monitoring/observability (Prometheus,Grafana,ELK), incident management and on-call experience.
Kubernetes, OpenShift, ECE, Confluent Kafka, Ansible, Ansible Tower, Terraform, Python, Prometheus, Grafana, ELK, Dynatrace, PagerDuty, ServiceNow, AWS, Azure, GCP, Linux, Red Hat OS
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (Senior or Staff)
Toronto or New York City or North America
$144k-$200k/yr HybridFull Time
MongoDB
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development and distributed systems experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud platform expertise (AWS, GCP, Azure); Linux and networking knowledge; participate in 24/7 on-call.
Python, Go, Argo Workflows, ArgoCD, Kubernetes, AWS, Google Cloud Platform (GCP), Azure, Linux, CI/CD
1w
Save
Mark Applied
Hide
Senior Platform Reliability Engineer
Toronto or Waterloo
$113k-$163k/yr HybridFull Time
Manulife
ManulifeTSX: MFC: Provides insurance, wealth management, and investment services globally.
Hands-on experience with Azure, AKS, Kubernetes, Terraform, Helm, Flux, GitHub Actions and cloud-native production platform support; strong troubleshooting, automation and observability skills.
Azure, AKS, APIM, Redis, Key Vault, Data Factory, Functions, GitHub, GitHub Actions, Terraform, Helm, Flux, New Relic, Azure Monitor, ADX, PostgreSQL, MongoDB, Cosmos DB, Azure SQL, SQL MI, AI Foundry
1w
Save
Mark Applied
Hide
Senior Staff Cloud Security Engineer – Financial Services - Office of the CISO
Toronto, Ontario, Canada
RemoteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
12+ YOE12+ years IT experience with 5+ years in cloud security, expertise in cloud operations and enterprise cyber defense, excellent communication and presentation skills, eligible for Canadian Reliability Status.
3w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
Toronto, Ontario, Canada
$123k-$165k/yr HybridFull Time
Docebo
DoceboNASDAQ: DCBO: Cloud platform for enterprise learning management and training delivery.
6+ YOE6+ years SRE/DevOps experience with leadership in incident response, cloud (AWS preferred), monitoring/observability, hiring/mentoring, and process ownership.
AWS
1mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability (Scenario Testing) (262286)
Toronto, Ontario, Canada
OnsiteFull Time
Scotiabank
ScotiabankToronto Stock Exchange: BNS: Provides global personal, commercial, and investment banking services.
7+ YOE7+ years in SRE/platform/production engineering, strong incident analysis, proficiency with SQL, Python, Jupyter, PowerBI, cloud (GCP/Azure), Kubernetes, observability and resilience engineering practices; strong leadership and communication.
SQL, Python, Jupyter, PowerBI, ServiceNow, GCP, Azure, Kubernetes
3h
Save
Mark Applied
Hide
Staff Software Engineer, Quality & Reliability
Los Angeles or San Francisco or Toronto or Raleigh or United States
$172k-$229k/yr HybridFull Time
BuildOps
BuildOps: SaaS platform for managing commercial contracting businesses.
Extensive experience solving cross-cutting reliability and quality problems, leading multi-team initiatives, systems thinking, cloud (AWS) experience, strong programming in TypeScript or Java, observability and CI/CD familiarity, and strong communication.
TypeScript, Java, AWS, CI/CD
6d
Save
Mark Applied
Hide
Director, Application Reliability Engineering - Operations
Toronto, Ontario, Canada
OnsiteFull Time
CIBC
CIBCToronto Stock Exchange: CM: Provides personal, commercial, and investment banking and wealth management.
10+ YOE10+ years in application support/incident/problem management with SRE, cloud, and enterprise Tier 1 critical applications; strong leadership, automation, CI/CD, and vendor management experience.
1w
Save
Mark Applied
Hide
Senior SRE, Managed Gateways
Toronto or Canada
$118k-$167k/yr RemoteFull Time
Kong
Kong: Provides cloud-native API management and service mesh platforms.
Experienced SRE with deep Kubernetes and multi-cloud expertise, strong Golang skills, CI/CD and IaC (Terraform/Ansible) experience, monitoring/alerting knowledge, and ability to lead enterprise implementations.
Kubernetes, Golang, CI/CD, Terraform, Ansible, Prometheus, Grafana, ELK stack, Datadog, Istio, Linkerd, PostgreSQL, Cassandra, Kong Konnect
2mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering and DevOps
Toronto, Ontario, Canada
$142k-$186k/yr HybridFull Time
Moneris
Moneris: Processes merchant payments and provides point-of-sale commerce solutions.
8+ YOE3+ Mgmt8+ years senior distributed systems; 3+ years leading technical teams; SRE principles; cloud (Azure) and Kubernetes; IaC; automation; observability tools (Dynatrace, Datadog, New Relic, AppDynamics); Linux production; scripting
Azure, Kubernetes, Infrastructure as Code, Dynatrace, Datadog, New Relic, AppDynamics, Scripting