48 platform reliability engineer jobs at 36 companies in North Tonawanda, NY

4d
Save
Mark Applied
Hide
Staff Platform Site Reliability Engineer
Toronto, Ontario, Canada
HybridFull Time
Index Exchange
Index Exchange: Independent ad-tech supply-side platform helping media owners monetize digital content and enabling brands to buy programmatic advertising.
8+ YOERequires 8+ years in platform engineering, SRE, infrastructure engineering, or DevOps; advanced Linux, Kubernetes, infrastructure-as-code, networking, and Go or Python expertise.
Kubernetes, Terraform, Ansible, GitOps, ArgoCD, Go, Python, Linux, EKS, GKE, Ceph, Hadoop, Spark, HBase, Kafka, Prometheus, Grafana, ELK, Mimir, Loki, Tempo, Vault, AWS, GCP
3w
Save
Mark Applied
Hide
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Toronto, Ontario, Canada
$113k-$210k/yr HybridFull Time
Manulife
ManulifeTSX / NYSE: MFC: International financial services and insurance provider.
5+ YOE5+ years platform/DevOps experience operating cloud-native distributed systems; experience with Azure, Kubernetes, Terraform/Ansible; on-call and incident response; knowledge of LLM/AI infrastructure and backend services.
Azure, Kubernetes, Terraform, Ansible, Python, Java, Scala, TypeScript, CI/CD, GitOps
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
HybridFull Time
iManage
iManage: AI knowledge-work software helping legal, accounting, and financial-services organizations manage documents, email, and governed knowledge.
Experience in reliability engineering with automation, cloud platforms, observability, and on-call responsibility; strong collaboration and architectural skills.
Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK, EFK, CI/CD, Bash, Python, Java
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Buffalo, New York, United States
$140k-$233k/yr OnsiteFull Time
M&T Bank
M&T BankNYSE: MTB: A diversified financial services providing banking and wealth management.
7+ YOEExpert in reliability engineering, SLO/SLI frameworks, incident and problem management, observability, automation, cloud platforms, and production operations; 7+ years systems analysis/application development or equivalent.
AWS, Azure, CI/CD, SDLC, SLO/SLI
1mo
Save
Mark Applied
Hide
Staff Software Reliability Engineer - Data Platform
Toronto, Ontario, Canada
$160k-$220k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
5+ YOE5+ years experience building and operating scalable distributed data-platform services; strong OO skills (Java); experience with messaging, data processing, storage systems, reliability, observability, and incident management.
Kinesis, Flink, ElasticSearch, Snowflake, Kafka, Spark, Beam, Databricks, Hadoop, Kubernetes, Mesos, Java, AWS
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Canada or United States or Toronto
RemoteFull Time
BeyondTrust
BeyondTrust: Cybersecurity software providing privileged access management and identity security solutions.
7+ YOERequires 7+ years in SRE, DevOps, or platform engineering, including 2+ years at senior or staff level; expertise in cloud and on-prem infrastructure, Docker, Kubernetes, CI/CD, observability, GitOps, and a systems language.
AWS, Azure, api gateways, service meshes, Terraform, OpenTofu, Ansible, GitOps, Grafana Cloud, Datadog, OpenTelemetry, Docker, Kubernetes, Go, Java, C#, Linux, Windows
2mo
Save
Mark Applied
Hide
Head of Platform Engineering, Reliability & Control
Toronto or Vancouver
$170k-$185k/yr HybridFull Time
Connor, Clark & Lunn Financial Group
Connor, Clark & Lunn Financial Group: Privately owned Canadian asset manager serving individuals, institutional investors and advisors with traditional and alternative investment strategies.
Senior leadership in platform engineering/SRE with experience scaling shared engineering capabilities and establishing reliability standards.
Datadog, Grafana, Prometheus, Elasticsearch, OpenSearch, CI/CD, Infrastructure as Code, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (265177)
Toronto, Ontario, Canada
OnsiteFull Time
Scotiabank
ScotiabankTSX: BNS: Global financial institution providing personal, commercial, and investment banking services.
3+ YOE3+ years experience in ETL platforms and application support, Unix shell scripting, Java, and SQL. Experience with observability tools, incident management, automation, IaC, and on-call rotations. Undergraduate degree in CS or equivalent required.
ETL, iWay, Informatica, Talend, DataStage, Unix Shell Scripting, Java, SQL, Dynatrace, Grafana, Splunk, Infrastructure as Code (IaC)
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Observability
United States or Vancouver or Toronto or Buenos Aires or Brazil or Colombia or Mexico
$129k-$304k/yr RemoteFull Time
Chainlink Labs
Chainlink Labs: Blockchain infrastructure providing open-source oracle and interoperability services to financial institutions, developers, and decentralized applications.
7+ YOE7+ years in DevOps/SRE/platform roles; experience with observability (metrics, logs, traces), Kubernetes, monitoring stacks, real-time systems, and proficiency in one or more languages (C, C++, Java, Python, Go, Perl, Ruby).
OTEL, Prometheus, Grafana, ELK Stack, Splunk, Grafana Stack, AWS, Terraform, Terragrunt, Kubernetes, Calico, ArgoCD, GitHub Actions, Packer, GitOps, C, C++, Java, Python, Go, Perl, Ruby
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
$140k-$182k/yr RemoteFull Time
Movable Ink
Movable Ink: AI-powered marketing software helping marketers personalize customer experiences across email, mobile, and web.
4+ YOE4+ years SRE/Software Engineering experience on cloud platforms, observability and SLO design, IaC and Kubernetes, on-call readiness, strong Linux and programming skills.
AWS, GCP, Prometheus, Thanos, Grafana Alloy, Loki, Tempo, Terraform, Kubernetes, EKS, GKE, NodeJS, Go, Ruby, Python, Shell, Linux
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (Senior or Staff)
Toronto or New York City or North America
$144k-$200k/yr HybridFull Time
MongoDB
MongoDBNASDAQ: MDB: Unified data platform for building modern applications.
6+ YOE6+ years software development and distributed systems experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud platform expertise (AWS, GCP, Azure); Linux and networking knowledge; participate in 24/7 on-call.
Python, Go, Argo Workflows, ArgoCD, Kubernetes, AWS, Google Cloud Platform (GCP), Azure, Linux, CI/CD
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Toronto, Ontario, Canada
$140k-$155k/yr RemoteFull Time
Caseware
Caseware: Private Canadian AI-powered audit and accounting software serving accounting firms, corporations, and government regulators.
8+ YOERequires 8+ years in SRE, platform engineering, DevOps, or related roles; advanced AWS and Kubernetes expertise; Istio, IaC, CI/CD, observability, TypeScript, Node.js, and incident management experience.
AWS, Amazon EKS, AWS IAM, Amazon VPC, AWS Lambda, Amazon CloudFront, Amazon S3, Kubernetes, Istio, AWS CDK, GitHub Actions, AWS CloudWatch, OpenTelemetry, AWS X-Ray, TypeScript, Node.js, Gateway API, mTLS, Certn.co
1w
Save
Mark Applied
Hide
Data Platform Engineer
United States or Canada or San Francisco or Toronto
$190k-$220k/yr RemoteFull Time
Owner.com
Owner.com: AI restaurant software platform helping independent U.S. restaurants grow online sales and manage digital operations.
5+ YOERequires 5+ years in data or analytics engineering, strong SQL and Python, production Snowflake, dbt and modern orchestration experience, plus reliable systems design and cross-functional communication.
Fivetran, Portable, dbt, Hex, Snowflake, Dagster, Metaplane, Select.dev, PostHog, GA4, Hightouch, Census, Sigma, SQL, Python, Airflow, Databricks
5d
Save
Mark Applied
Hide
Staff, Site Reliability Engineer(Global Security)
Toronto or Halifax or Vancouver
OnsiteFull Time
Royal Bank of Canada
Royal Bank of CanadaTSX: RY: Diversified multinational financial services and banking institution.
5+ YOERequires 5+ years in SRE, DevOps, or platform engineering; staff-level technical leadership; production software engineering; highly available systems; CI/CD; Kubernetes; observability; incident response; disaster recovery; AWS or Azure; and strong collaboration.
Terraform, Ansible, Puppet, Kubernetes, Helm, Stonebranch, Jenkins, GitLab CI, GitHub Actions, Docker, Prometheus, Grafana, Dynatrace, ELK, Splunk, SIEM, AWS, Azure, Microsoft Entra, Okta, Idira, SailPoint, Ping, HashiCorp Vault, Python, Go, Java, PowerShell, Bash, OAuth2, OIDC, SAML, LDAP, SCIM, AIOps, ML
2w
Save
Mark Applied
Hide
Senior AI Platform Engineer
Toronto, Ontario, Canada
$126k-$175k/yr HybridFull Time
Epiq
Epiq: Private legal-services provider helping law firms, corporations, financial institutions, and government agencies manage complex matters.
7+ YOE7+ years building production software; strong Python and TypeScript experience; production AI/LLM systems, retrieval-augmented generation, distributed systems, reliability, and incident response experience required.
Python, TypeScript, LangGraph, LangChain, MCP, Langfuse, OpenTelemetry, PostgreSQL, React
1mo
Save
Mark Applied
Hide
Quality Platform Lead
Welland, Ontario, Canada
$105k-$155k/yr OnsiteFull Time
Waukesha
Waukesha: Gas-engine manufacturer serving oil-and-gas compression and distributed-power operators under INNIO’s Waukesha brand.
3+ YOEBachelor's in engineering or related field preferred, 3+ years relevant experience (5+ preferred), Lean Six Sigma, RCA and reliability methods, strong data analysis and stakeholder skills.
Microsoft Office, Microsoft Excel, Microsoft PowerPoint, Creo
1mo
Save
Mark Applied
Hide
Lead, Site Reliability Engineering (Application Support)
Toronto, Ontario, Canada
$86k-$130k/yr HybridFull Time
OMERS
OMERS: Canadian jointly sponsored defined-benefit pension plan administering retirement income for Ontario municipal and other public-sector employees.
5+ YOE5+ years SRE/Platform/DevOps experience with strong Azure, incident response, CI/CD (GitHub Actions), observability (Datadog/Azure Monitor/Log Analytics), container and networking knowledge, and scripting (PowerShell/Bash/Python).
Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), Azure Functions, GitHub Actions, Datadog, Azure Monitor, Log Analytics, PowerShell, Bash, Azure CLI, Python, Kubernetes
3d
Save
Mark Applied
Hide
Senior Backend Software Engineer, Integrations Platform [Canada]
Canada or Toronto or San Francisco or New York City or London or Dublin or Tel Aviv or Sydney
RemoteFull Time
Vanta
Vanta: Private software that helps businesses automate compliance, manage risk, and prove security trust.
8+ YOE8+ years building reliable, scalable backend systems; expertise in RESTful APIs, GraphQL, authentication, data integrations, security, and platform design; experience leading complex projects and mentoring engineers.
TypeScript, React, Node.js, CI/CD, RESTful APIs, GraphQL, OAuth, JWT, AI Skills
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Storage Platform
Bellevue or Menlo Park or Toronto
$230k-$270k/yr HybridFull Time
Robinhood
RobinhoodNASDAQ: HOOD: Provides brokerage, crypto, advisory, and banking services.
5+ YOEDeep expertise in PostgreSQL/Aurora, distributed systems (sharding, replication, transactions), proficiency in Go or Rust, experience with Kubernetes and AWS services, and strong reliability/performance engineering skills.
PostgreSQL, Aurora PostgreSQL, Go, Rust, Kubernetes, RDS, DynamoDB