29 infrastructure reliability engineer jobs at 28 companies in North Tonawanda, NY

1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Canada or United States or Toronto
RemoteFull Time
BeyondTrust
BeyondTrust: Cybersecurity software providing privileged access management and identity security solutions.
7+ YOERequires 7+ years in SRE, DevOps, or platform engineering, including 2+ years at senior or staff level; expertise in cloud and on-prem infrastructure, Docker, Kubernetes, CI/CD, observability, GitOps, and a systems language.
AWS, Azure, api gateways, service meshes, Terraform, OpenTofu, Ansible, GitOps, Grafana Cloud, Datadog, OpenTelemetry, Docker, Kubernetes, Go, Java, C#, Linux, Windows
4d
Save
Mark Applied
Hide
Staff Platform Site Reliability Engineer
Toronto, Ontario, Canada
HybridFull Time
Index Exchange
Index Exchange: Independent ad-tech supply-side platform helping media owners monetize digital content and enabling brands to buy programmatic advertising.
8+ YOERequires 8+ years in platform engineering, SRE, infrastructure engineering, or DevOps; advanced Linux, Kubernetes, infrastructure-as-code, networking, and Go or Python expertise.
Kubernetes, Terraform, Ansible, GitOps, ArgoCD, Go, Python, Linux, EKS, GKE, Ceph, Hadoop, Spark, HBase, Kafka, Prometheus, Grafana, ELK, Mimir, Loki, Tempo, Vault, AWS, GCP
2w
Save
Mark Applied
Hide
AI Infrastructure Engineer
Toronto, Ontario, Canada
OnsiteFull Time
Palona AI
Palona AI: AI platform helping restaurants capture demand, convert revenue, and manage operations through voice, text, and visual agents.
3+ YOERequires 3+ years of relevant technical experience, distributed-systems expertise, cloud experience, infrastructure automation, production debugging, software development in Python or another modern language, and strong reliability judgment.
Python, Docker, AWS, Microsoft Azure, ECS, Lambda, API Gateway, OpenTofu, Terraform, Datadog, CI/CD
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Toronto, Ontario, Canada
$90k-$133k/yr HybridFull Time
Morningstar
MorningstarNASDAQ: MORN: Provider of independent investment research and financial data.
5+ years in SRE, DevOps, or cloud infrastructure; bachelor's degree or equivalent; AWS, IaC, CI/CD, Docker, Linux, networking, scripting, monitoring, and SRE expertise required.
AWS, Terraform, AWS CDK, CloudFormation, Docker, Amazon ECS, Amazon EKS, Splunk, Amazon CloudWatch, New Relic, Harness, Python, Bash, GitHub Copilot, Claude Code, Amazon EC2, Amazon S3, AWS Lambda, Amazon RDS, Amazon VPC, AWS IAM, Amazon Route 53, Jenkins, GitHub Actions, Linux, Unix, Datadog, AWS SAM, Serverless Framework
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Cloud Networking & Infrastructure as Code)
Waterloo or Toronto or Ottawa
$120k-$170k/yr HybridFull Time
Magnet Forensics
Magnet Forensics: Canadian digital investigation software helping public-safety agencies and enterprises acquire, analyze, and manage digital evidence.
Requires networking or computer science education or equivalent experience, strong AWS networking expertise, multi-account and multi-region architecture experience, IaC, CI/CD, scripting, troubleshooting, and communication skills.
AWS, VPC, Transit Gateway, Route 53, VPN, Terraform, AWS CDK, CloudFormation, CI/CD, Python, Bash, PowerShell, TCP/IP, DNS, ISO 27001, SOC 2, NIST
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Edmonton or Vancouver or Kitchener-Waterloo or Toronto
$146k-$197k/yr HybridFull Time
Jobber
Jobber: Private Canadian SaaS providing quoting, scheduling, invoicing, payments, and customer-management software to home-service businesses.
Senior cloud infrastructure engineer with AWS, Terraform, continuous deployment, programming, incident management, automation, and collaboration experience; Azure, Kubernetes, security, and observability are advantageous.
AWS, Infrastructure-as-Code, Terraform, CircleCI, Azure, Ruby, Python, Bash, Ruby on Rails, GQL, React, Kubernetes, Cloudflare, DNS, AI
3d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Richmond Hill or Toronto or Canada
OnsiteFull Time
Johnson Controls
Johnson ControlsNYSE: JCI: Global leader in smart, healthy, and sustainable building solutions.
7+ YOERequires 7+ years in SRE, production, L3 support, or infrastructure engineering; production Terraform, Azure, AWS, Kubernetes, Datadog, and Grafana experience; Canadian work authorization without sponsorship.
OpenBlue, Airwall, Datadog, Grafana, Terraform, Azure, AWS, Kubernetes, Java, C#, Claude, Microsoft Copilot, Codex, Cursor
3w
Save
Mark Applied
Hide
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Toronto, Ontario, Canada
$113k-$210k/yr HybridFull Time
Manulife
ManulifeTSX / NYSE: MFC: International financial services and insurance provider.
5+ YOE5+ years platform/DevOps experience operating cloud-native distributed systems; experience with Azure, Kubernetes, Terraform/Ansible; on-call and incident response; knowledge of LLM/AI infrastructure and backend services.
Azure, Kubernetes, Terraform, Ansible, Python, Java, Scala, TypeScript, CI/CD, GitOps
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Canada or Toronto
$145k-$193k/yr RemoteFull Time
Penn Interactive
Penn Interactive: PENN Entertainment’s wholly owned digital gaming division serving online betting, casino, and sports-media customers.
5+ YOERequires 5+ years in a similar role, production Kubernetes and Linux experience, cloud and distributed systems expertise, proficiency in at least two of Go, Python, or Bash, and infrastructure project leadership.
Kubernetes, Linux, GCP, AWS, ArgoCD, Helm, GitHub Actions, Datadog, Go, Python, Bash, Shell, Terraform, Istio, Cilium, Ceph, Talos OS, PostgreSQL, PgBouncer
1w
Save
Mark Applied
Hide
Senior DevOps / Blockchain Infrastructure Engineer - Operate
Toronto, Ontario, Canada
$72k-$138k/yr HybridContract, Temporary
Deloitte Canada
Deloitte Canada: Professional services firm providing audit, consulting, tax, and advisory services.
5+ YOERequires 5+ years in DevOps, platform, infrastructure, or site reliability engineering; Kubernetes, Terraform, cloud, Linux, CI/CD, blockchain, monitoring, networking, security, and automation expertise; bachelor's degree required.
Kubernetes, Terraform, Azure, AWS, Google Cloud Platform, Linux, Prometheus, Grafana, Datadog, Splunk, Hyperledger Besu, Quorum, Ethereum, HashiCorp Vault
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
Toronto or Victoria or Ontario or British Columbia
$180k-$233k/yr RemoteFull Time
Thumbtack
Thumbtack: Home services marketplace helping homeowners find and hire local professionals for repairs, maintenance, and improvements.
5+ YOE5+ years managing infrastructure; extensive AWS and Linux fluency; coding in Python, Go, PHP, JavaScript; expertise in DNS, TLS, HTTP/S, TCP/IP; experience operating and observing distributed microservices; rotating on-call.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
2w
Save
Mark Applied
Hide
Head of Commissioning and Reliability
Austin or California or New York or Denver or Calgary or Toronto
$221k-$260k/yr RemoteFull Time
Intersect
Intersect: Energy and data center infrastructure-locating industrial demand with dedicated gas and renewable power.
12+ YOEBachelor's degree in a related engineering field and 12+ years commissioning and reliability experience with high-voltage infrastructure, utility-scale power, or mission-critical industrial facilities.
Isograph, ReliaSoft, SCADA, CMMS, Reliability-Centered Maintenance (RCM)
2w
Save
Mark Applied
Hide
Software Engineer, Infrastructure
New York City or Toronto
$145k-$195k/yr HybridFull Time
Epiq
Epiq: Private legal-services provider helping law firms, corporations, financial institutions, and government agencies manage complex matters.
3+ YOERequires 3+ years in infrastructure, platform, or site reliability engineering; production cloud and Kubernetes experience; Terraform, CI/CD, observability, security, incident response, system design, and Python or Go proficiency.
Terraform, Kubernetes, Docker, Prometheus, Grafana, OpenTelemetry, PostgreSQL, RabbitMQ, Python, AWS, GCP, Azure, GitHub Actions, Azure DevOps
1w
Save
Mark Applied
Hide
Senior Software Engineer - Site Reliability
Toronto, Ontario, Canada
$140k-$180k/yr RemoteFull Time
Funded.club
Funded.club: Fixed-fee startup recruiting agency providing managed headhunting services to startups and scale-ups worldwide.
5+ YOEBachelor's degree in software engineering, computer science, or similar; 5+ years of software development; production distributed-systems experience; automation, Linux, containers, infrastructure as code, observability, Git, testing, code review, and CI.
Go, Python, Rust, C, C++, Shell, YAML, Git, Linux, Terraform, Ansible, Prometheus, Grafana, BIND, PowerDNS, Unbound, DNSSEC, BGP, iptables, nftables, eBPF, OpenVPN, WireGuard, IKEv2, JunOS, VyOS, MySQL, Postgres, Redis, HAProxy, nginx
3d
Save
Mark Applied
Hide
Senior Software Engineer, Infrastructure
United States or Canada or New York City or San Francisco or Boston or Toronto or Chicago or Los Angeles or Washington
$160k-$190k/yr RemoteFull Time
Voltus
Voltus: Privately held virtual power plant operator that connects commercial, industrial, residential, and transportation energy resources to electricity markets.
6+ YOE6+ years of software engineering experience, including DevOps/SRE production operations. Requires Go and/or Python, deep AWS and Kubernetes expertise, Terraform, observability, stateful systems, reliability ownership, and on-call experience.
GitHub, Kubernetes, HashiCorp Nomad, HashiCorp Consul, HashiCorp Vault, AWS, Python, Postgres, Go, FastAPI, Temporal, Delta Lake, ClickHouse, TypeScript, React, Docker, Terraform, Buildkite, Prometheus, Grafana, Elasticsearch, OpenSearch, Argo CD, Flux, Helm, Jenkins, OpenTelemetry, Amazon MSK, AWS Organizations, AWS Control Tower, Auth0, Okta, Amazon Cognito, Keycloak, Java, C++, Claude Code, MCP, AWS Bedrock, GitOps
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Agents
Vancouver or Toronto or London
$140k-$196k/yr HybridFull Time
Klue
Klue: AI-powered competitive enablement and win-loss analysis platform for enterprise revenue teams.
Senior-level software engineering experience building backend and retrieval systems, working with LLM/agentic systems, Python, cloud infrastructure, search/vector DBs, and production reliability/observability.
OpenAI, Anthropic, Gemini, PydanticAI, Logfire, Elasticsearch, Pinecone, PostgreSQL, Docker, Kubernetes, GCP, Temporal, Python, Git, CI/CD, PGVector, AWS, Azure, Copilot, Curssor, Claude Code
3w
Save
Mark Applied
Hide
DevOps SRE
Toronto, Ontario, Canada
$80k-$130k/yr OnsiteFull Time
CGI
CGITSX: GIB.A: Global IT consulting and business services firm.
8+ YOERequires 8+ years in SRE, DevOps, cloud, or infrastructure engineering; Azure, Kubernetes, IaC, DevSecOps, CI/CD, observability, automation, troubleshooting, and AI-assisted development experience.
Microsoft Azure, Kubernetes, Microsoft Visual Studio Code, GitLab, JFrog Artifactory, Nexus, Ansible, Infrastructure as Code (IaC), OpenShift Pipelines, Argo CD, SonarQube, Terraform, Bicep, Azure Resource Manager (ARM), Azure Monitor, Prometheus, Grafana, Elastic, GitHub, Azure DevOps, BMAD
1w
Save
Mark Applied
Hide
Senior AI/ML Engineer
Lisbon or Atlanta or London or San Francisco or Santiago or Sydney or Tokyo or Toronto or Australia or Canada or United States
HybridFull Time
PagerDuty
PagerDutyNYSE: PD: Public software providing AI-powered digital operations and incident management software to business teams.
5+ YOE5+ years of software engineering experience building production distributed systems and AI systems, with expertise in LLMs, agents, retrieval, cloud infrastructure, containers, Kubernetes, reliability, and evaluation.
LLM, Kubernetes, AWS, GCP, Azure, LangChain, LlamaIndex, Kafka, Airflow, Spark
5d
Save
Mark Applied
Hide
AI Evaluation Engineer (QA)
San José or Medellín or Bogotá or Mexico City or Buenos Aires or São Paulo or Toronto
HybridFull Time, Contract
Appnovation
Appnovation: Full-service digital consultancy serving organizations through digital strategy, design, development, and support.
8+ YOEBachelor's degree in Computer Science or Engineering and 8–12 years of relevant experience. Requires enterprise data integration, cloud infrastructure, IaC, CI/CD, data modeling, AI/MLOps, reliability, and architectural leadership expertise.
AWS, Azure, GCP, Terraform, CloudFormation, CI/CD, ETL, ELT, SLOs, SLIs, AI, ML