29 infrastructure reliability engineer jobs at 28 companies in North Tonawanda, NY
1mo
Save
Mark Applied
Hide
1mo
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
7+ YOERequires 7+ years in SRE, DevOps, or platform engineering, including 2+ years at senior or staff level; expertise in cloud and on-prem infrastructure, Docker, Kubernetes, CI/CD, observability, GitOps, and a systems language.
AWS, Azure, api gateways, service meshes, Terraform, OpenTofu, Ansible, GitOps, Grafana Cloud, Datadog, OpenTelemetry, Docker, Kubernetes, Go, Java, C#, Linux, Windows
Index Exchange: Independent ad-tech supply-side platform helping media owners monetize digital content and enabling brands to buy programmatic advertising.
8+ YOERequires 8+ years in platform engineering, SRE, infrastructure engineering, or DevOps; advanced Linux, Kubernetes, infrastructure-as-code, networking, and Go or Python expertise.
Palona AI: AI platform helping restaurants capture demand, convert revenue, and manage operations through voice, text, and visual agents.
3+ YOERequires 3+ years of relevant technical experience, distributed-systems expertise, cloud experience, infrastructure automation, production debugging, software development in Python or another modern language, and strong reliability judgment.
Python, Docker, AWS, Microsoft Azure, ECS, Lambda, API Gateway, OpenTofu, Terraform, Datadog, CI/CD
MorningstarNASDAQ: MORN: Provider of independent investment research and financial data.
5+ years in SRE, DevOps, or cloud infrastructure; bachelor's degree or equivalent; AWS, IaC, CI/CD, Docker, Linux, networking, scripting, monitoring, and SRE expertise required.
Senior Site Reliability Engineer (Cloud Networking & Infrastructure as Code)
Waterloo or Toronto or Ottawa
$120k-$170k/yrHybridFull Time
Magnet Forensics: Canadian digital investigation software helping public-safety agencies and enterprises acquire, analyze, and manage digital evidence.
Requires networking or computer science education or equivalent experience, strong AWS networking expertise, multi-account and multi-region architecture experience, IaC, CI/CD, scripting, troubleshooting, and communication skills.
Johnson ControlsNYSE: JCI: Global leader in smart, healthy, and sustainable building solutions.
7+ YOERequires 7+ years in SRE, production, L3 support, or infrastructure engineering; production Terraform, Azure, AWS, Kubernetes, Datadog, and Grafana experience; Canadian work authorization without sponsorship.
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Toronto, Ontario, Canada
$113k-$210k/yrHybridFull Time
ManulifeTSX / NYSE: MFC: International financial services and insurance provider.
5+ YOE5+ years platform/DevOps experience operating cloud-native distributed systems; experience with Azure, Kubernetes, Terraform/Ansible; on-call and incident response; knowledge of LLM/AI infrastructure and backend services.
Penn Interactive: PENN Entertainment’s wholly owned digital gaming division serving online betting, casino, and sports-media customers.
5+ YOERequires 5+ years in a similar role, production Kubernetes and Linux experience, cloud and distributed systems expertise, proficiency in at least two of Go, Python, or Bash, and infrastructure project leadership.
Senior Software Engineer, Site Reliability Engineering
Toronto or Victoria or Ontario or British Columbia
$180k-$233k/yrRemoteFull Time
Thumbtack: Home services marketplace helping homeowners find and hire local professionals for repairs, maintenance, and improvements.
5+ YOE5+ years managing infrastructure; extensive AWS and Linux fluency; coding in Python, Go, PHP, JavaScript; expertise in DNS, TLS, HTTP/S, TCP/IP; experience operating and observing distributed microservices; rotating on-call.
Austin or California or New York or Denver or Calgary or Toronto
$221k-$260k/yrRemoteFull Time
Intersect: Energy and data center infrastructure-locating industrial demand with dedicated gas and renewable power.
12+ YOEBachelor's degree in a related engineering field and 12+ years commissioning and reliability experience with high-voltage infrastructure, utility-scale power, or mission-critical industrial facilities.
Epiq: Private legal-services provider helping law firms, corporations, financial institutions, and government agencies manage complex matters.
3+ YOERequires 3+ years in infrastructure, platform, or site reliability engineering; production cloud and Kubernetes experience; Terraform, CI/CD, observability, security, incident response, system design, and Python or Go proficiency.
Funded.club: Fixed-fee startup recruiting agency providing managed headhunting services to startups and scale-ups worldwide.
5+ YOEBachelor's degree in software engineering, computer science, or similar; 5+ years of software development; production distributed-systems experience; automation, Linux, containers, infrastructure as code, observability, Git, testing, code review, and CI.
United States or Canada or New York City or San Francisco or Boston or Toronto or Chicago or Los Angeles or Washington
$160k-$190k/yrRemoteFull Time
Voltus: Privately held virtual power plant operator that connects commercial, industrial, residential, and transportation energy resources to electricity markets.
6+ YOE6+ years of software engineering experience, including DevOps/SRE production operations. Requires Go and/or Python, deep AWS and Kubernetes expertise, Terraform, observability, stateful systems, reliability ownership, and on-call experience.
Klue: AI-powered competitive enablement and win-loss analysis platform for enterprise revenue teams.
Senior-level software engineering experience building backend and retrieval systems, working with LLM/agentic systems, Python, cloud infrastructure, search/vector DBs, and production reliability/observability.
CGITSX: GIB.A: Global IT consulting and business services firm.
8+ YOERequires 8+ years in SRE, DevOps, cloud, or infrastructure engineering; Azure, Kubernetes, IaC, DevSecOps, CI/CD, observability, automation, troubleshooting, and AI-assisted development experience.
Microsoft Azure, Kubernetes, Microsoft Visual Studio Code, GitLab, JFrog Artifactory, Nexus, Ansible, Infrastructure as Code (IaC), OpenShift Pipelines, Argo CD, SonarQube, Terraform, Bicep, Azure Resource Manager (ARM), Azure Monitor, Prometheus, Grafana, Elastic, GitHub, Azure DevOps, BMAD
Lisbon or Atlanta or London or San Francisco or Santiago or Sydney or Tokyo or Toronto or Australia or Canada or United States
HybridFull Time
PagerDutyNYSE: PD: Public software providing AI-powered digital operations and incident management software to business teams.
5+ YOE5+ years of software engineering experience building production distributed systems and AI systems, with expertise in LLMs, agents, retrieval, cloud infrastructure, containers, Kubernetes, reliability, and evaluation.
San José or Medellín or Bogotá or Mexico City or Buenos Aires or São Paulo or Toronto
HybridFull Time, Contract
Appnovation: Full-service digital consultancy serving organizations through digital strategy, design, development, and support.
8+ YOEBachelor's degree in Computer Science or Engineering and 8–12 years of relevant experience. Requires enterprise data integration, cloud infrastructure, IaC, CI/CD, data modeling, AI/MLOps, reliability, and architectural leadership expertise.