Oxford Properties
Posted 3w ago

Lead, Site Reliability Engineering (Application Support)

Oxford Properties
Toronto, Ontario, Canada
$86k-$130k/yrHybridFull Time
Responsibilities
  • monitoring applications
  • responding incidents
  • supporting deployments
Requirements
  • 5+ years SRE/Platform/DevOps experience with strong Azure
  • Incident response
  • CI/CD (GitHub Actions)
  • Observability (Datadog/Azure Monitor/Log Analytics)
  • Container and networking knowledge, and scripting (PowerShell/Bash/Python)
Technical tools mentioned
Azure Container AppsAzure Active Directory (Entra ID)Key VaultStorage AccountsAzure SQLAPI Management (APIM)Azure FunctionsGitHub ActionsDatadogAzure MonitorLog AnalyticsPowerShellBashAzure CLIPythonKubernetes

Job description

Choose a workplace that empowers your impact. 

Join a global workplace where employees thrive. One that embraces diversity of thought, expertise and experience. A place where you can personalize your employee journey to be — and deliver — your best.  

We are a purpose-driven, dynamic and sustainable pension plan. An industry leading global investor with teams in Toronto to London, New York, Singapore, Sydney and other major cities across North America and Europe. We embody the values of our 665,000 members, placing their best interests at the heart of everything we do.

Join us to accelerate your growth & development, prioritize wellness, build connections, and support the communities where we live and work.

Don’t just work anywhere — come build tomorrow together with us.

Know someone at OMERS or Oxford Properties? Great! If you're referred, have them submit your name through Workday first. Then, watch for a unique link in your email to apply.

 

Role Summary

The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to guarantee the ongoing stability of all systems. The Lead is also expected to implement and maintain the infrastructure and tools that are necessary to manage the software development process.

Reporting to the SRE, Developer Platform Engineering, the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal applications. This position combines platform engineering, site reliability engineering, and production support responsibilities to ensure applications remain reliable, secure, and easy to operate.

This is an excellent opportunity for someone who is passionate about Site Reliability Engineering and enjoys improving the reliability, availability, performance, and operability of cloud-native platforms. The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while partnering with cross-functional teams to onboard, deploy, monitor, and support applications across the organization.

You Will Be Responsible For

  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.

  • Respond to incidents, lead triage activities, and work closely with SRE, Platform, Network, Security, and application teams to restore service and resolve issues.

  • Support deployments, release activities, change management, and CI/CD pipelines, including GitHub Actions workflows.

  • Check, troubleshoot, and resolve user and developer access issues, including Azure AD groups, SSO, application permissions, and firewall rules.

  • Configure and support platform components such as Azure Container Apps, App Registrations, Key Vault, DNS, certificates, networking, and shared cloud services.

  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, Log Analytics, and related observability tools.

  • Support onboarding of new applications and teams to the DEV platform by helping with setup, access, deployment readiness, monitoring, and operational handover.

  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation to improve support effectiveness and knowledge sharing.

  • Contribute to automation and continuous improvement initiatives that reduce manual effort, improve reliability, and strengthen operational processes.

  • Provide technical guidance to team members and stakeholders while promoting Site Reliability Engineering and platform support best practices.

Required Skills & Experience

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.

  • Strong hands-on experience with Microsoft Azure services, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), and Azure Functions.

  • Experience supporting production environments, including incident response, troubleshooting, problem management, and operational support processes.

  • Experience with container technologies and cloud-native application architectures.

  • Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms.

  • Strong understanding of identity, networking, and access management concepts, including SSO, OAuth, application registrations, and security groups.

  • Experience with observability and monitoring platforms such as Datadog, Azure Monitor, and Log Analytics.

  • Understanding of cloud networking concepts, including DNS, certificates, firewalls, private endpoints, and network security controls.

  • Experience with scripting and automation using technologies such as PowerShell, Bash, Azure CLI, Python, or similar tools.

  • Strong knowledge of operating systems and cloud infrastructure concepts.

  • Proven ability to work effectively in cross-functional environments and collaborate with technical and business stakeholders.

  • Strong communication, problem-solving, and organizational skills.

Preferred Skills & Experience

  • Experience with container apps and Kubernetes container orchestration platforms.

  • Experience with different pipelines.

  • Experience supporting enterprise developer platforms or internal platform engineering teams.

  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and reliability engineering practices.

  • Experience with enterprise API integrations and platform services.

  • Exposure to AI, LLM, or agent-based technology platforms.

  • Experience with Azure networking and security best practices in enterprise environments.

  • Azure, Network, DevOps, or cloud-related certifications.

  • Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline.

We believe that time together in the office is important for OMERS and Oxford, the strength of our employees, and the work we do for our pension members. In delivering on our pension promise, keeping us connected to our work and each other, our flexible hybrid work guideline requires teams to come in to the office 4 days per week. 

  

This posting is for an existing vacancy.

 

The expected salary range for this position is $86,000.00 - $130,000.00 per year.

 

You may also be eligible to receive an annual Incentive Award pursuant to our Short-term Incentive plan and our Long-Term Incentive plan (if applicable), and to participate in our group benefits and retirement plans – details on these elements of compensation are included within OMERS & Oxford offer letters.

 

As one of Canada’s largest defined benefit pension plans, our people-first culture is at its best when our workforce reflects the communities where we live and work — and the members we proudly serve.

From hire to retire, we are an equal opportunity employer committed to an inclusive, barrier-free recruitment and selection process that extends all the way through your employee experience. This sense of belonging and connection is cultivated up, down and across our global organization thanks to our vast network of Employee Resource Groups with executive leader sponsorship, our Purpose@Work committee and employee recognition programs.

 

Artificial intelligence (AI) tools are used to support certain stages of the OMERS recruitment process. While AI assists us in our process, human judgment and decision-making remain central to our candidate experience.

About Oxford Properties

Global real estate investment, development, and management.

Year founded
1960
Employees
2000
Organization type
Private
Latest investment
Raised $515.00M Debt Financing (2025) — led by CIBC Capital Markets, TD Securities, RBC Capital Markets
Headquarters
CA

Similar jobs

Site Reliability Engineer roles near Toronto, Ontario
3h
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Toronto, Ontario, Canada
$140k-$155k/yr RemoteFull Time
Caseware
Caseware: AI-powered audit and financial reporting software platform.
8+ YOERequires 8+ years in SRE, platform engineering, DevOps, or related roles; advanced AWS and Kubernetes expertise; Istio, IaC, CI/CD, observability, TypeScript, Node.js, and incident management experience.
AWS, Amazon EKS, AWS IAM, Amazon VPC, AWS Lambda, Amazon CloudFront, Amazon S3, Kubernetes, Istio, AWS CDK, GitHub Actions, AWS CloudWatch, OpenTelemetry, AWS X-Ray, TypeScript, Node.js, Gateway API, mTLS, Certn.co
1d
Save
Mark Applied
Hide
Site Reliability Engineer
Toronto, Ontario, Canada
OnsiteFull Time
Royal Bank of Canada
Royal Bank of CanadaTSX: RY: Provides personal, commercial, and investment banking services worldwide.
Experienced Level 2/3 production support professional with Unix/Linux, SQL, .NET, monitoring, file transfer, PGP, certificate management, ITIL processes, ServiceNow, Jira, and Confluence expertise.
AppDynamics, LogicMonitor, Splunk, SQL, .NET, Unix, Linux, SFTP, Connect Direct, AS2, FTP, FTPS, PGP, ServiceNow, Jira, Confluence, Microsoft SharePoint, ITIL
1d
Save
Mark Applied
Hide
Site Reliability Engineer - Cloud & Platform Engineering, Manulife Bank Technology
Waterloo or Toronto
$86k-$136k/yr HybridFull Time
Manulife
ManulifeTSX: MFC: Provides insurance, wealth management, and investment services globally.
2+ YOERequires 2–5+ years in SRE, DevOps, platform engineering, cloud operations, or related technology; cloud, automation, observability, troubleshooting, communication, and collaboration experience.
Microsoft Azure, AWS, Google Cloud Platform, Python, Bash, PowerShell, New Relic, Grafana, Azure Data Explorer (ADX), Docker, Kubernetes, ITIL, Agile, Microsoft Power BI
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Cloud Networking & Infrastructure as Code)
Waterloo or Toronto or Ottawa
$120k-$170k/yr HybridFull Time
Magnet Forensics
Magnet Forensics: Provides software for digital forensics and evidence recovery.
Requires networking or computer science education or equivalent experience, strong AWS networking expertise, multi-account and multi-region architecture experience, IaC, CI/CD, scripting, troubleshooting, and communication skills.
AWS, VPC, Transit Gateway, Route 53, VPN, Terraform, AWS CDK, CloudFormation, CI/CD, Python, Bash, PowerShell, TCP/IP, DNS, ISO 27001, SOC 2, NIST
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Edmonton or Vancouver or Kitchener-Waterloo or Toronto
$146k-$197k/yr HybridFull Time
Jobber
Jobber: Software for scheduling, invoicing, and managing home service businesses.
Senior cloud infrastructure engineer with AWS, Terraform, continuous deployment, programming, incident management, automation, and collaboration experience; Azure, Kubernetes, security, and observability are advantageous.
AWS, Infrastructure-as-Code, Terraform, CircleCI, Azure, Ruby, Python, Bash, Ruby on Rails, GQL, React, Kubernetes, Cloudflare, DNS, AI
6d
Save
Mark Applied
Hide
Site Reliability Engineer
Mississauga, Ontario, Canada
$95k-$130k/yr HybridFull Time
Finastra
Finastra: Develops software for global retail and transaction banking.
Experience in application, production support, or systems engineering; Docker and Kubernetes troubleshooting; Bash and Python scripting; Java, Linux, log analysis, incident management, and disaster recovery knowledge.
Docker, Kubernetes, Bash, Python, Grafana, Loki, Jenkins, GitHub CI, ArgoCD, Oracle DB, Azure, AWS, GCP, YAML, XML, AKS, Flux, Ansible, Azure CI/CD, GitHub, Kafka, ElasticSearch, OpenSearch, IBM MQ, DNS, Networking, Red Hat AMQ, Grafana
6d
Save
Mark Applied
Hide
Site Reliability Engineer - Application Support
Toronto, Ontario, Canada
$80k-$90k/yr OnsiteFull Time
Capgemini
CapgeminiEuronext Paris: CAP: Provides global IT consulting and digital transformation services.
Experienced application support or SRE professional with Windows Server, UNIX/Linux, OpenShift, PostgreSQL, incident management, troubleshooting, monitoring, and communication skills; bachelor's degree preferred.
Windows Server, UNIX/Linux, OpenShift Container Platform (OCP), PostgreSQL, SQL, PowerShell, Bash, Python, Shell, Splunk, Dynatrace, AppDynamics, Grafana, Prometheus, CI/CD
6d
Save
Mark Applied
Hide
Senior Site Reliability Developer
Toronto, Ontario, Canada
$107k-$157k/yr OnsiteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
5+ YOEBachelor’s degree in computer science, engineering, or related field; 5+ years in SRE, DevOps, or similar; AWS, Terraform, CI/CD, containers, monitoring, Linux, scripting, and production troubleshooting expertise.
Amazon Web Services (AWS), Terraform, Serverless, CloudFormation, Jenkins, GitHub, Artifactory, Docker, Kubernetes, Amazon ECS, Dynatrace, Grafana, DataDog, ELK Stack, CloudWatch, Java, SpringBoot, Amazon ElastiCache, AWS Lambda, Amazon Kinesis, Amazon DynamoDB, Amazon VPC, AWS IAM, Amazon API Gateway, Kafka, Flink, Jira, Google Apigee, ServiceNow, Splunk, OpenTelemetry, Redis, Gradle, Python, Go, Bash, Groovy, Node.js, UNIX, Linux, Kibana, OpenSearch