CloudFactory
Posted 2d ago

Senior SRE

CloudFactory
Canada
RemoteFull Time, Contract
Responsibilities
  • designing infrastructure
  • automating pipelines
  • supporting reliability
Requirements
  • Requires 5+ years building and operating production infrastructure
  • Python
  • Docker
  • Kubernetes
  • GCP or AWS
  • Terraform, CI/CD, and site-reliability practices. Degree in computing
  • Engineering, or quantitative field preferred, or equivalent experience
Technical tools mentioned
PythonDockerKubernetesGCPAWSTerraformCI/CDPrometheusGrafanaAnsibleChefPuppet

Job description

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. 

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Senior SRE, you will design and build scalable infrastructure, working closely with cross-functional teams to develop systems and pipelines that support the automation, reliability, and scalability of our production environments. You will bring a high degree of autonomy to designing new infrastructure components and applying site-reliability practices across our systems, while communicating complex technical issues clearly to stakeholders across the business. This is an exciting opportunity to make a real impact while working alongside talented people from developing nations.

Please note: This is a full-time, fixed-term employee position with an expected duration of 6 months.

Responsibilities:

Infrastructure design and automation

  • Design and implement new core infrastructure components with a high degree of autonomy.
  • Optimize and improve existing systems and operations, such as deployment pipelines, environment provisioning, and high-throughput batch jobs.
  • Use Infrastructure as Code (IaC) tools, such as Terraform, to manage and scale complex infrastructure.

CI/CD automation

  • Develop CI/CD pipelines to automate build, test, deployment, and monitoring processes.
  • Create and manage multi-step CI/CD pipelines, including environment setup and artifact handling.

Reliability and availability

  • Support the reliability, availability, and performance of production systems, applying site-reliability practices across the infrastructure.
  • Set up monitoring, alerting, and observability tooling to maintain visibility into system health.

Collaboration and communication

  • Collaborate closely with software engineers, product, and business stakeholders on the design and delivery of infrastructure and deployment systems.
  • Communicate complex technical issues clearly to stakeholders from technical and non-technical backgrounds alike.

Requirements

Must-have skills (required)

  • 5+ years of experience building and operating infrastructure in production environments.
  • Fluent in Python, with strong experience writing production-ready code.
  • Experience with Docker and Kubernetes.
  • Knowledgeable about cloud platforms such as GCP or AWS.
  • Experience using Infrastructure as Code (IaC) tools such as Terraform.
  • Experience using CI/CD platforms to automate build, test, and deployment pipelines.
  • Comfortable applying site-reliability principles, such as availability, observability, and automation, across production systems.

Academic and professional requirements

  • Degree in Computer Science, Engineering, or another quantitative or computational field, or equivalent practical experience.

Nice-to-have skills (preferred)

  • Familiarity with monitoring tools such as Prometheus or Grafana.
  • Experience with configuration management tools (e.g., Ansible, Chef, Puppet).
  • Exposure to multi-cloud or hybrid-cloud environments.

Benefits

  • Great Mission and Culture
  • Meaningful Work
  • Market competitive salary
  • Quarterly variable compensation
  • Hybrid Working Model
  • Comprehensive medical cover 
  • Group life insurance
  • Personal development and growth opportunities

At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work—we can’t wait to meet you!

About CloudFactory

Provides human-powered data labeling and AI training solutions.

Year founded
2010
Employees
3600
Organization type
Private
Latest investment
Raised $65.00M Series C (2019) — led by FTV Capital, Weatherford Capital
Headquarters
GB

Similar jobs

Site Reliability Engineer roles
1d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Canada
$131k-$164k/yr RemoteFull Time
Blackpoint Cyber
Blackpoint Cyber: Provides managed detection and response cybersecurity services to businesses.
5+ YOE5+ years in senior SRE or equivalent roles, with expertise in cloud infrastructure, automation, Terraform, AWS, Kubernetes, data streaming, monitoring, incident response, and production troubleshooting.
Terraform, Terragrunt, AWS, Kubernetes, Helm, ArgoCD, Istio, Kustomize, Confluent Cloud, Apache Kafka, Redis, Prometheus, Grafana, Alert Manager, OpsGenie, PagerDuty, LaunchDarkly, PostHog, Amazon RDS, OpenSearch, Elasticsearch, ChaosSearch, Google Cloud Platform, Microsoft Azure, Jenkins, GitHub Actions, Node.js, Python, Go
2d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Toronto, Ontario, Canada
$140k-$155k/yr RemoteFull Time
Caseware
Caseware: AI-powered audit and financial reporting software platform.
8+ YOERequires 8+ years in SRE, platform engineering, DevOps, or related roles; advanced AWS and Kubernetes expertise; Istio, IaC, CI/CD, observability, TypeScript, Node.js, and incident management experience.
AWS, Amazon EKS, AWS IAM, Amazon VPC, AWS Lambda, Amazon CloudFront, Amazon S3, Kubernetes, Istio, AWS CDK, GitHub Actions, AWS CloudWatch, OpenTelemetry, AWS X-Ray, TypeScript, Node.js, Gateway API, mTLS, Certn.co
2d
Save
Mark Applied
Hide
Développeur principal en Ingénierie des Sites
Montreal, Quebec, Canada
HybridFull Time
National Bank of Canada
National Bank of CanadaTSX: NA: Canadian financial institution providing comprehensive banking and investment solutions.
10+ YOERequires 10 years of software development, site reliability, observability, or automation experience; expertise with Datadog, Splunk, Vector, Python, AWS, Terraform, GitHub, and Docker; plus technical leadership.
Datadog, Splunk, Vector, Python, Amazon Web Services, Amazon Elastic Kubernetes Service (EKS), Helm, Amazon EC2, AWS Lambda, Terraform, GitHub, Docker, GitOps, Scrum, SAFe
3d
Save
Mark Applied
Hide
Site Reliability Engineer
Toronto, Ontario, Canada
OnsiteFull Time
Royal Bank of Canada
Royal Bank of CanadaTSX: RY: Provides personal, commercial, and investment banking services worldwide.
Experienced Level 2/3 production support professional with Unix/Linux, SQL, .NET, monitoring, file transfer, PGP, certificate management, ITIL processes, ServiceNow, Jira, and Confluence expertise.
AppDynamics, LogicMonitor, Splunk, SQL, .NET, Unix, Linux, SFTP, Connect Direct, AS2, FTP, FTPS, PGP, ServiceNow, Jira, Confluence, Microsoft SharePoint, ITIL
3d
Save
Mark Applied
Hide
Site Reliability Engineer - Cloud & Platform Engineering, Manulife Bank Technology
Waterloo or Toronto
$86k-$136k/yr HybridFull Time
Manulife
ManulifeTSX: MFC: Provides insurance, wealth management, and investment services globally.
2+ YOERequires 2–5+ years in SRE, DevOps, platform engineering, cloud operations, or related technology; cloud, automation, observability, troubleshooting, communication, and collaboration experience.
Microsoft Azure, AWS, Google Cloud Platform, Python, Bash, PowerShell, New Relic, Grafana, Azure Data Explorer (ADX), Docker, Kubernetes, ITIL, Agile, Microsoft Power BI
5d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Burnaby, British Columbia, Canada
$94k-$139k/yr HybridFull Time
2K
2KNASDAQ: TTWO: Publishes and develops global video game franchises and entertainment.
5+ YOE5+ years in SRE or platform engineering; deep Kubernetes, cloud, IaC, observability, automation, networking, coding, incident response, and post-mortem leadership experience required.
Terraform, Pulumi, ArgoCD, Flux, Kubernetes, EKS, GKE, Istio, Cilium, Prometheus, Grafana, Datadog, GitHub Actions, Jenkins, PasswordState, 1Password, AWS Secrets Manager, OPA/Gatekeeper, Helm, Terragrunt, AWS, GCP, VMware, Ansible, Puppet, AWS Systems Manager, OpenTelemetry, Go, Python, TypeScript, Linux, TCP/IP, DNS, TLS
6d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Cloud Networking & Infrastructure as Code)
Waterloo or Toronto or Ottawa
$120k-$170k/yr HybridFull Time
Magnet Forensics
Magnet Forensics: Provides software for digital forensics and evidence recovery.
Requires networking or computer science education or equivalent experience, strong AWS networking expertise, multi-account and multi-region architecture experience, IaC, CI/CD, scripting, troubleshooting, and communication skills.
AWS, VPC, Transit Gateway, Route 53, VPN, Terraform, AWS CDK, CloudFormation, CI/CD, Python, Bash, PowerShell, TCP/IP, DNS, ISO 27001, SOC 2, NIST
6d
Save
Mark Applied
Hide
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
Montreal, Quebec, Canada
RemoteFull Time, Contract
Software Mind
Software Mind: Provides software engineering and digital transformation services to businesses.
5+ YOERequires 5+ years in SRE, DevOps, platform, or production engineering; Kubernetes production experience; incident response; Splunk, Prometheus, Grafana, CI/CD, Linux, networking, Node.js or JVM troubleshooting.
Kubernetes, Splunk, Prometheus, Grafana, Helm, ArgoCD, Flux, Linux, Node.js, JVM, Java, mTLS, JWT, KEDA, Lit