Quickbase
Posted 3w ago

Sr Site Reliability Engineer

Quickbase
Sofia, Sofia City Province, Bulgaria
RemoteFull Time
Responsibilities
  • designing architecture
  • automating infrastructure
  • mentoring engineers
Requirements
  • 7+ years SRE/Cloud/Platform experience with 4+ years in Microsoft Azure
  • IaC (Terraform/Azure Bicep/ARM), CI/CD
  • Scripting
  • Observability
  • Security, and technical leadership
Technical tools mentioned
Microsoft AzureAzure MonitorLog AnalyticsApplication InsightsAzure Key VaultMicrosoft Entra IDTerraformAzure BicepARM templatesAzure DevOpsGitHub ActionsJenkinsPowerShellPythonBashDockerKubernetesAzure Kubernetes Service (AKS)AWSIAMVPCEC2S3CloudWatchAzure Policy

Job description

Senior Site Reliability Engineer (SRE) - Azure Platform

Location: Remote
Department: Site Reliability Engineering (Bulgaria)
Employment Type: Full-Time
Level: Senior Individual Contributor / Technical Lead

Position Summary

We are seeking a senior, hands-on Site Reliability Engineer (SRE) to strengthen our Azure platform capabilities and help mature cloud operations across the organization. This role will serve as a primary technical lead for Microsoft Azure, supporting two existing Azure environments while helping establish standards, operating procedures, and readiness for future development use.

The ideal candidate combines deep Azure expertise, strong SRE discipline, infrastructure automation experience, and practical technical leadership. This person will be expected to lead by example, document and standardize platform practices, mentor other team members, and help train the team so others can support routine Azure administration and operations within the first six months. The role will also provide occasional AWS administration, operational support, and technical guidance as needed.

Key Responsibilities

Cloud Architecture & Platform Engineering

·    Serve as a primary Azure subject matter expert for current Azure environments and related corporate Azure usage.

·    Design, implement, and enhance Azure architecture with emphasis on reliability, security, scalability, observability, cost management, and operational simplicity.

·    Manage and optimize Azure services such as virtual machines, storage, virtual networks, Azure Monitor, Log Analytics, Application Insights, Key Vault, application services, and related platform components.

·    Establish and improve Azure standards for subscriptions, resource groups, networking, naming, tagging, identity, monitoring, backup/recovery, and operational procedures.

·    Provide occasional AWS support as needed, including administration, operational troubleshooting, architecture review, and workload placement guidance.

Infrastructure as Code & Automation

·    Develop, maintain, and govern Infrastructure as Code (IaC) using Terraform, Azure Bicep, ARM templates, or equivalent tooling.

·    Build reusable IaC modules/templates and environment-specific configurations to support secure, repeatable Azure deployments.

·    Build and improve CI/CD automation using Azure DevOps, GitHub Actions, Jenkins, or similar tools, ensuring secure and reliable deployment workflows.

·    Automate provisioning, configuration, monitoring, access management, and operational workflows to reduce manual effort and improve reliability.

·    Use Git-based workflows, pull requests, peer reviews, and change controls to manage infrastructure changes.

Reliability, Observability & Operations

·    Improve logging, monitoring, alerting, and operational dashboards using Azure-native tools and complementary systems.

·    Define and promote reliability practices such as incident response, root-cause analysis, remediation tracking, service health reviews, and operational readiness checks.

·    Act as an escalation resource for significant cloud infrastructure issues and participate as backup in on-call or incident response activities as needed.

·    Identify recurring operational issues and drive long-term fixes through automation, architecture improvements, and better observability.

Security, Compliance & Access Management

·    Administer Azure Key Vault, including secrets management, certificate lifecycle, rotation practices, and governance controls.

·    Implement and manage Azure RBAC, Microsoft Entra ID integrations, managed identities, least-privilege access, and related identity controls.

·    Partner with Security and Compliance teams to ensure infrastructure aligns with SOC 2 controls, internal audit standards, and security best practices.

·    Support policy-driven governance, including tagging, access reviews, configuration standards, and evidence collection for compliance activities.

Technical Leadership, Mentoring & Documentation

·    Lead technical design discussions, platform reviews, and operational improvement efforts across SRE, engineering, security, DevOps, and product teams.

·    Create and maintain clear documentation, runbooks, architecture diagrams, operating procedures, and best-practice guides for Azure operations.

·    Mentor and train other team members through pairing, knowledge-transfer sessions, standards documentation, and operational walkthroughs.

·    Help establish a sustainable support model so routine Azure administration and operations can be shared by additional team members within six months.

·    Communicate technical concepts clearly to both technical and non-technical stakeholders.

Required Qualifications

·    7+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, Systems Engineering, or DevOps roles.

·    4+ years of hands-on experience designing, operating, and improving Microsoft Azure environments.

·    Strong proficiency with Azure architecture and operational services, including networking, compute, storage, identity, monitoring, Key Vault, and application platform services.

·    Hands-on experience with Infrastructure as Code using Terraform, Azure Bicep, ARM templates, or equivalent tooling.

·    Experience building and maintaining CI/CD workflows using Azure DevOps, GitHub Actions, Jenkins, or comparable platforms.

·    Scripting and automation experience using PowerShell, Python, Bash, or similar languages.

·    Demonstrated expertise in logging, alerting, monitoring, incident response, root-cause analysis, and reliability engineering practices.

·    Knowledge of cloud security practices, RBAC, secrets management, managed identities, and least-privilege access models.

·    Familiarity with SOC 2 or similar compliance-driven environments.

·    Proven ability to lead technical initiatives, influence standards, document best practices, and mentor other engineers.

Preferred Qualifications

·    Experience providing AWS administration or operational support, including IAM, VPC/networking, EC2, S3, CloudWatch, and account/resource governance.

·    Experience with Azure landing-zone concepts, Azure Policy, cost management, backup/recovery patterns, and cloud governance models.

·    Relevant certifications such as Microsoft Azure Solutions Architect Expert (AZ-305), Azure DevOps Engineer Expert (AZ-400), Azure Administrator Associate (AZ-104), or equivalent cloud/DevOps certifications.

·    Experience with containerization and orchestration technologies such as Docker, Kubernetes, or Azure Kubernetes Service (AKS).

·    Familiarity with Zero Trust, identity governance, FinOps/cost optimization, and compliance-driven cloud operations.

Success Criteria

Within the first 6 months, success in this role will include:

·    Documented current-state assessment of the two Azure environments.

·    Defined and communicate Azure operating standards, including naming, tagging, access, monitoring, backup/recovery, and change-management practices.

·    Improved runbooks, support procedures, and operational documentation that enable other team members to assist with routine Azure administration and operations.

·    Meaningful knowledge transfer through mentoring, pairing, and practical training sessions with the SRE and/or engineering teams.

·    Clear ownership of Azure reliability, observability, security, and automation improvement priorities.

Within 6-12 months, success in this role will include:

·    Measurable improvements in reliability, performance, observability, security, and operational maturity of Azure workloads.

·    Increased automation coverage across provisioning, CI/CD, monitoring, access controls, and recurring operational tasks.

·    Modernized Azure resources and platform patterns aligned with architectural, security, and compliance standards.

·    A more distributed support model in which trained team members can handle routine Azure tasks with reduced dependence on a single subject matter expert.

·    Reliable contribution to occasional AWS administration, operational support, and cross-cloud guidance when needed.


How We Think About AI
At Quickbase, we view AI as a tool to accelerate how work gets done — not replace it. We encourage thoughtful use of AI to improve speed, quality, and decision-making, while maintaining strong judgment, accountability, and data integrity.


Benefits

  • Unlimited remote work policy
  • 25 days of annual leave, 2 additional days off for volunteering
  • Competitive remuneration package incl. an annual bonus
  • Top-notch IT setup.
  • Mental health support, life insurance, food vouchers
  • Additional health insurance - for you and your loved ones
  • Annual wellness support allowance
  • External Professional Learning Opportunities

Equal Opportunity Statement

Quickbase is committed to building a diverse and inclusive workplace. We encourage candidates from all backgrounds to apply – even if you don’t meet every qualification listed. We are proud to be an equal opportunity employer.

About Quickbase

Provides a no-code platform for building custom business applications.

Year founded
1999
Employees
700
Organization type
Private
Latest investment
Private Equity (2019) — led by Vista Equity Partners
Headquarters
US

Similar jobs

Site Reliability Engineer roles near Sofia, Sofia City Province
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Sr. SRE)
Sofia, Sofia City Province, Bulgaria
HybridFull Time
hosting.com
hosting.com: Global provider of web hosting and cloud infrastructure solutions.
5+ YOEFive years of Linux systems administration experience, Bash scripting, virtualization, Ansible, databases, web servers, complex incident resolution, and fluent written and spoken English.
cPanel, Plesk, Ansible, Linux, KVM, Proxmox, AWX, Salt, MariaDB, MySQL, Apache, LiteSpeed, NGINX, Bash, Docker, Podman, Python, PHP, Go, Xen, Hyper-V, RAID, Windows, IIS, Microsoft DNS, PowerShell
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Sofia, Sofia City Province, Bulgaria
HybridFull Time
Experian
ExperianLondon Stock Exchange: EXPN: Provides data and analytical tools to manage credit risk.
1+ YOE1+ years SRE experience, strong English, Kubernetes and cloud familiarity, Linux and networking troubleshooting, incident management, observability and IaC experience.
Kubernetes, EKS, Splunk, Dynatrace, Thousand Eyes, ServiceNow, Jira, Jenkins, Python, Java, Cassandra, Redis, Apigee, Okta, Postgres, AWS, Microsoft Azure, GCP, Infrastructure as Code, Git Ops
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sofia or Macedonia or Poland or Portugal
HybridFull Time
Valtech
Valtech: Experience innovation providing digital transformation and consulting services.
5+ years engineering experience (minimum 2 years as SRE), production incident management, programming/scripting, cloud and monitoring expertise, English C1+, experience with CI/CD and Kubernetes.
Datadog, New Relic, Dynatrace, Prometheus, Grafana, GitHub, Azure DevOps, GitLab, Jenkins, Docker, Kubernetes, EKS, Argo CI/CD, Java, Springboot, Kafka, AWS, Azure, GCP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer – Middle
Poznań or Warszawa or Sofia
RemoteFull Time
SOFTSWISS
SOFTSWISS: Provider of comprehensive software solutions for the iGaming industry.
3+ YOE3+ years engineering experience; scripting/OOP (Ruby, Python, Java); config management (Ansible/Saltstack/Terraform); monitoring (DataDog/ELK/GrayLog); Kubernetes, DB internals, debugging; English and Russian proficiency.
Ruby, Python, Java, Ansible, Saltstack, Terraform, DataDog, ELK, GrayLog, K8S, Ruby on Rails, SQL, CI/CD
1mo
Save
Mark Applied
Hide
Sr Staff Site Reliability Engineer
Sofia, Sofia City Province, Bulgaria
OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
5+ YOE5+ years SRE experience, strong Kubernetes and Terraform skills, hands-on experience with a major cloud (GCP or AWS), Prometheus/Grafana, CI/CD and GitOps tooling, Python proficiency, incident response and distributed systems troubleshooting.
GCP, AWS, Azure, PagerDuty, Prometheus, Grafana, Kubernetes, Terraform, CI/CD, GitOps, Python, GitLab CI, GitHub Actions, Jenkins, Flux
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center
Sofia, Sofia City, Bulgaria
HybridFull Time
A1
A1Vienna Stock Exchange: TKA: Provides mobile, fixed-line, internet, and digital television services.
SRE with Kubernetes and observability, multi-cloud, and AIOps experience.
Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK/PLG, Datadog, Dynatrace, New Relic, Python, GitHub Actions
3mo
Save
Mark Applied
Hide
Site Reliability Engineer
Sofia, Sofia, Bulgaria
HybridFull Time
Man Group
Man GroupLondon Stock Exchange: EMG: Global alternative investment management firm serving institutional clients.
Strong SRE knowledge with observability, automation, and incident management; hands-on with Prometheus/Grafana/ELK/Loki, Ansible/Terraform, Python/Go/PowerShell; Kubernetes experience; on-call and incident response.
Prometheus, Grafana, ELK, Loki, Ansible, Terraform, Python, Go, PowerShell, Kubernetes
5mo
Save
Mark Applied
Hide
Site Reliability Engineer
Sofia, Sofia, Bulgaria
HybridFull Time
Flutter Entertainment
Flutter EntertainmentNYSE: FLUT: Operates online sports betting and digital casino gaming platforms.
5+ YOEOwn and maintain observability, monitoring, and reliability for AWS-based systems; strong SRE experience; cloud platforms AWS/Azure/GCP; on-call; 24/7 support.
Prometheus, Grafana, ELK, AWS, Azure, GCP, Jenkins, GitLab CI, Azure DevOps, Docker, Kubernetes, Terraform