Great Eastern Life Assurance (Malaysia) Berhad
Posted 2mo ago

Lead Site Reliability Engineer

Great Eastern Life Assurance (Malaysia) Berhad
Singapore or Malaysia
OnsiteFull Time
Responsibilities
  • designing infrastructure
  • monitoring performance
  • troubleshooting issues
Requirements
  • Design
  • Implement, and maintain VMware Cloud Foundation infrastructure
  • Manage resources and backups
  • Monitor health via vROps/Dynatrace
  • Perform security remediation
  • 15 years IT infrastructure experience
  • Scripting (Python/PowerShell)
  • Ansible/Terraform automation
Technical tools mentioned
VMware Cloud Foundation (VCF)NSX-TvCenterESXiVMware vSpherevSANNetBackup (NBU)vRealize AutomationvRealize OperationsvROPSDynatraceHPSAPythonPowerShellAnsibleTerraformPrometheusGrafana

Job description

A Site Reliability Engineer (SRE) for VMware Cloud Foundation (VCF) focuses on ensuring the reliability, availability, and performance of the VCF platform through automation, monitoring, and proactive problem-solving. This role involves developing and implementing strategies to improve the platform's stability, collaborating with development and operations teams, and contributing to the overall VCF roadmap.

·        VCF Infrastructure Design and Maintenance:

Experience to design, implement, and maintain VMware Cloud Foundation (VCF) infrastructure to support GE’s organizational requirements.

·        Resource Management and Troubleshooting:

Manage and troubleshoot VCF resource availability, including compute, memory, and storage (SAN and vSAN) up to 160 ESXi and more than 1200 VMs across multiple clusters located in both SG and MY, using tools such as NSX-T, vCenter, ESXi 8.x, and VMware vSphere Cluster availability. 

·        Backup Services:

Experience in integration with backup services using NetBackup (NBU) such as HotAdd for image backup/restore and Media to file level backup/restore to support business application VMs requirements, including full, incremental, and ad-hoc backups.

·        Health Monitoring and Performance Management:

Conduct daily health checks and monitor VCF infrastructure metrics via vROPS, vCenter, Dynatrace to ensure optimal workload performance and timely issue resolution.

·       Security Compliance and Remediation:

Analyze VCF components and perform NVA security remediation to maintain compliance across vSphere, NSX, vSAN, and other VCF elements. Maintains awareness of industry trends on regulatory MAS (SG) and BNC (MY) compliance, emerging threats and technologies to understand the risk and better safeguard the company. Experience with HPSA scanning tools is a plus.         

·        Operational Documentation:

Develop and maintain comprehensive Standard Operating Procedures (SOPs) for VCF operations, including ESXi uptime/downtime records, VCF inventory, recovery procedures, and disaster recovery plans.

·        Patch Management:

Apply updates, service packs and patching to ESXi hosts and vSphere components to ensure security and product currency.

·        Security Policy Implementation:

Collaborate with the security team to implement required policies, including hardening measures to protect VCF nodes, NSX firewall, DSA on VM level etc. Takes accountability in considering business and regulatory compliance risks and takes appropriate steps to mitigate the risks.

·        Project Execution:

Execute VCF-related infrastructure projects, ensuring timely delivery and alignment with business requirements.

·        Vendor Collaboration:

Work with vendors and third-party contractors to manage projects and implementation of VCF-related products and services. Partner with the project delivery team to identify business application requirements and support deployment on the VCF platform.


·   Bachelor’s degree in computer science, Information Technology, Computer Engineering, or a related field. A Master’s degree or relevant certifications (e.g., ITIL, TOGAF, Cloud certifications or VMware VCP) is a plus.

·        Extensive Knowledge of VMware Cloud Foundation (VCF) and NSX:

Strong understanding and hands-on experience with VCF components, including vSphere, vSAN, NSX, and the vRealize Suite (e.g., vRealize Automation, vRealize Operations).

·        Industry Experience:

15 years of experience in IT infrastructure roles, with a significant portion focused on VMware VCF solutioning, hand-on deployment experiences, and be able to work on enterprise level capabilities.

·        Scripting Proficiency:

Skilled in scripting languages such as Python or PowerShell.

·        Automation Expertise:

Hand-on Experience with automation tools and frameworks, including Ansible and Terraform.

·        Monitoring and Alerting Tools:

Familiarity with implementing monitoring and alerting solutions such as Prometheus, Grafana, Dynatrace, or vRealize Operations.

·        Problem-Solving Skills:

Demonstrated ability to troubleshoot and resolve complex technical issues effectively.

·        Communication and Collaboration:

Strong communication skills with the ability to collaborate across various levels of stakeholders.



Job Details

To All Recruitment Agencies: Great Eastern does not accept unsolicited agency resumes. Please do not forward resumes to our email or our employees. We will not be responsible for any fees related to unsolicited resumes.

About Great Eastern Life Assurance (Malaysia) Berhad

Malaysian life insurer providing life, health, investment-linked, employee-benefit, and group protection products to individuals, families, and employers.

Similar jobs

Site Reliability Engineer roles
5d
Save
Mark Applied
Hide
Site Reliability Engineer / SRE (Senior / Lead)
Singapore
OnsiteFull Time
Reolink
Reolink: Private security-camera manufacturer and online retailer serving households and businesses worldwide.
5+ YOEBachelor's degree or higher preferred in computer science or related field; 5+ years in cloud services and application operations; experience with AWS, Azure, Google Cloud, Docker, Kubernetes, CI/CD, Jenkins, or GitLab CI.
Amazon Web Services (AWS), Microsoft Azure, Google Cloud, Docker, Kubernetes, CI/CD, Jenkins, GitLab CI
1w
Save
Mark Applied
Hide
SVP, Site Reliability Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
Singapore Exchange
Singapore ExchangeSGX-ST: S68: SGX Group is a publicly traded Singaporean multi-asset exchange providing listing, trading, clearing, settlement, depository and data services.
Executive-level SRE or large-scale technology operations experience in mission-critical environments; expertise in enterprise reliability, observability, cloud infrastructure, resilience, governance, and engineering leadership; bachelor's degree required.
Datadog, New Relic, Grafana Cloud, AWS, Kubernetes, EKS, Terraform, DORA, SLIs, SLOs, SLA, MTTD, MTTR, AI
1w
Save
Mark Applied
Hide
Senior SRE, Infrastructure & Platform
Singapore
HybridFull Time
F5
F5NASDAQ: FFIV: Delivering and securing applications across any multi-cloud environment.
5+ YOERequires 5+ years in SRE, DevOps, or infrastructure engineering; strong Python, Ansible, Linux, CI/CD, Kubernetes, AWS, Azure, API, security automation, and software engineering expertise.
Python, Go, Ansible, NetBox, HashiCorp Vault, Proxmox, Proxmox API, iLO Redfish, GitLab CI, Molecule, ansible-lint, Kubernetes, Helm, GitOps, AWS, Azure, RESTful APIs, RHEL, CentOS, VMware vSphere, libvirt, IPMI, operator-sdk, kubebuilder, controller-runtime, OSTree, Pulp, Backstage
1w
Save
Mark Applied
Hide
Site Reliability Engineer, Security Engineering
Sydney or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteFull Time
TikTok
TikTok: Short-form mobile video and social media platform.
3+ YOEBachelor's degree in computer science or related field, 3+ years relevant experience, programming in Go, Java, or Python, web framework experience, Linux and networking knowledge, Kubernetes and SRE tooling experience.
Go, Java, Python, Gin, Django, Spring, Linux, TCP/IP, HTTP, Kubernetes, Ansible, Argo CD, Prometheus, Grafana
1w
Save
Mark Applied
Hide
Site Reliability Engineer Intern
Singapore, Singapore, Singapore
OnsiteInternship, Full Time
ShopBack
ShopBack: Singapore-headquartered private shopping, rewards, and payments platform helping consumers earn cashback on everyday purchases.
0+ YOEStudents in year 3 or 4 of a bachelor's program or above, or fresh graduates, with software engineering fundamentals and experience in networking, Docker, GitOps, Kubernetes, programming, scripting, AWS, Postgres, and Linux.
Docker, GitOps, Kubernetes, Java, C, C++, C#, Objective-C, Python, JavaScript, Go, Shell, AWS, Postgres, Linux, ChatGPT, Gemini, Claude
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Singapore
OnsiteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Diversified health care helping people live healthier lives.
5+ YOERequires a bachelor's degree or equivalent certification, 5+ years of software engineering and SRE experience, Python and Terraform, AI/ML production experience, Kubernetes, and rotating 24x7 on-call availability.
Terraform, GitHub Actions, Python, Node.js, GCP, AWS, Azure, Kubernetes, EKS, AKS, GKE
2w
Save
Mark Applied
Hide
Global Banking & Markets, Site Reliability Engineer, Vice President, Singapore
Singapore, Singapore, Singapore
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities and investment management firm.
8+ YOERequires 8+ years of software or reliability engineering experience, strong programming skills, cloud and distributed-systems expertise, production incident management, automation, observability, risk management, and stakeholder coordination.
Java, GCP, AWS, Kubernetes, Docker, Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist, Apache Kafka, Terraform, Helm, Spring Boot, gRPC, Protocol Buffers, Apache Camel, Spring Integration, Prometheus, Grafana, OpenTelemetry, SQL, NoSQL, Vert.x, Netty
2w
Save
Mark Applied
Hide
Lead Platform Site Reliability Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services and investment banking firm.
5+ YOEBachelor’s degree in a related discipline, SRE certification or training, 5+ years’ applied experience, programming expertise, observability and CI/CD experience, container orchestration, networking, and AI-assisted reliability workflows.
Python, Java Spring Boot, .Net, Grafana, Dynatrace, Prometheus, Datadog, Splunk, Jenkins, GitLab, Terraform, ECS, Kubernetes, Docker