UST
Posted 2d ago

Lead I - DevOps Engineering

UST
Bengaluru, Karnataka, India
OnsiteFull Time
Responsibilities
  • operating platforms
  • designing pipelines
  • automating deployments
Requirements
  • Requires 5+ years in monitoring
  • Observability
  • Platform engineering, SRE, or IT operations
  • Observability platforms
  • OpenTelemetry, cloud
  • Linux/Windows
  • Kubernetes
  • Scripting, IaC, CI/CD, and a bachelor's degree or equivalent
Technical tools mentioned
DynatraceDatadogNew RelicSplunkElasticPrometheusGrafanaOpenTelemetryJiraMicrosoft TeamsSlackPagerDutyOpsgeniePythonPowerShellBashTerraformBicepCloudFormationAnsibleAzure MonitorLog AnalyticsAWS CloudWatchGoogle Cloud OperationsKubernetesTempoLokiLinuxWindowsDNSTCP/IPCMDBITSMCI/CDSLOSLASREMTTDMTTR

Job description

Job Details

  • Location: Bangalore
  • Experience: 5 - 7 Years
  • Job Type: Full Time
  • Openings: 1

Role Description

Job Code 15023 Locations Bengaluru Minimum Experience 5 Maximum Experience 8 Mandatory Skills Observability,Telemetry,Cloud,Containers,Infrastructure as Code,Automation,CI/CD Skill to Evaluate Observability,Telemetry,Cloud,Containers,Infrastructure as Code,Automation,CI/CD Experience 5 to 8 Years Location Bengaluru Job Description Job Title: Platform Engineer Job Description: As a Platform Engineer, you will build, operate, and continuously improve enterprise monitoring and observability platforms that enable reliable service delivery. You will design scalable telemetry pipelines for metrics, logs, and traces; improve data quality and effectiveness; and develop dashboards that provide clear insight into service health and performance. Working closely with application, infrastructure, security, compliance, and incident-response teams, you will automate platform operations, establish monitoring standards, and improve detection-to-diagnosis workflows. Your work will help reduce noise, accelerate incident recovery, and promote consistent adoption of observability practices across cloud and on-premises environments. Key Responsibilities: • Engineer and operate enterprise monitoring and observability platforms such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, and Grafana. • Design and maintain telemetry collection using agents, exporters, collectors, and API-based integrations across applications, servers, networks, databases, containers, storage, and cloud services. • Manage the performance, capacity, high availability, security, and lifecycle upgrades of monitoring platforms and supporting infrastructure. • Implement platform security controls, including role-based access control, single sign-on, encryption, secrets management, and auditability. • Establish ing standards for severity, thresholds, anomaly detection, suppression, correlation, deduplication, and notification routing. • Reduce fatigue by tuning signals, removing redundant rules, and creating actionable s with links to dashboards, runbooks, and diagnostic context. • Build standardized dashboards for service health, application performance, infrastructure capacity, availability, and SLO/SLA tracking. • Partner with service owners to define and implement service-level indicators, service-level objectives, and error budgets where applicable. • Report operational trends such as availability, mean time to detect, mean time to restore, recurring incidents, and platform adoption. • Automate platform configuration, deployment, and service onboarding using infrastructure-as-code, configuration-management, and CI/CD tools. • Integrate monitoring platforms with Jira, CMDB systems, Microsoft Teams, Slack, PagerDuty, Opsgenie, and other incident-management tools. • Develop reusable s, dashboards, templates, policies, and self-service onboarding patterns. • Define monitoring standards covering minimum telemetry requirements, metadata, tagging, naming conventions, and data governance. • Maintain technical documentation, operational procedures, troubleshooting guides, and onboarding runbooks. • Participate in incident investigations and post-incident reviews to improve detection, prevent recurrence, and strengthen monitoring coverage. • Collaborate with Security and Compliance teams to meet privacy, data-retention, regulatory, and audit requirements. • Participate in a scheduled 24×7 on-call rotation supporting monitoring infrastructure and services. Skills and Tools Required: • Three to seven or more years of experience in monitoring and observability, systems or platform engineering, site reliability engineering, or IT operations. • Hands-on experience with at least one enterprise observability platform, such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, or Grafana. • Strong understanding of metrics, logs, traces, telemetry pipelines, and distributed-system troubleshooting. • Experience with OpenTelemetry concepts, collectors, and instrumentation best practices. • Experience administering and troubleshooting Linux and Windows systems in large-scale environments. • Working knowledge of networking fundamentals, including DNS, TCP/IP, firewalls, load balancers, and common infrastructure components. • Experience with cloud-monitoring services such as Azure Monitor and Log Analytics, AWS CloudWatch, or Google Cloud Operations. • Knowledge of container and Kubernetes observability using technologies such as Prometheus, Grafana, Tempo, Loki, and OpenTelemetry. • Scripting and automation skills using Python, PowerShell, or Bash. • Experience with infrastructure-as-code and configuration-management tools such as Terraform, Bicep, CloudFormation, or Ansible. • Familiarity with version control, CI/CD pipelines, automated deployments, and repeatable build practices. • Experience integrating observability platforms with ITSM, CMDB, event-management, and incident-response workflows. • Understanding of SRE principles, including SLIs, SLOs, error budgets, reliability engineering, MTTD, and MTTR. • Strong ability to distinguish meaningful signals from noise and design s that support effective action. • Systems-level troubleshooting skills across application, platform, infrastructure, database, and network layers. • Excellent written and verbal communication skills, including the ability to work with technical and non-technical stakeholders. • A customer-focused approach to self-service enablement, documentation, standardization, and platform adoption. • The ability to prioritize effectively, work in a fast-paced operational environment, and handle incident escalation calmly. • A bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent professional experience. Preferred Certifications: • Monitoring-platform certification from vendors such as Datadog, Dynatrace, or New Relic. • Cloud certification for Microsoft Azure, Amazon Web Services, or Google Cloud. • ITIL Foundation certification. Join Sony as a Platform Engineer and help establish reliable, scalable, and actionable observability capabilities across the organization. You will work within an inclusive, collaborative, and global technology community while contributing to service reliability, operational efficiency, and continuous improvement. Education Qualificaiton Bachelor’s Degree in Computer Science or a related field Project Details The Monitoring Platform Engineering team builds and operates scalable, secure, and reliable monitoring solutions for enterprise applications and infrastructure across cloud and on-premises environments. The role involves managing telemetry pipelines, dashboards, s, and platform integrations while automating deployment and service onboarding through IaC and CI/CD. The project aims to reduce noise, improve incident detection and resolution, and strengthen service reliability and operational efficiency. Shift Timings 9am - 5PM

Skills

Observability, OpenTelemetry, Cloud Infrastructure, CI/CD

About UST

Global provider of digital transformation and IT services.

Year founded
1999
Employees
30000
Organization type
Private
Latest investment
Raised $1.00B Private Equity (2018) — led by Temasek
Headquarters
US

Similar jobs

Platform Engineer roles near Bengaluru, Karnataka
1d
Save
Mark Applied
Hide
Platform Engineer - ServiceNow (Global Role)
Bangalore or Chennai
OnsiteFull Time
Atos
AtosEuronext Paris: ATO: Provides global IT services, cybersecurity, and digital transformation solutions.
5+ YOERequires 5+ years of ServiceNow experience, 8+ years of platform engineering or enterprise applications experience, and mandatory CSA and CAD certifications; CTA and additional CIS certifications preferred.
ServiceNow, App Engine, Flow Designer, Workflow Studio, Integration Hub, REST, SOAP, Scripted REST APIs, MID Server, ETL, Virtual Agent, Now Assist, AI Search, AI Agents, OAuth, SSO, MFA, CI/CD, Workflow Data Fabric (WDF), JavaScript, Glide APIs, Import Sets, Transform Maps, CMDB, Security, Performance Analytics, API Gateways, ESB/iPaaS, RPA, IoT, OT
1d
Save
Mark Applied
Hide
Senior Platform Engineer - Directory Services
Bengaluru or Bangalore or Bengaluru
HybridFull Time
The Walt Disney Company
The Walt Disney CompanyNYSE: DIS: Produces media content and operates global theme parks.
8+ YOE8+ years engineering large-scale Active Directory environments; expertise in AD architecture, trusts, replication, DNS, PKI, multi-domain forests, security hardening, RadiantLogic, RBAC, PIM, and PAM. Bachelor's degree or equivalent experience.
Active Directory, RadiantLogic, Virtual Directory System, RBAC, Privileged Identity Management (PIM), Privileged Access Management (PAM), Entra ID, Azure AD, DNS, PKI, Privileged Access Workstations (PAW)
1d
Save
Mark Applied
Hide
Platform Engineer, Java Fmwk
Bengaluru or Bangalore or India or United States
OnsiteFull Time
eBay
eBayNASDAQ: EBAY: Global online marketplace for buying and selling diverse products.
4+ YOERequires 4–7 years of Java software development, Spring Boot, Maven, REST APIs, Git, CI/CD, Docker, Kubernetes, microservices, distributed systems, automation, and cloud-native experience.
Java, Spring Boot, Raptor.io, REST, GraphQL, Maven, Jenkins, Tekton, Git, Docker, Kubernetes, CI/CD
2d
Save
Mark Applied
Hide
Senior Platform Engineer- DevOps+ Azure Red Hat OpenShift
Bangalore, Karnataka, India
OnsiteFull Time
CGI
CGINYSE: GIB: Provides information technology and business consulting services.
5+ YOE5–9 years in cloud, platform engineering, DevOps, or SRE; Azure Red Hat OpenShift, Kubernetes, CI/CD, GitOps, Terraform, cloud infrastructure, security, troubleshooting, and technical leadership experience required.
Microsoft Azure, Azure Red Hat OpenShift, AWS, Google Cloud Platform, Terraform, CI/CD, GitOps, Linux, Windows, Kubernetes, OpenShift, Docker, Podman, ARM, Bicep, Ansible, CloudFormation, CloudBees CI, Jenkins, Azure DevOps, GitHub Actions, Argo CD, GitHub, Azure Repos, Azure DevOps Services, Bitbucket, Nexus Repository Manager, Python, Bash, PowerShell, JavaScript, Snyk, Veracode, Red Hat Advanced Cluster Security, Microsoft Azure certification, Red Hat OpenShift certification, Kubernetes certification
3d
Save
Mark Applied
Hide
Senior Platform Engineer I - Evergreen
Bangalore, Karnataka, India
OnsiteFull Time
Booking Holdings
Booking HoldingsNASDAQ: BKNG: Global provider of online travel and related services.
8+ YOEMaster's degree and 8+ years of relevant knowledge in software application development, system design, service ownership, incident management, observability, architecture, communication, and mentoring.
SMART
3d
Save
Mark Applied
Hide
Platform Engineer - Open Shift Platform
Bengaluru, Karnataka, India
OnsiteFull Time
Airbus
AirbusEuronext Paris: AIR: Designs and manufactures commercial and military aircraft and space systems.
Professional platform engineering role developing, integrating, testing, deploying, and operating scalable software systems while following architectural standards and managing compliance risks.
OpenShift
5d
Save
Mark Applied
Hide
Technical Lead, Platform Engineering (Observability)
Bengaluru, Karnataka, India
HybridFull Time
Mercari
MercariTokyo Stock Exchange: 4385: C2C marketplace for buying and selling pre-owned items.
9+ YOERequires 9+ years building scalable production systems, observability platforms, Kubernetes, Go or Python, GCP or AWS, Terraform, distributed tracing, alerting, SLOs, and technical leadership.
Datadog, Prometheus, Grafana, Kubernetes, Go, Python, GCP, AWS, Terraform, OpenTelemetry
6d
Save
Mark Applied
Hide
Engineering Division - Global Cyber Defense & Intel - Associate - Bengaluru
Bengaluru, Karnataka, India
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
4+ YOEBachelor's degree or equivalent practical experience; 4+ years in infrastructure, platform, or cloud engineering; production AWS, Kubernetes, Linux, containerization, DevSecOps, IaC, and observability experience.
Kafka, Spark, Kubernetes, BigQuery, AWS, Azure, GCP, Terraform, Ansible, Git, VPC, IAM, EC2, EKS, RDS, S3, SQS, SNS, CloudWatch, Docker, Prometheus, Grafana, Loki, FluentBit, OpenTelemetry, Helm, GitOps, Flink, Linux, Bash, Python, Packer, OPA, Kyverno, GitLab CI, Harness, CloudFormation, Go, R, SLOs, SLIs