ByteDance
Posted 1mo ago

Site Reliability Engineer - Big Data Computer Platform

ByteDance
Singapore
OnsiteFull Time
Responsibilities
  • ensuring reliability
  • troubleshooting incidents
  • automating operations
Requirements
  • 2+ years SRE experience
  • Bachelor's in CS/Engineering
  • Strong Linux
  • Networking
  • Database knowledge
  • Kubernetes and monitoring tool proficiency
  • Coding in Python/Shell/Java/Go
  • Strong problem-solving and communication
Technical tools mentioned
ClickHouseSparkPrestoDorisHadoopKubernetesPythonShellJavaGo

Job description

About the team
Our Compute Platform SRE team supports all Big Data services and products across the company. We are a newly established team and waiting for talents like you to shape the team's future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains within TikTok. We look forward to welcoming you to the team.

Responsibilities:
- Ensure the reliability of all TikTok's major data warehouse products, services, and query engines, such as ClickHouse, Spark, Presto, Doris, etc.
- Ensure that all service level objectives and agreements from ByteDance's Data Platform services are met; respond promptly to any system outages or issues.
- Analyze service performance and reliability patterns to identify potential performance bottlenecks. Implement proactive measures to prevent service disruptions. Work with development teams to optimize application performance, ensuring that services run efficiently and that resources are utilized effectively.
- Build robust incident management mechanism. Lead efforts to troubleshoot and resolve service incidents and postmortems. Coordinate with cross-functional teams to manage and mitigate service-impacting events.
- Develop highly efficient toolchains covering end-to-end deployment and reliability assurance operations. Automate infrastructure provisioning, scaling, and management processes to reduce manual interventions and improve service quality. Develop and enhance system capabilities such as auto-failure-detection, auto-healing, chaotic engineering, and perform systematic disaster drills.
- Engage with product and development teams to integrate reliability and performance considerations into the software lifecycle.
- Assess and forecast infrastructure needs based on growth patterns and upcoming initiatives.
- Keep current with industry trends, best practices, and emerging technologies related to site reliability and infrastructure engineering.

Minimum Qualifications:
- Bachelor's Degree or above in Computer Science, Engineering, or a related field. Passionate about computer science and Internet technology.
- At least 2 years of experience in the SRE domain.
- At least 2 years of experience and in-depth understanding of Linux, computer networking, and databases. Proficient in common SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration platforms like Kubernetes.
- Familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, Presto and Kubernetes.
- At least 2 years of experience in coding in at least one scripting or programming language, including but not limited to Python, Shell, Java, Go, etc.

Preferred Qualifications:
- Excellent problem-solving skills and the ability to think critically. Start with the end state in mind, and be willing to take a moonshot.
- Strong written and verbal communication skills, with great customer-first mindset. Strong sense of ownership and easy to collaborate with.
- Able to collaborate effectively with partners and team members across time zones in different countries.

About ByteDance

Developing AI-driven content platforms and mobile applications.

Similar jobs

Site Reliability Engineer roles
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Singapore
OnsiteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
5+ YOERequires a bachelor's degree or equivalent certification, 5+ years of software engineering and SRE experience, Python and Terraform, AI/ML production experience, Kubernetes, and rotating 24x7 on-call availability.
Terraform, GitHub Actions, Python, Node.js, GCP, AWS, Azure, Kubernetes, EKS, AKS, GKE
1d
Save
Mark Applied
Hide
AVP/VP - Site Reliability Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
SGX Group
SGX GroupSingapore Exchange: S68: Operates Singapore's primary securities and derivatives exchange markets.
5+ YOERequires 5+ years in SRE, platform, or infrastructure engineering; production programming in a modern language; Kubernetes, Terraform or equivalent IaC, CI/CD, observability, cloud, distributed systems, and incident response expertise.
Kafka, Kubernetes, CI/CD, Terraform, GCP, AWS, Go, Python, Kotlin, TypeScript, Rust, FIX protocol
3d
Save
Mark Applied
Hide
Global Banking & Markets, Site Reliability Engineer, Vice President, Singapore
Singapore, Singapore, Singapore
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
8+ YOERequires 8+ years of software or reliability engineering experience, strong programming skills, cloud and distributed-systems expertise, production incident management, automation, observability, risk management, and stakeholder coordination.
Java, GCP, AWS, Kubernetes, Docker, Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist, Apache Kafka, Terraform, Helm, Spring Boot, gRPC, Protocol Buffers, Apache Camel, Spring Integration, Prometheus, Grafana, OpenTelemetry, SQL, NoSQL, Vert.x, Netty
5d
Save
Mark Applied
Hide
Lead Platform Site Reliability Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOEBachelor’s degree in a related discipline, SRE certification or training, 5+ years’ applied experience, programming expertise, observability and CI/CD experience, container orchestration, networking, and AI-assisted reliability workflows.
Python, Java Spring Boot, .Net, Grafana, Dynatrace, Prometheus, Datadog, Splunk, Jenkins, GitLab, Terraform, ECS, Kubernetes, Docker
5d
Save
Mark Applied
Hide
Associate Engineer/Engineer, SRE
Singapore
OnsiteFull Time
Rakuten Viki
Rakuten VikiTokyo Stock Exchange: 4755: Global video streaming platform for Asian movies and shows.
2+ YOEBachelor's degree in computer science, engineering, or equivalent; 2+ years in SRE, DevOps, or backend development, with software, Linux, networking, Docker, Kubernetes, cloud, IaC, CI/CD, and security skills.
Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Amazon Web Services (AWS), Spinnaker, Cloud Build, Datadog, PostgreSQL, RabbitMQ, Redis, Docker, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Linux/Unix
6d
Save
Mark Applied
Hide
Principal Site Reliability Engineer
London or Singapore or Tokyo or Houston or Boston
HybridFull Time
Veson Nautical
Veson Nautical: Develops enterprise software for global maritime freight management.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
6d
Save
Mark Applied
Hide
Site Reliability Engineer
Singapore
OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOERequires 5–12 years of SRE, L1 production support, or IT operations experience in a financial institution, plus programming, scripting, and relevant operations tool experience.
C++, Java, Shell, Python, Kafka, CDH, OpenShift, J2EE, Spring Boot, Ansible, Kubernetes, MySQL, Hive, Tableau, Druid, Jenkins Pipeline
6d
Save
Mark Applied
Hide
SRE L1 Support/Cloud Platform Ops Engineers
Singapore, Singapore, Singapore
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
2+ YOERequires 2+ years in NOC, data center operations, or IT support; basic Linux administration, monitoring and ticketing systems experience, and ability to perform physical data center tasks and 12-hour shifts.
Linux, Prometheus, Grafana, Nagios, ServiceNow, Jira Service Management, DCGM