ByteDance
Posted 1mo ago

Big Data Site Reliability Engineer - Computer Platform

ByteDance
Singapore
OnsiteFull Time
Responsibilities
  • ensuring stability
  • managing operations
  • troubleshooting issues
Requirements
  • Bachelor's in a computer field
  • Experience in Big Data SRE or technical support for toB products
  • Familiarity with big data components
  • Troubleshooting skills, and proficiency in at least one programming language
Technical tools mentioned
ClickHouseHadoopSparkFlinkHivePresto/TrinoDorisKafkaHBaseHudiShellPythonJavaScala

Job description

About the team
ByteDance and affiliate are developing the next-generation high-performance analytical database, with a mission to enable efficient and real-time data-driven decision-making on PB-level data sets. The initial product was forked from Clickhouse, after which large re-architecture had been taken place. The product now not only improves the efficiency of Clickhouse but also fits into the elastic cloud-native infrastructure with better scalability and resource utilization. With years of polishment in the internal EB-level scenarios, we are now ready to serve our business partners via various cloud vendors.

Responsibilities:
- Ensure the stability of ByteDance's data platform, including building and maintaining the operations and maintenance system for detection, emergency response, recovery, and to guarantee business continuity.
- Manage automated operations and maintenance for ByteDance's in-house big data and open-source products, improving the efficiency of delivery, operations and maintenance, and technical support.
- Promote the accumulation of big data operations experience towards documentation, tooling and standardization, enhancing the operational capabilities of multiple operation and maintenance centers.

Minimum Qualifications:
- Bachelor's degree or above in a computer-related field.
- Experience in Big Data SRE operations or technical support for toB (business-facing) products.
- Familiarity with one or more open-source components, such as Hadoop, Spark, Flink, Hive, Presto/Trino, Doris, Kafka, HBase, Hudi, ClickHouse, etc.
- Practical experience in troubleshooting big data product issues, with methodology for investigating online big data product problems and the ability to quickly locate issues.
- Familiarity with at least one programming language, including but not limited to Shell, Python, Java, Scala, etc.

Preferred Qualification:
- Possess good communication skills, teamwork abilities, and self-driven capabilities for continuous self-improvement.

About ByteDance

Developing AI-driven content platforms and mobile applications.

Similar jobs

Site Reliability Engineer roles
1d
Save
Mark Applied
Hide
Global Banking & Markets, Site Reliability Engineer, Vice President, Singapore
Singapore, Singapore, Singapore
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
8+ YOERequires 8+ years of software or reliability engineering experience, strong programming skills, cloud and distributed-systems expertise, production incident management, automation, observability, risk management, and stakeholder coordination.
Java, GCP, AWS, Kubernetes, Docker, Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist, Apache Kafka, Terraform, Helm, Spring Boot, gRPC, Protocol Buffers, Apache Camel, Spring Integration, Prometheus, Grafana, OpenTelemetry, SQL, NoSQL, Vert.x, Netty
3d
Save
Mark Applied
Hide
Lead Platform Site Reliability Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOEBachelor’s degree in a related discipline, SRE certification or training, 5+ years’ applied experience, programming expertise, observability and CI/CD experience, container orchestration, networking, and AI-assisted reliability workflows.
Python, Java Spring Boot, .Net, Grafana, Dynatrace, Prometheus, Datadog, Splunk, Jenkins, GitLab, Terraform, ECS, Kubernetes, Docker
3d
Save
Mark Applied
Hide
Associate Engineer/Engineer, SRE
Singapore
OnsiteFull Time
Rakuten Viki
Rakuten VikiTokyo Stock Exchange: 4755: Global video streaming platform for Asian movies and shows.
2+ YOEBachelor's degree in computer science, engineering, or equivalent; 2+ years in SRE, DevOps, or backend development, with software, Linux, networking, Docker, Kubernetes, cloud, IaC, CI/CD, and security skills.
Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Amazon Web Services (AWS), Spinnaker, Cloud Build, Datadog, PostgreSQL, RabbitMQ, Redis, Docker, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Linux/Unix
4d
Save
Mark Applied
Hide
Principal Site Reliability Engineer
London or Singapore or Tokyo or Houston or Boston
HybridFull Time
Veson Nautical
Veson Nautical: Develops enterprise software for global maritime freight management.
5+ YOEBachelor's degree or equivalent experience; 5+ years of GCP experience, production Kubernetes/GKE, Terraform, cloud networking, and Python, Go, or TypeScript programming skills.
Google Cloud Platform, Bigtable, Cloud SQL, Dataflow, Datastore, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Google Cloud Key Management Service (KMS), Pub/Sub, Amazon Web Services, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Terraform, Terragrunt, Atlantis, GitLab Pipelines, ArgoCD, Octopus Deploy, ElasticSearch, Kubernetes Operator, PostgreSQL, SQL Server, BigQuery, Splunk, Grafana, Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry, Claude, Amazon Bedrock, Gemini, Vertex AI, Python, Go, TypeScript, GitLab CI
4d
Save
Mark Applied
Hide
Site Reliability Engineer
Singapore
OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOERequires 5–12 years of SRE, L1 production support, or IT operations experience in a financial institution, plus programming, scripting, and relevant operations tool experience.
C++, Java, Shell, Python, Kafka, CDH, OpenShift, J2EE, Spring Boot, Ansible, Kubernetes, MySQL, Hive, Tableau, Druid, Jenkins Pipeline
4d
Save
Mark Applied
Hide
SRE L1 Support/Cloud Platform Ops Engineers
Singapore, Singapore, Singapore
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
2+ YOERequires 2+ years in NOC, data center operations, or IT support; basic Linux administration, monitoring and ticketing systems experience, and ability to perform physical data center tasks and 12-hour shifts.
Linux, Prometheus, Grafana, Nagios, ServiceNow, Jira Service Management, DCGM
5d
Save
Mark Applied
Hide
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore
Singapore, Singapore, Singapore
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
5+ YOEAt least 5 years in SRE or production operations; incident command, Linux, networking, distributed systems, public cloud, observability, automation, SLOs, capacity planning, and strong communication skills required.
Linux, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry, ELK, PagerDuty, Opsgenie, Slack, Microsoft Teams, Terraform, CloudFormation, Ansible, Go, Python, Kubernetes
5d
Save
Mark Applied
Hide
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore
Singapore, Central Region, Singapore
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Provides investment banking, securities, and wealth management services globally.
5+ YOERequires 5+ years in SRE or production operations, incident command experience, Linux, networking, distributed systems, cloud, observability, infrastructure automation, programming, and SLO/error-budget expertise.
Linux, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry, ELK, PagerDuty, Opsgenie, Slack, Microsoft Teams, Terraform, CloudFormation, Ansible, Go, Python, Kubernetes, ITIL