ByteDance
Posted 9h ago

Site Reliability Engineer Intern (Data Infra) - 2027 Fall

ByteDance
Seattle, Washington, United States
OnsiteInternship
Responsibilities
  • enhancing service lifecycle
  • implementing platforms
  • managing infrastructure
Requirements
  • Currently pursuing a bachelor's degree in computer science or related field
  • Programming experience in C, C++, Java, Python, Go, or Rust
  • Knowledge of Unix/Linux internals
  • Networking, and distributed systems
Technical tools mentioned
CC++JavaPythonGoRustUnixLinuxKubernetesRedisMySQLFlinkNginxDockerOpenStackHadoopSpark

Job description

About the Team:
Our data infrastructure Site Reliability Engineering (SRE) team is a pioneer in innovation. We seamlessly merge software development and infrastructure operations to design, build, and manage large-scale, highly distributed systems.

We take pride in overseeing one of the industry's most extensive cloud infrastructures. As software development evolves, building systems from a mix of components has become the new standard. In this era, SRE takes a central role. This role demands the ability to design, develop, and operate these components, transforming them into cloud-managed, scalable, and reliable elements. Our professionals play a critical role as connectors, ensuring the seamless integration of these diverse components to deliver high-performing systems.

Our dynamic SRE field is about actively shaping the future of technology, not just keeping pace with it. We contribute significantly to the next chapter of data infrastructure. We're currently in the process of building global teams around the world. Join us today and embark on this transformative journey!

We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.

Candidates who pass resume screening will be invited to participate in Our Company's technical online assessment.

Responsibilities:
- Participate in and enhance the complete service lifecycle, from inception and design, through development, capacity planning, launch reviews, deployment, operation, and refinement.
- Design and implement software platforms and monitoring frameworks to govern service-oriented architecture (SOA) efficiently, automatically, and intelligently.
- Develop and manage components of cloud-managed data infrastructure, encompassing technologies such as Kubernetes, Redis, MySQL, Flink, and more.
- Establish sustainable mechanisms for scaling systems, such as automation, to drive enhancements in reliability, efficiency, and velocity.
- Provide sustainable user support, manage incident responses, and conduct blameless postmortems as part of our ongoing efforts to improve our systems.

Minimum Qualifications:
- Currently pursuing a Bachelor's degree in Computer Science or a related technical discipline.
- Experience programming in one of the following Languages: C, C++, Java, Python, Go, and Rust
- Knowledge of Unix/Linux system internals, networking, and distributed systems

Preferred Qualifications:
- Currently pursuing a Master's degree in Computer Science or a related technical discipline.
- Experience in MySQL, Redis, Ngnix, Kubernetes, Docker, OpenStack, Hadoop, Spark, Flink, etc.
- Strong skills in fast learning and communication

About ByteDance

Developing AI-driven content platforms and mobile applications.

Similar jobs

Site Reliability Engineer roles near Seattle, Washington
2d
Save
Mark Applied
Hide
Senior Engineer Site Reliability Engineer
Bellevue, Washington, United States
$84k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
10+ YOERequires 10 years of experience, Azure expertise, scripting with PowerShell, Python, or Bash, strong communication, monitoring and alerting experience, and familiarity with Scrum or Kanban.
Microsoft Azure, Microsoft SQL Server 2019, PowerShell, Python, Bash, Scrum, Kanban
5d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOERequires 7+ years in SRE, HPC engineering, DevOps, or similar; expertise in AI infrastructure, Linux, distributed systems, networking, Python, Go, monitoring, and automation tools.
Linux, Ansible, Terraform, Python, Go, Prometheus, Grafana, ClickHouse, PyTorch, TensorFlow, DeepSpeed, MLPerf, Docker, Kubernetes, InfiniBand, RoCE, NCCL, GPU-direct, CLOS, 100GbE, Ethernet, SOC 2, ISO 27001
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Top Secret Clearance (Starlink)
Redmond or Hawthorne
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor's degree in a relevant discipline and 5 years of software development experience, or 7+ years' professional software experience; Linux and active Top Secret or Top Secret/SCI clearance required.
Linux, Kubernetes, Istio, Apache Kafka, Spark, HBase, HDFS, Flink, Python, C#, Java, Scala, Go
1w
Save
Mark Applied
Hide
Alibaba Cloud-ECS Site Reliability Engineer-Bellevue
Bellevue, Washington, United States
$133k-$220k/yr OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE: BABA: Global cloud computing infrastructure and services provider
3+ YOEBachelor's degree in computer science, information technology, or related field; 3+ years in system operations or SRE; cloud computing, ECS, K8S, cloud architecture, and complex issue resolution experience.
ECS, ACK, ACS, OOS, Compute Nest, K8S
1w
Save
Mark Applied
Hide
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEProduction change-management experience, high-stakes release and on-call experience, AWS/GCP operations, Python proficiency, and a bachelor's degree or equivalent experience.
Python, Rust, AWS, GCP, AWS Bedrock, GCP Vertex, Claude
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, AI Infrastructure
Seattle or Bellevue
$178k-$342k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
3+ YOEBachelor's degree in computer science, software engineering, or related field; 3+ years of SRE, DevOps, or systems engineering experience; Linux, networking, distributed systems, programming, and automation skills.
Linux, Go, Python, C, C++, Java, Bash, Kubernetes, AWS, GCP, Azure, Terraform, Prometheus, Grafana
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux