ByteDance
Posted 1mo ago

Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)

ByteDance
Singapore
OnsiteFull Time
Responsibilities
  • optimizing inference
  • developing operators
  • designing parallelism
Requirements
  • Bachelor's in Computer Science or equivalent,3+ years experience,proficient in C/C++ and Python,CUDA and GPU architecture knowledge,deep learning operator and inference optimization experience
Technical tools mentioned
C/C++PythonCUDANsightProfilervLLMTensorRT-LLMSGLang

Job description

About The Team
The mission of our AML team is to push the next-generation AI infrastructure and recommendation platform for the ads ranking, search ranking, live & e-Commerce ranking in our company. We also drive substantial impact on core businesses of the company.

Responsibilities
- Responsible for the iteration of the underlying architecture of the large model inference engine and end-to-end GPU performance optimization, through means such as operator fusion and compilation optimization, deeply optimizing GPU memory access, computing pipeline, and Stream asynchronous scheduling, eliminating inference computing bottlenecks, improving single-card inference throughput, and reducing inference latency.
- Adapt to all series of GPU/NPU hardware architectures, refine the universality of the inference engine and hardware adaptability, and build a high-performance, low-loss underlying base for large model inference.
- Lead the design, development, and optimization of distributed parallel solutions for large model inference scenarios, with a focus on implementing multi-dimensional parallel strategies such as tensor parallelism (TP), pipeline parallelism (PP), sequence parallelism, and MoE expert parallelism, to address core issues such as multi-card splitting and deployment of ultra-large models, high cross-card communication overhead, load imbalance, and low parallel efficiency.
- Follow up on cutting-edge technologies such as global large model inference, GPU high-performance computing, distributed parallelism, and cache optimization, benchmark against mainstream inference frameworks such as vLLM and TensorRT-LLM, complete the implementation of solutions and technological innovation, continuously iterate and optimize the performance and cost advantages of the inference system, and build the core technological barriers of the team.

Minimum Qualification(s)
- Bachelor’s degree in Computer Science or equivalent with 3+ years of relevant experience
- Solid foundation in computer low-level knowledge, proficient in C/C++ and Python programming, skilled in CUDA programming and familiar with GPU hardware architecture principles, and well-versed in GPU memory models, computing scheduling, and communication mechanisms;
- Proficiently master the underlying development and implementation of various basic operators in Deep learning, be well-versed in GPU adaptation and optimization of core operators such as matrix operations, normalization, and activation functions, and be able to independently complete operator handwritten reconstruction, memory access optimization, vectorization acceleration, and precision alignment to ensure high performance and high stability of operator inference.
- Familiar with the end-to-end process of deep learning inference compilation, understand core compilation technologies such as computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation, and be able to simplify the inference process, reduce GPU memory usage, and decrease inference latency through compilation-level improvements, thereby significantly enhancing the throughput efficiency of model inference.
- Proficient in using GPU performance analysis tools such as Nsight and Profiler, able to accurately identify performance bottlenecks such as computing power waste, memory access blockage, and scheduling redundancy during the inference process, possess the thinking of software-hardware collaborative optimization, capable of outputting systematic optimization solutions and completing implementation iterations, and adaptable to the requirements of industrial-level high-concurrency, low-latency inference business.

Preferred Qualification(s)
- Thoroughly understand the core principles of large model inference, proficiently master the core technologies of model parallelism, have experience in implementing distributed inference solutions such as tensor parallelism, pipeline parallelism, and sequence parallelism, and be familiar with multi-card communication, load balance, and parallel efficiency optimization methods.
- Those with experience in secondary development and Performance optimization of mainstream large model inference frameworks such as vLLM, SGLang, TensorRT-LLM, etc. are preferred.
- Familiarity with model computation efficiency optimization solutions for mainstream deep learning frameworks.

About ByteDance

Developing AI-driven content platforms and mobile applications.

Similar jobs

Backend Engineer roles
1d
Save
Mark Applied
Hide
Senior Python Backend Engineer (LLM/RAG)
Singapore
HybridFull Time
EPAM Systems
EPAM SystemsNYSE: EPAM: Provides global digital platform engineering and software development services.
Production Python backend development, microservices, APIs, event-driven systems, search and retrieval, PostgreSQL, vector databases, LLM/RAG applications, and AML screening knowledge are required.
Python, PostgreSQL, Elasticsearch, OpenSearch, Pinecone, Azure AI Search
2d
Save
Mark Applied
Hide
Backend Engineer / Senior Backend Engineer
Singapore, Central Singapore, Singapore
OnsiteFull Time
PatSnap
PatSnap: AI-powered platform for intellectual property and R&D intelligence.
3+ YOEBachelor’s or Master’s degree in computer science, software engineering, or related field; 3+ years of Java backend development; Spring, RESTful APIs, microservices, databases, CI/CD, and containerization experience.
Java, Spring, Spring Boot, SQL, NoSQL, Docker, Kubernetes
5d
Save
Mark Applied
Hide
Backend Engineer - AI Solutions - A26300
Singapore, Singapore, Singapore
OnsiteContract
Activate Interactive
Activate Interactive: Provides custom application development and digital transformation services.
1+ YOEBachelor's degree in a relevant IT field and 1–2 years' experience. Requires Python, backend/API development, AI/ML, cloud, testing, monitoring, security, troubleshooting, and communication skills.
Python, AWS, GCC, SHIPS-HATS, GitLab CI/CD, CI/CD, DevSecOps, MLOps, LLMOps, GitLab, APIs
6d
Save
Mark Applied
Hide
Backend Engineer Graduate (TikTok Vertical Recommendation Architecture) - 2027 Start
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$128k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
0+ YOEBachelor's degree in computer science or related field; strong programming in Go, C++, Java, or Python; knowledge of algorithms, operating systems, networking, and distributed systems.
Go, C++, Java, Python, Redis, Kafka
1w
Save
Mark Applied
Hide
Senior Backend Engineer - Asian Timezones
Taipei City or Singapore or Hong Kong
RemoteFull Time
Hermeneutic Investments
Hermeneutic Investments: Proprietary trading firm and crypto-focused hedge fund
5+ YOERequires 5+ years backend experience, 3+ years Go experience or comparable depth, distributed systems, microservices, cloud-native applications, networking, message brokers, databases, and production operations expertise.
Go, Docker, ECS, Kubernetes, Kafka, Redpanda, PostgreSQL, Redis, ClickHouse, TimescaleDB, Rust, C++, Python, Terraform, CloudFormation, AWS, GCP, Azure, Prometheus, Grafana, OpenTelemetry, TCP, UDP, WSS, REST, gRPC, CI/CD
1w
Save
Mark Applied
Hide
Backend Engineer, Agent Development (EC AI)
Singapore, Singapore, Singapore
OnsiteFull Time
Rakuten Group
Rakuten GroupTokyo Stock Exchange: 4755: Global ecosystem of internet, financial, and mobile communication services.
Requires 5+ years building production backend services, strong Python, ML production systems, transactional backends, distributed systems, LLM systems, agentic patterns, and cloud/container deployment.
1w
Save
Mark Applied
Hide
Senior Backend Engineer — Data Platform & AI Agents
Paris or New York City or London or Singapore
HybridFull Time
Jus Mundi
Jus Mundi: AI-powered search engine for international law and arbitration.
Requires 5+ years building production backend or data services, strong Python and Postgres expertise, production agentic LLM experience, document processing, DevOps, and fluent written and spoken English.
Python, Postgres, LangGraph, LangChain, LlamaIndex, Kubernetes, Elasticsearch, OpenSearch, TypeScript, React, PDF, HTML, OCR, MCP
1w
Save
Mark Applied
Hide
服务端工程师
Singapore, Singapore, Singapore
OnsiteFull Time
Xiaomi
XiaomiHKEX: 1810: Designs and manufactures smartphones, consumer electronics, and smart devices.
Bachelor's degree or higher in computer science or a related field; strong computer and distributed systems fundamentals; proficiency in Java, Go, or Python; and cross-team collaboration skills.
Java, Go, Python