Job Description
• Lead the architecture exploration and performance analysis for next-generation AI server systems to define optimal and competitive solutions.
• Perform in-depth analysis of key AI workloads, such as Large Language Models (LLMs), to accurately identify and pinpoint system performance bottlenecks.
• Develop and maintain system-level performance models to simulate various combinations of AI accelerators, memory systems, and interconnect topologies.
• Collaborate closely with system architects to deliver quantitative analysis reports based on simulation data, driving key architectural decisions.
#LI-JK1
Main Requirements and Qualifications
- • Computer Architecture Knowledge: Solid understanding of Computer Architecture, particularly in AI Accelerators (NPU/GPU/ASIC), Memory Subsystems, and System Interconnects.
- • Programming Skills: Proficient in Python. Experience with modeling frameworks or data analysis libraries is a plus.
- • Modeling & Analysis Experience: Experience in system-level modeling, performance analysis, or architectural simulation is a plus.
- • Communication: Excellent communication skills to articulate complex modeling results and technical concepts to cross-functional teams and decision-makers.
- • (Preferred Plus) AI Deployment Experience: Hands-on experience in deploying or optimizing large-scale network models (e.g., Transformer, LLM) on real-world systems is a strong plus.