System Failure Analysis Engineer (GPU Servers / Data Center)
Austin, Texas, United States
$85k-$145k/yrOnsiteFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Senior failure analysis engineer with practical experience debugging GPU-accelerated servers in rack/data center environments; adept with BIOS/BMC/IPMI; capable of AI-assisted debugging and cross-functional collaboration.
oscilloscopes, logic analyzers, protocol analyzers, power analyzers, BIOS/UEFI debugging tools
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Requires hardware or AI infrastructure program leadership, GPU and datacenter expertise, cross-functional execution, strong communication, program management tools, and a bachelor's or master's degree in systems, electrical engineering, computer science, or related engineering.
Jira, Confluence, Microsoft Excel, Microsoft PowerPoint, GPU, CPU, BIOS, BMC, HSIO, FW
Pegatron TechnologiesTWSE: 4938: Electronics manufacturing and system integration services provider.
0+ YOEB.S. or M.S. in EE/CS or related, 0-3 years in system design/board bring-up or technical support, familiarity with Linux, Python, BIOS/BMC, GPUs/NVLink/PCIe, and ability to travel to datacenters in the Austin area.
Chicago or Detroit or Minneapolis or St. Louis or Kansas City or Omaha or Columbus or Tulsa or Nashville or Austin
$185k-$235k/yrFieldFull Time
World Wide Technology: Global technology solutions provider and systems integrator.
10+ YOE10+ years in technical pre-sales/solutions architecture; hands-on AI/ML infrastructure design across GPU, storage, networking and MLOps; NVIDIA and cloud platform experience; bachelor's degree required; strong presentation and whiteboarding skills.
NVIDIA DGX, NVIDIA HGX, CUDA, AI Enterprise, NeMo, Omniverse, Advanced Technology Center (ATC), AWS, Azure, GCP, Dell, HPE, Cisco, NetApp, Pure Storage, Vast Data, MLOps
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced in system- and SoC-level HW/FW debug, data center hardware experience, RAS knowledge, ability to root-cause complex GPU/system issues, and lead junior engineers; Bachelors or Masters in electrical or computer engineering required.
Knowmadics: Develops AI-driven electronic warfare and situational awareness software platforms.
7+ YOE7-10 years in software or ML engineering; Python ML pipelines; deep learning libraries (e.g., PyTorch); C or systems programming; ETL with Kafka/Spark; ML for unstructured data; GPUs/CUDA; production model deployment; US citizenship and eligible for clearance.
EmersonNYSE: EMR: Engineering industrial automation and software solutions for global industries.
10+ YOEBS in EE/CE/CS (MS/PhD preferred), 10+ years in system/platform/compute architecture, deep knowledge of heterogeneous computing (CPUs, GPUs, FPGAs, accelerators), HW/SW co-design experience, programming in C/C++, Python, CUDA/OpenCL/SYCL or HDL, US work authorization required.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Principal-level mechanical engineering experience in server/GPU systems, expert CAD (Creo or SolidWorks), DFM and manufacturing knowledge, cross-functional leadership, and strong communication skills.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years engineering experience leading teams to design, scale, and operate large distributed GPU/cloud infrastructures; expertise in reliability, security, IaC, telemetry, and incident response.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Leadership of validation architecture for AI platforms; hands-on when needed; manages a small team of senior architects; cross-domain collaboration across CPU, GPU, networking, firmware, and software.
Scripting, Automation, Test Content Development, CPU Architecture Validation, GPU Architecture Validation
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years leading large cross-cutting programs, strong technology background in cloud and AI infrastructure, strategic thinking, leadership, negotiation, and knowledge of Data Center GPU architecture.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years leading large cross-cutting programs with strategic thinking, technology background in cloud/GPU/AI/HPC, leadership, negotiation, and ability to drive alignment across senior stakeholders.
Austin or Dallas or Salt Lake City or Fremont or Dacula or Pittsburgh
$140k-$244k/yrHybridFull Time
WescoNYSE: WCC: Distributes electrical and industrial products and supply chain solutions.
13+ YOEBachelor's required; 13+ years in HPC/data center/infrastructure; experience with AI/GPU systems and vendor ecosystems; leadership and pre-sales experience; proficiency with Microsoft Office; travel up to 25%.
Microsoft Office, NVIDIA, Dell, Lenovo, Supermicro, Broadcom, Qualcomm, Intel
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
On-site data center manager to lead deployment and availability of large GPU platforms (500–1000+ systems); manage engineers/technicians, drive root-cause debugging, and present to executive stakeholders. Bachelors in engineering required.
Dell TechnologiesNYSE: DELL: Provides information technology hardware, software, and services.
5+ YOEProficiency in hardware design and PCB debugging; experience with PCIe, DDR, SAS or Ethernet; BS in Electrical/Computer Engineering preferred; typically 5+ years engineering experience.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOEPeople manager with 10+ years experience in software engineering, strong software architecture background, experience with RDMA/RoCE network fabrics and cloud infrastructure, and proven leadership of engineering teams.
Pegatron TechnologiesTWSE: 4938: Electronics manufacturing and system integration services provider.
0+ YOEBachelor's in CE/EE or related, 0–3 years hardware/server testing experience, proficiency in Python/Bash/C++, Linux, knowledge of PCIe, Ethernet/InfiniBand, test scripting, data analysis, and using ticketing systems like Jira.