System Failure Analysis Engineer (GPU Servers / Data Center)
Austin, Texas, United States
$85k-$145k/yrOnsiteFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Senior failure analysis engineer with practical experience debugging GPU-accelerated servers in rack/data center environments; adept with BIOS/BMC/IPMI; capable of AI-assisted debugging and cross-functional collaboration.
oscilloscopes, logic analyzers, protocol analyzers, power analyzers, BIOS/UEFI debugging tools
Chicago or Detroit or Minneapolis or St. Louis or Kansas City or Omaha or Columbus or Tulsa or Nashville or Austin
$185k-$235k/yrFieldFull Time
World Wide Technology: Global technology solutions provider and systems integrator.
10+ YOE10+ years in technical pre-sales/solutions architecture; hands-on AI/ML infrastructure design across GPU, storage, networking and MLOps; NVIDIA and cloud platform experience; bachelor's degree required; strong presentation and whiteboarding skills.
NVIDIA DGX, NVIDIA HGX, CUDA, AI Enterprise, NeMo, Omniverse, Advanced Technology Center (ATC), AWS, Azure, GCP, Dell, HPE, Cisco, NetApp, Pure Storage, Vast Data, MLOps
Pegatron TechnologiesTWSE: 4938: Electronics manufacturing and system integration services provider.
0+ YOEB.S. or M.S. in EE/CS or related, 0-3 years in system design/board bring-up or technical support, familiarity with Linux, Python, BIOS/BMC, GPUs/NVLink/PCIe, and ability to travel to datacenters in the Austin area.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced in system- and SoC-level HW/FW debug, data center hardware experience, RAS knowledge, ability to root-cause complex GPU/system issues, and lead junior engineers; Bachelors or Masters in electrical or computer engineering required.
Knowmadics: Develops AI-driven electronic warfare and situational awareness software platforms.
7+ YOE7-10 years in software or ML engineering; Python ML pipelines; deep learning libraries (e.g., PyTorch); C or systems programming; ETL with Kafka/Spark; ML for unstructured data; GPUs/CUDA; production model deployment; US citizenship and eligible for clearance.
EmersonNYSE: EMR: Engineering industrial automation and software solutions for global industries.
10+ YOEBS in EE/CE/CS (MS/PhD preferred), 10+ years in system/platform/compute architecture, deep knowledge of heterogeneous computing (CPUs, GPUs, FPGAs, accelerators), HW/SW co-design experience, programming in C/C++, Python, CUDA/OpenCL/SYCL or HDL, US work authorization required.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Principal-level mechanical engineering experience in server/GPU systems, expert CAD (Creo or SolidWorks), DFM and manufacturing knowledge, cross-functional leadership, and strong communication skills.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Leadership of validation architecture for AI platforms; hands-on when needed; manages a small team of senior architects; cross-domain collaboration across CPU, GPU, networking, firmware, and software.
Scripting, Automation, Test Content Development, CPU Architecture Validation, GPU Architecture Validation
Austin or Dallas or Salt Lake City or Fremont or Dacula or Pittsburgh
$140k-$244k/yrHybridFull Time
WescoNYSE: WCC: Distributes electrical and industrial products and supply chain solutions.
13+ YOEBachelor's required; 13+ years in HPC/data center/infrastructure; experience with AI/GPU systems and vendor ecosystems; leadership and pre-sales experience; proficiency with Microsoft Office; travel up to 25%.
Microsoft Office, NVIDIA, Dell, Lenovo, Supermicro, Broadcom, Qualcomm, Intel
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years leading large cross-cutting programs, technical background in cloud architecture, knowledge of data center GPU architecture and AI workloads, strong leadership, negotiation, analytical, and communication skills.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years in operational leadership within hyperscale data center or cloud environments; experience in infrastructure financial management, cost optimization, AI/GPU infrastructure, executive-level communication; MBA or Master’s preferred.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
On-site data center manager to lead deployment and availability of large GPU platforms (500–1000+ systems); manage engineers/technicians, drive root-cause debugging, and present to executive stakeholders. Bachelors in engineering required.
Senior GenAI & High Performance Computing (HPC) Delivery Engineer
Round Rock or Austin or Chicago or Illinois or Florida or United States
$145k-$199k/yrFieldFull Time
Dell TechnologiesNYSE: DELL: Provides information technology hardware, software, and services.
7+ YOE7+ years with HPC/GenAI clusters and GPU systems; deep experience with NVIDIA Base Command Manager, benchmarking tools (HPL, STREAM, NCCL, RCCL, OSU), RH distros or RHCSA/RHCE, InfiniBand/RoCE, Linux at scale, containers/orchestration, and customer-facing deployments with up to 70% travel.
Dell TechnologiesNYSE: DELL: Provides information technology hardware, software, and services.
5+ YOEProficiency in hardware design and PCB debugging; experience with PCIe, DDR, SAS or Ethernet; BS in Electrical/Computer Engineering preferred; typically 5+ years engineering experience.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOEPeople manager with 10+ years experience in software engineering, strong software architecture background, experience with RDMA/RoCE network fabrics and cloud infrastructure, and proven leadership of engineering teams.
Pegatron TechnologiesTWSE: 4938: Electronics manufacturing and system integration services provider.
0+ YOEBachelor's in CE/EE or related, 0–3 years hardware/server testing experience, proficiency in Python/Bash/C++, Linux, knowledge of PCIe, Ethernet/InfiniBand, test scripting, data analysis, and using ticketing systems like Jira.