4 gpu inference performance engineer jobs at 4 companies in Austin, TX

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
2w
Save
Mark Applied
Hide
Research Kernel Engineer
Singapore or Austin
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
Requires a degree in computer science, electrical engineering, or related field; CUDA or Triton, Python, C++, GPU architecture, profiling, inference optimization, and high-performance computing experience.
CUDA, Triton, Python, C++, Nsight Compute, Nsight Systems, vLLM, SGLang, TensorRT-LLM, MLIR, TVM, TorchInductor
2mo
Save
Mark Applied
Hide
Senior Deep Learning Framework Communications Engineer
Santa Clara or Austin or Westford or Durham or United States
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years software engineering experience in HPC/AI, experience with PyTorch/JAX and inference engines, Python/C++/CUDA development, performance benchmarking and profilers, understanding of multi-GPU communication and compilers.
PyTorch, TRT-LLM, vLLM, SGLang, JAX, NCCL, NVSHMEM, GPUDirect, torch.compile, PyTorch profiler, NVIDIA Nsight Systems, Python, C++, CUDA, Triton, cuTe, MPI
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python