11 inference optimization engineer jobs at 6 companies in Austin, TX

4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Austin, Texas, United States
HybridFull Time
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
Experience productionizing ML models with inference optimization, benchmarking, deployment, and reliability; proficiency in Python and modern ML frameworks; experience with GPU/accelerator optimization and distributed systems.
Python, PyTorch, TensorFlow, JAX, SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp
2w
Save
Mark Applied
Hide
Research Kernel Engineer
Singapore or Austin
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
Requires a degree in computer science, electrical engineering, or related field; CUDA or Triton, Python, C++, GPU architecture, profiling, inference optimization, and high-performance computing experience.
CUDA, Triton, Python, C++, Nsight Compute, Nsight Systems, vLLM, SGLang, TensorRT-LLM, MLIR, TVM, TorchInductor
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years software engineering experience with GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL
1w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Causal & Decision Systems
Austin, Texas, United States
RemoteFull Time
CSC Generation
CSC Generation: Acquires and transforms retail brands into digital-first businesses.
Exceptional technical ability and judgment with experience in machine learning, statistical modeling, causal inference, decision systems, optimization, production ML, Python, SQL, and behavioral datasets.
Python, SQL
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
3mo
Save
Mark Applied
Hide
Senior Deep Learning Communication Architect
Santa Clara or Austin
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEPhD/Masters/BS in CS/EE or equivalent; 6+ years building/scaling DNNs and optimizing LLM training/inference; strong C++ and Python skills; experience with PyTorch, CUDA/OpenCL, communication libraries (MPI, NCCL, UCX, UCC, NVSHMEM) and high-speed interconnects.
NVLink, InfiniBand, SPC-X, MPI, NCCL, UCX, UCC, NVSHMEM, PyTorch, TensorRT-LLM, vLLM, SGLang, C++, Python, CUDA, OpenCL, Dynamo, Triton, FSDP
2w
Save
Mark Applied
Hide
Research Scientist-Model Efficiency
Singapore or Austin
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
Bachelor's, Master's, or PhD in computer science, electrical engineering, or related field; Python and PyTorch expertise; hands-on LLM inference, model optimization, or ML systems experience.
Python, PyTorch, C++, CUDA, Triton, vLLM, SGLang, TensorRT-LLM