15 inference optimization engineer jobs at 7 companies in Texas

4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEMaster's/PhD or equivalent experience, 5+ years software development, strong C/C++ skills, GPU/CUDA experience, DL inference optimization and production deployment experience.
CUTLASS, OAI TRITON, NCCL, CUDA, Python, C, C++, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Austin, Texas, United States
HybridFull Time
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
Experience productionizing ML models with inference optimization, benchmarking, deployment, and reliability; proficiency in Python and modern ML frameworks; experience with GPU/accelerator optimization and distributed systems.
Python, PyTorch, TensorFlow, JAX, SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp
2w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
2w
Save
Mark Applied
Hide
Research Kernel Engineer
Singapore or Austin
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
Requires a degree in computer science, electrical engineering, or related field; CUDA or Triton, Python, C++, GPU architecture, profiling, inference optimization, and high-performance computing experience.
CUDA, Triton, Python, C++, Nsight Compute, Nsight Systems, vLLM, SGLang, TensorRT-LLM, MLIR, TVM, TorchInductor
2mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure)
New York or San Francisco or McLean or Cambridge or San Jose or Plano
$209k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE +6 years or Master's +4 years; 6+ years programming with Python/Go/Scala/Java; experience deploying scalable AI on cloud; LLM, inference, similarity search, VectorDBs, guardrails, model evaluation, and optimization experience; leadership and research literacy.
Python, Go, Scala, Java, C++, C#, Golang, AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, AWS, Google Cloud, Azure
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years software engineering experience with GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL
1w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Causal & Decision Systems
Austin, Texas, United States
RemoteFull Time
CSC Generation
CSC Generation: Acquires and transforms retail brands into digital-first businesses.
Exceptional technical ability and judgment with experience in machine learning, statistical modeling, causal inference, decision systems, optimization, production ML, Python, SQL, and behavioral datasets.
Python, SQL
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
3mo
Save
Mark Applied
Hide
Senior Deep Learning Communication Architect
Santa Clara or Austin
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEPhD/Masters/BS in CS/EE or equivalent; 6+ years building/scaling DNNs and optimizing LLM training/inference; strong C++ and Python skills; experience with PyTorch, CUDA/OpenCL, communication libraries (MPI, NCCL, UCX, UCC, NVSHMEM) and high-speed interconnects.
NVLink, InfiniBand, SPC-X, MPI, NCCL, UCX, UCC, NVSHMEM, PyTorch, TensorRT-LLM, vLLM, SGLang, C++, Python, CUDA, OpenCL, Dynamo, Triton, FSDP
2w
Save
Mark Applied
Hide
Research Scientist-Model Efficiency
Singapore or Austin
OnsiteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
Bachelor's, Master's, or PhD in computer science, electrical engineering, or related field; Python and PyTorch expertise; hands-on LLM inference, model optimization, or ML systems experience.
Python, PyTorch, C++, CUDA, Triton, vLLM, SGLang, TensorRT-LLM