21 inference engineer jobs at 8 companies in Austin, TX

1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
2mo
Save
Mark Applied
Hide
Lead ML Inference Engineer, Advertising
San Jose or Austin
$247k-$486k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE5+ MgmtLead the design and development of a state-of-the-art inference platform; 10+ years in distributed systems; ML serving; leadership experience.
High-performance languages, ML frameworks, GPU acceleration, HPC, Distributed systems, Inference platforms, Monitoring tooling
1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
4w
Save
Mark Applied
Hide
Staff AI Infrastructure Engineer
Austin or Reston
HybridFull Time
Seekr
Seekr: Transparent AI platform for enterprise and government decision-making.
8+ YOE8+ years building distributed systems and cloud-native AI infrastructure; strong Python and systems programming skills; Kubernetes, GPU inference, and platform engineering experience; leadership and architecture experience.
Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, Infrastructure-as-Code, GitOps, CI/CD, AWS, Azure, Oracle Cloud Infrastructure, Google Cloud Platform
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Austin, Texas, United States
HybridFull Time
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
Experience productionizing ML models with inference optimization, benchmarking, deployment, and reliability; proficiency in Python and modern ML frameworks; experience with GPU/accelerator optimization and distributed systems.
Python, PyTorch, TensorFlow, JAX, SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp
1w
Save
Mark Applied
Hide
Senior Field Application Engineer – AI
Austin, Texas, United States
$170k-$292k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Established AI background with hands-on training/inference on GPUs, experience with Pytorch/Tensorflow/JAX, Linux administration, strong communication, willingness to travel ~10-20%, bachelor\u0002s degree in technical field preferred; must be US-work-authorized.
Pytorch, Tensorflow, JAX, MLperf, Hugging Face, KVM, Kubernetes, OpenStack, OpenShift, HIP, CUDA, Python, C/C++, Fortran, OpenACC, OpenMP, JIRA
2mo
Save
Mark Applied
Hide
Software Developer 4
Santa Clara or Seattle or New York or Austin or Nashville or United States
$100k-$235k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
7+ YOERequires 7+ years building software systems and AI applications, strong Python and ML framework experience (PyTorch/TensorFlow), experience with LLMs, data engineering (Spark/Kafka/Flink/OCI), distributed training/inference, and network automation tools.
Python, Go, PyTorch, TensorFlow, Spark, Kafka, Flink, OCI Streaming/Data Flow, NetFlow, Terraform, Ansible, NAPALM, Batfish, CI/CD
2mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
3w
Save
Mark Applied
Hide
AI Systems Architect (Models & Hardware Co-Design)
Santa Clara or Austin or Boston
$200k-$500k/yr OnsiteFull Time
Velaura AI
Velaura AI: Developing ultra-low-power silicon and IP for AI accelerators.
Deep knowledge of modern ML architectures (e.g., transformers), strong mathematical foundations, experience with ML training/inference, and familiarity with ML frameworks and hardware co-design.
PyTorch, JAX, TensorFlow