34 inference engineer jobs at 16 companies in Texas

1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
2mo
Save
Mark Applied
Hide
Lead ML Inference Engineer, Advertising
San Jose or Austin
$247k-$486k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE5+ MgmtLead the design and development of a state-of-the-art inference platform; 10+ years in distributed systems; ML serving; leadership experience.
High-performance languages, ML frameworks, GPU acceleration, HPC, Distributed systems, Inference platforms, Monitoring tooling
1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
4w
Save
Mark Applied
Hide
Staff Backend Engineer, ML Inference Systems
Mountain View or Montreal or Texas or New York
$245k-$318k/yr RemoteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years building and operating distributed systems; expertise in Golang, cloud (GCP), Kubernetes, monitoring with Prometheus/Grafana; experience with high-throughput, low-latency inference systems.
Golang, GCP, Kubernetes, Prometheus, Grafana, Docker, NVIDIA Triton Inference Server
5h
Save
Mark Applied
Hide
GPU Engineer
Houston or California
OnsiteFull Time
Bot Auto
Bot Auto: Operates a fleet of autonomous trucks for freight transportation.
3+ YOEBachelor’s or Master’s in computer science, electrical engineering, or related field; strong GPU and parallel computing knowledge; C/C++ and Python proficiency; GPU profiling, neural inference, and embedded systems experience.
CUDA, NVIDIA Nsight Systems, NVIDIA Nsight Compute, PyTorch, ONNX, TensorRT, C, C++, Python, OpenCL, Vulkan, NVIDIA Jetson Thor, NVIDIA DRIVE Thor, FP8, NVFP4, NVIDIA Multi-Process Service (MPS), Multi-Instance GPU (MIG)
4w
Save
Mark Applied
Hide
Staff AI Infrastructure Engineer
Austin or Reston
HybridFull Time
Seekr
Seekr: Transparent AI platform for enterprise and government decision-making.
8+ YOE8+ years building distributed systems and cloud-native AI infrastructure; strong Python and systems programming skills; Kubernetes, GPU inference, and platform engineering experience; leadership and architecture experience.
Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, Infrastructure-as-Code, GitOps, CI/CD, AWS, Azure, Oracle Cloud Infrastructure, Google Cloud Platform
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Austin, Texas, United States
HybridFull Time
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
Experience productionizing ML models with inference optimization, benchmarking, deployment, and reliability; proficiency in Python and modern ML frameworks; experience with GPU/accelerator optimization and distributed systems.
Python, PyTorch, TensorFlow, JAX, SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp
5d
Save
Mark Applied
Hide
Embedded Software Senior Engineer
Irving, Texas, United States
$113k-$169k/yr OnsiteFull Time
Caterpillar
CaterpillarNYSE: CAT: Manufactures construction and mining equipment, engines, and gas turbines.
5+ YOEDeep experience with Linux-based embedded systems, C++, AI toolchains (CUDA, TensorRT, DeepStream, JetPack), system-level debugging, and embedded AI inference at the edge.
Linux, C++, CUDA, TensorRT, DeepStream, JetPack
3w
Save
Mark Applied
Hide
Embedded Software Senior Engineer
Irving, Texas, United States
$113k-$169k/yr OnsiteFull Time
Caterpillar
CaterpillarNYSE: CAT: Manufactures construction equipment, mining machinery, and industrial engines.
5+ YOEExpertise in Linux-based embedded systems, C++, AI toolchains and inference optimization; experience with CUDA/TensorRT/DeepStream/JetPack; strong debugging, validation, and system-integration skills.
Linux, C++, CUDA, TensorRT, DeepStream, JetPack
1w
Save
Mark Applied
Hide
Senior Field Application Engineer – AI
Austin, Texas, United States
$170k-$292k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Established AI background with hands-on training/inference on GPUs, experience with Pytorch/Tensorflow/JAX, Linux administration, strong communication, willingness to travel ~10-20%, bachelor\u0002s degree in technical field preferred; must be US-work-authorized.
Pytorch, Tensorflow, JAX, MLperf, Hugging Face, KVM, Kubernetes, OpenStack, OpenShift, HIP, CUDA, Python, C/C++, Fortran, OpenACC, OpenMP, JIRA
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure)
New York or San Francisco or McLean or Cambridge or San Jose or Plano
$209k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE +6 years or Master's +4 years; 6+ years programming with Python/Go/Scala/Java; experience deploying scalable AI on cloud; LLM, inference, similarity search, VectorDBs, guardrails, model evaluation, and optimization experience; leadership and research literacy.
Python, Go, Scala, Java, C++, C#, Golang, AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, AWS, Google Cloud, Azure
1w
Save
Mark Applied
Hide
Principal Engineer
Eagan or Frisco or New York City or Toronto or Ann Arbor
$159k-$295k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
Expertise in cloud cost and inference optimization, FinOps practices, production LLM cost tuning, Python proficiency, and cross-team leadership; Bachelor's in CS/CE or equivalent required.
AWS, Terraform, CloudFormation, Python, FastAPI, Kubernetes
2mo
Save
Mark Applied
Hide
Software Developer 4
Santa Clara or Seattle or New York or Austin or Nashville or United States
$100k-$235k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
7+ YOERequires 7+ years building software systems and AI applications, strong Python and ML framework experience (PyTorch/TensorFlow), experience with LLMs, data engineering (Spark/Kafka/Flink/OCI), distributed training/inference, and network automation tools.
Python, Go, PyTorch, TensorFlow, Spark, Kafka, Flink, OCI Streaming/Data Flow, NetFlow, Terraform, Ansible, NAPALM, Batfish, CI/CD
1mo
Save
Mark Applied
Hide
Gen AI developer
Irving, Texas, United States
OnsiteFull Time
Virtusa
Virtusa: Global provider of digital engineering and IT outsourcing services.
7+ YOEBachelor's degree, 7+ years building scalable AI/ML systems; deep Python and data library expertise (NumPy, Pandas); experience with LLMs/GenAI, ML training/inference/monitoring, distributed systems, data engineering, and technical leadership.
Python, NumPy, Pandas, LLMs
2mo
Save
Mark Applied
Hide
Director, Machine Learning Engineering
Palo Alto or New York City or Dallas or Bethesda or Seattle
$150k-$300k/yr OnsiteFull Time
GEICO
GEICO: Provides vehicle and property insurance services to consumers.
10+ YOE10+ years engineering or AI/ML leadership experience; proven track record building scalable distributed systems and personalization platforms; expertise in RAG, context/memory architectures, and real-time inference; strong business acumen and cross-functional leadership.
Retrieval-Augmented Generation (RAG), LLMs
2mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
3w
Save
Mark Applied
Hide
AI Systems Architect (Models & Hardware Co-Design)
Santa Clara or Austin or Boston
$200k-$500k/yr OnsiteFull Time
Velaura AI
Velaura AI: Developing ultra-low-power silicon and IP for AI accelerators.
Deep knowledge of modern ML architectures (e.g., transformers), strong mathematical foundations, experience with ML training/inference, and familiarity with ML frameworks and hardware co-design.
PyTorch, JAX, TensorFlow
1mo
Save
Mark Applied
Hide
Director of Product Management – AI Essentials
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yr HybridFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides global edge-to-cloud technology solutions and IT infrastructure services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product management experience with 5+ years AI/ML product leadership; experience building AI/ML platforms, inference/model serving, GPU ecosystem, executive communication, and GTM strategy.
GreenLake, GenAI, GPU, agent frameworks, AI tools, MLOps, DevOps, SaaS