Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years in LLM/VLM/Audio LM, deep learning, and related fields; publications in top venues; real-time generative models experience; strong Python and ML framework skills; dataset curation experience.
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in computer science or related field; proficiency in Golang, Java, C++, or Python; software development experience; and knowledge of databases, networks, operating systems, and distributed systems.
Golang, Java, C++, Python, Kubernetes, Docker, Istio, Envoy, Service Mesh, Function Calling, MCP
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent experience in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership; strong C++ or Python; expertise in LLM/VLM/inference and production-quality software.
TensorRT LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership, strong C++ or Python skills, expertise in LLM/VLM/inference and delivering production-quality software.
TensorRT LLM, TensorRT-LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
4+ YOEMaster's in CS/EE or related, 4+ years AI research experience (LLM/Transformers), strong deep learning background, Python and PyTorch skills, experience with LLM inference or on-device deployment.
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
Research Scientist Intern, Monetization Generative AI - LLM (PhD)
Bellevue or Menlo Park or Seattle or New York
$8k-$12k/moHybridInternship
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
Ph.D. in CS/AI/NLP/Speech/CV or related field; work authorization; experience in Python/C/C++/Java; experience with PyTorch or TensorFlow; ML/NLP system building.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
San Diego, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Senior hands-on individual contributor leading automated LLM model evaluation, building regression-detection infrastructure, and partnering with modeling, framework, and infrastructure teams.
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
8+ YOE8+ years building and launching AI/LLM products; knowledge of ElasticSearch, LangChain/LangGraph/OpenLLM, SFT, RLHF, RAG, Agents, MLOps, cloud (AWS); strong collaboration and product leadership skills.
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
Senior Research Engineer, LLM Training & Post-Training
New York City or San Francisco or Seattle or London
$165k-$310k/yrHybridFull Time
Lightning AI: Unified platform to build, train, and deploy AI models.
Requires significant PyTorch LLM training experience, distributed multi-GPU systems expertise, Python software engineering, experiment design, and a master's degree, PhD, or equivalent experience in a related field.