AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Preferred bachelor's degree in a technical field and master's degree. Requires C++ development, functional modeling, virtual platforms, validation, distributed systems, computer architecture, and interconnect technology experience.
C++, CI/CD, PCIe, CXL, AXI, CHI, ACE, UCIe, HBM, DDR, SystemC/TLM-2.0, Arm Fast Models, Simics, QEMU
NetAppNasdaq: NTAP: Provides intelligent data infrastructure for hybrid cloud environments.
15+ YOEExpert in AI inferencing and distributed systems at scale with 15+ years experience; hands-on with inference engines, model optimization, GPU/TPU orchestration, Kubernetes, RDMA/DPDK; strong architecture, communication, and mentorship skills.
Distinguished Engineer - AI (San Jose, CA, US, 95128)
San Jose, California, United States
$266k-$396k/yrOnsiteFull Time
NetAppNASDAQ: NTAP: Sells enterprise data storage and cloud management software.
15+ YOE15+ years building low-latency, fault-tolerant distributed systems and AI/ML inference platforms; expertise with inference engines, model optimization, storage for AI, RDMA/DPDK, and Kubernetes-based orchestration.
San Francisco or New York City or San Jose or Seattle or Austin or Boston
$115k-$200k/yrHybridFull Time
Absentia Labs: AI-native toxicology platform accelerating drug safety and discovery.
5+ YOE5+ years ML industry experience, proven production-scale model training, expertise with LLMs/diffusion/GNNs, strong PyTorch skills, distributed training and data-pipeline experience, and solid software engineering practices.
Distributed Systems Engineer 4 - Content & Business Products
Los Gatos or United States
$250k-$413k/yrRemoteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
2+ YOE2+ years working on distributed systems; proficiency in Java or C# and OO design; experience with multithreading, microservices, data modeling, API design; participate in on-call rotation and lead incident reviews.
Tech Lead Software Engineer - AI Compute Infrastructure
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOE5+ years experience building cloud/ML infrastructure, strong knowledge of large-model inference, distributed systems, scheduling, and container orchestration; proficiency in Go/Rust/Python/C++.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
15+ YOE15+ years systems and architecture experience, 10+ years in distributed computing, expertise in IAM, cryptography, network/cloud/application security, threat modeling, and strong communication; bachelor’s or equivalent.
Machine Learning Engineer, TikTok - Business Governance
San Jose, California, United States
$156k-$388k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
Strong ML/DL knowledge with Transformer/LLM familiarity, hands-on Python and PyTorch, distributed training and large-scale data processing, experience productionizing models for content safety and cross-team collaboration.
Tessera Labs: Automates complex enterprise workflows with multi-agent AI systems.
Significant language-model training or post-training experience, RL tuning, Python and PyTorch or JAX proficiency, distributed GPU training, empirical experimentation, and strong software engineering and writing skills.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
Machine Learning Engineer - Ads Core and Commerce Ads
San Jose or Los Angeles County
$137k-$360k/yrOnsiteFull Time
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOERequires SQL and Python, data manipulation, Hadoop or Spark, distributed computing, machine learning, deep learning, feature engineering, model evaluation, optimization, and 5+ years of relevant experience preferred.
Distributed Systems Engineer 4 - Content & Business Products
Los Gatos, California, United States
$250k-$413k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global video streaming and media production service.
2+ YOE2+ years in distributed systems; proficient in Java or C#; strong knowledge of multithreading, observability; experience with microservices, data modeling, API design; on-call experience; good cross-functional communication.
Java, C#, OO design, Microservices, API design, gRPC, GraphQL, observability
Austin or San Jose or United States or Canada or Mexico
HybridFull Time
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
5+ YOE5+ years applied ML experience; strong software development skills in Spark, Python, or Java; experience with distributed ML frameworks, low-latency model evaluation, experimentation, statistics, and large-scale data systems.
DiDi GlobalOTC Markets: DIDIY: Global technology platform providing mobility, delivery, and financial services.
3+ YOEMaster's degree in a relevant field, 3+ years of deep learning R&D experience, pretrained model development, Transformer expertise, distributed training, ablation studies, and strong model engineering skills.
Software Development Manager, AWS Neuron SDK - Distributed Training
Cupertino, California, United States
$213k-$288k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtExperience with PyTorch or JAX, distributed training at scale, 7+ years engineering experience, 3+ years engineering team management, 3+ years designing/architecting systems, and partnering with product teams; deep learning model training experience.