TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
Vast.ai: Decentralized marketplace for GPU cloud computing resources.
Expertise in systems and GPU engineering, GPU architectures, neural network performance, C++/CUDA/Python proficiency, and strong research background with publications preferred.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
3+ YOE3+ years capacity planning or systems engineering experience; hyperscaler cloud experience; GPU topology knowledge (NVIDIA H100/B200); Bachelor’s or Master’s in quantitative field; strong cross-functional communication and modeling skills.
YieldNest: Liquid restaking protocol for risk-adjusted DeFi yields.
Senior infrastructure/DevOps engineer with expertise in bare-metal provisioning, GPU scheduling, Terraform/Pulumi, CI/CD for infrastructure, storage for AI/ML workloads, and cloud-init provisioning.
HPNYSE: HPQ: Manufactures personal computers, printers, and 3D printing hardware.
10+ YOEDesign GPU/NPU driver architectures and memory management for shared CPU/GPU/NPU use; 10+ years experience recommended; degree in electrical or related engineering preferred; expertise in memory/GPU design, hardware architecture, debugging, and laboratory tools.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
2+ YOE2+ years software engineering experience; proficiency in Go and/or Python; production Kubernetes experience; develop performance tests, automation, and platform tooling; participate in on-call rotation.
HPNYSE: HPQ: Manufacturer of personal computers, printers, and imaging devices.
10+ YOE10+ years experience in electrical/hardware or driver architecture; expertise in GPU/memory design, OS-driver contracts (Windows/Linux), debugging, FPGA and hardware architecture; degree in electrical engineering or related preferred.
SkyPilot: Unified compute platform for orchestrating AI workloads across clouds.
Hands-on experience with GPU/accelerator systems and ML training or inference infrastructure; strong Python and systems-level skills; experience operating large-scale training or high-throughput inference.
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
10+ YOE10+ years building large-scale distributed systems; deep GPU and accelerator expertise; experience with driver/firmware lifecycle, CUDA, GPU scheduling, Kubernetes; strong Go proficiency and technical leadership.
Member of Technical Staff - GPU Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Liquid AI: Develops efficient general-purpose artificial intelligence foundation models.
Strong software engineering with production infrastructure tooling, deep distributed systems/Linux/networking/storage knowledge, experience operating shared compute clusters and supporting production users.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOEExpertise in GPU/CPU hardware and platform engineering, firmware and diagnostics (BMC, UEFI/BIOS, Linux), board-level tools, FPGA and server architectures (x86/ARM); 10+ years experience preferred; strong debugging and communication skills.
Together AI: Cloud platform for training and deploying artificial intelligence models.
Strong software engineering experience with Go, Python, or Rust; durable workflow orchestration (Temporal/Cadence); control-plane/orchestration and event-driven system experience; product mindset building internal platforms.
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years in infrastructure or systems engineering; strong Linux in production; GPU hardware and software experience; bare-metal provisioning; Python automation; debugging across hardware/OS/GPU.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE8+ years software engineering experience, strong C++ and Python skills, GPU inference experience, Linux, containers and Kubernetes, benchmarking and production optimization for latency-sensitive services.
System Software Engineer, Robot Platform — GPU & Accelerated Compute
Redwood City, California, United States
OnsiteFull Time
Sunday: Developing autonomous robots to perform household chores.
2+ YOE2+ years in GPU systems software; proficient in CUDA and a systems language (C++, C, or Rust); strong understanding of GPU architecture and time-slicing; experience with CUDA ecosystem and GPU sharing; solid Linux fundamentals.
CUDA, CUDA Graphs, CUDA IPC, Nsight Systems, Nsight Compute, NVDEC, NVENC, MPS, MIG, Linux
Toronto or San Francisco or New York City or London or Paris or Montreal
HybridFull Time
Cohere: Provides enterprise-grade large language models and AI software platforms.
Experience managing engineering teams focused on GPU/ML infrastructure, Kubernetes, IaC, observability, and collaboration with AI researchers; strong communication and mentorship skills.