Institute of Foundation Models: Develops open-source frontier-class AI foundation models and research.
2+ YOE2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC or infrastructure operations; strong Linux troubleshooting; scripting in Python or Bash; Bachelor’s in a related field.
KLANASDAQ: KLAC: Provides process control and yield management for semiconductor manufacturing.
8+ YOEExtensive Linux systems engineering in large-scale compute; HPC schedulers (Slurm), MPI, GPUs; scripting and automation; troubleshooting; collaboration.
KLANASDAQ: KLAC: Provides process control and yield management solutions for semiconductors.
0+ YOEMaster's degree (0 years) or Bachelor's + 2 years; foundational computer architecture and Linux knowledge; experience or exposure to HPC, distributed systems, or server platforms; Python/Bash scripting; strong problem-solving and teamwork skills.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years designing or operating large-scale storage infrastructure; Linux, Python and bash proficiency; experience with containers, distributed filesystems, storage performance tuning, and HPC/AI workloads.
Lawrence Berkeley National Laboratory: Conducts multidisciplinary scientific research for the U.S. Department of Energy.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; deep HPC expertise (Linux/Unix, Fortran, C/C++, MPI, OpenMP, CUDA, Python, containers, debuggers, performance tools); strong communication and teaching skills.
Lawrence Berkeley National Laboratory: Conducts scientific research to address global challenges.
8+ YOEAdvanced experience in HPC/AI performance engineering, data management, storage and I/O, scientific software development, and collaboration with domain scientists; typically 8+ years (CSE3) or 12+ years (CSE4) of related experience with relevant degree.
Anduril Industries: Defense technology building autonomous military hardware and software.
5+ YOE5+ years administering Linux/Unix in HPC or scientific computing; hands-on cluster, storage, and scientific code support; Bash/Python scripting; experience with MPI/OpenMP, CMake, compilers, filesystems, and secure computing; bachelor's degree in a technical field; ability to obtain US security clearance.
HPC Systems Administrator (Hardware & Infrastructure Operations)
Stanford, California, United States
$150k-$172k/yrOnsiteFull Time
Stanford University: A private research university providing higher education and research.
3+ YOEManage and maintain large-scale HPC hardware and infrastructure, perform diagnostics and root-cause analysis, collaborate with data center teams, automate provisioning and monitoring; bachelor's degree plus experience required.
Linux, Slurm, Lustre, InfiniBand, DCIM, NVIDIA H200, x86
Houston or New York City or San Francisco or Seattle
$230k-$343k/yrOnsiteFull Time
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
Requires leadership of HPC or infrastructure engineering organizations, Linux, distributed systems, production operations, GPU infrastructure, networking, budgeting, and people management; Slurm experience strongly preferred.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years programming experience; BS/MS or equivalent in CS or related field; strong Fortran/C/C++ and parallel programming knowledge; experience with OpenACC, OpenMP, MPI, CUDA; strong performance analysis and communication skills.
University of California, Los Angeles: A public research university providing higher education and research.
7+ YOE7+ years managing large-scale production storage for research or hyperscale environments; deep knowledge of scale-out, parallel, distributed, object, and federated storage; hands-on experience with Lustre, VAST Data, GPFS/Spectrum Scale, Ceph, BeeGFS, WekaFS, or MinIO; advanced Linux administration; Bash/Python scripting; Ansible and Git; identity ...
Sr. High Performance Computing (HPC) Systems Engineer
Brownsville or Hawthorne
OnsiteFull Time
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years systems engineering experience, hands-on Linux and HPC cluster administration, Kubernetes, scripting (Bash/Python), networking, and security; eligible for TS/SCI with polygraph.
General Atomics: Designs and manufactures unmanned aircraft and nuclear technology systems.
15+ YOEBachelor's degree or equivalent experience, 15+ years systems administration experience, deep Linux stack knowledge (NFS, ZFS, BTRFS, mdadm, LVM), scripting with Python/Ansible/Bash, troubleshooting distributed systems; SLURM and parallel file system experience desirable.
Albany or Arlington or Atlanta or Austin or Beaverton or Bentonville or Boston or Carmel or Charlotte or Chicago or Cincinnati or Cleveland or Columbus or Culver City or Denver or Des Moines or Detroit or Hartford or Houston or Irvine or Irving or Kirkland or Miami or Milwaukee or Minneapolis or Morristown or Mountain View or Nashville or New York City or Oklahoma City or Overland Park or Philadelphia or Pittsburgh or Raleigh or Redmond or Sacramento or San Diego or San Francisco or Scottsdale or Seattle or St. Louis or St. Petersburg or Walnut Creek or United States
$80k-$266k/yrHybridFull Time
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
5+ YOERequires 5+ years designing AI infrastructure, accelerated computing, clusters, orchestration, and automation; bachelor's degree or equivalent experience. Strong Kubernetes, Python, Terraform, and cloud expertise required.
Cirrascale: Provides specialized GPU-based cloud infrastructure for AI workloads.
5+ YOEBachelors in CS/CE or equivalent; 5+ years building distributed systems and modern observability tooling (Open Telemetry, Prometheus, Grafana, Datadog, ELK, etc.); 1+ year HPE OpsRamp; Bash and Python; cloud (AWS/GCP/OpenStack) and k8s; on-call participation.