This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

Lawrence Berkeley National Laboratory
Posted 1w ago

System Infrastructure / Platform Engineer, HPC Technology Department

Lawrence Berkeley National Laboratory
Berkeley, California, United States
$157k-$268k/yrHybridFull Time
Responsibilities
  • managing infrastructure
  • troubleshooting systems
  • developing automation
Requirements
  • Requires 4+ years managing large-scale Linux deployments
  • Strong Linux, Bash, and Python skills, and experience with HPC
  • Containers
  • Virtualization, cloud
  • Storage
  • Networking
  • Automation, and security
Technical tools mentioned
LinuxBashPythonDockerKubernetesProxmoxVMwareAWSAzureGCPiSCSINASLustreGPFSVASTInfiniBandSlingshotRoCEstracelsofeBPFgdbGitLabJiraChefForemanTerraformAnsiblePuppetSlurmTCP/IPEthernetBGPECMPScrumKanbanPerlmutterSpin KubernetesDoudnaCray/HPECSM/COS

Job description

The National Energy Research Scientific Computing Center (NERSC) is seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based infrastructure. NERSC operates some of the world’s largest supercomputers, supporting thousands of researchers tackling major scientific challenges.

 

In this role, you will manage high-performance computing environments, including HPC systems, containers, virtual machines, and core infrastructure services. You’ll work with cutting-edge technologies such as CPU/GPU clusters, parallel storage, high-speed networking, Slurm, and Kubernetes, balancing innovation with reliability, performance, and security at scale.

 

Collaborating with engineers, researchers, vendors, and open-source communities, you will help develop scalable solutions that advance scientific discovery and the future of HPC. If you have Linux experience, an interest in science, and enjoy fast-paced collaborative environments, NERSC would love to hear from you.

 

We’re here for the same mission, to bring science solutions to the world. Join our team and YOU will play a supporting role in our goal to address global challenges! Have a high level of impact and work for an organization associated with 17 Nobel Prizes!

 

Why join Berkeley Lab?

We invest in our employees by offering a total rewards package you can count on:

  • Exceptional health and retirement benefits, including pension or 401K-style plans

  • Opportunities to grow in your career - check out our Tuition Assistance Program

  • A culture where you’ll belong - we are invested in our teams! 

  • In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year.

  • Parental bonding leave (for both mothers and fathers)

  • Pet insurance

 

What You Will Do if hired at a Level 3:

  • Build and manage Linux systems and storage infrastructure

  • Troubleshoot complex technical issues with team members

  • Install, upgrade, and secure systems and services

  • Develop and maintain scripts and automation tools

  • Participate in a 24/7 on-call rotation

  • Lead small projects, upgrades, and service rollouts

  • Collaborate with vendors to improve technologies and user experience

  • Support reliable operations of NERSC’s Perlmutter supercomputer and Spin Kubernetes platform

  • Develop and integrate services across NERSC and DOE facilities, including the upcoming Doudna supercomputer

  • Present technical work to the HPC community at conferences and industry events

 

In Additional Responsibilities if hired at a Level 4:

  • Solve complex technical problems with independent judgment

  • Develop team strategies and project plans

  • Provide technical leadership and mentorship

  • Lead system improvements for performance, reliability, and security

  • Evaluate emerging HPC technologies and capabilities

  • Represent NERSC in HPC and DOE technical communities and advocacy groups

 

What is Required to be hired at a Level 3:

  • Typically, 8+ years of related experience with a Bachelor’s degree; alternatively, 6+ years with a Master’s degree; or equivalent career experience

  • 4+ years of experience managing large-scale Linux-based system deployments in a high-performance computing, cloud computing, or hyper-scale environment

  • Mastery of Linux concepts and operations (processes, networking, system logs, performance)

  • Proficiency with bash and Python scripting

  • Experience with some or all of our key technologies:

    • containers (such as Docker or Kubernetes)

    • virtualization (such as Proxmox or VMware)

    • cloud-based deployment (such as AWS, Azure or GCP)

    • identity and access management

    • database administration, tuning, and troubleshooting

    • storage systems technologies (such as iSCSI and NAS appliances)

    • parallel filesystems (such as Lustre, GPFS, or VAST)

    • high-speed networking/interconnect (such as InfiniBand, Slingshot, or RoCE)

    • advanced performance analysis and debugging tools (such as strace, lsof, ebpf, or gdb)

    • DevOps tools (such as Gitlab or Jira) and processes (such as issues, merge requests, and API/automation)

  • Familiarity with automated provisioning systems (such as Chef, Foreman, or Terraform)

  • Familiarity with configuration management systems (such as Ansible or Puppet)

  • Working knowledge of Linux system engineering and security practices

  • Ability to resolve complex issues in creative and effective ways and derive technical solutions in a collaborative environment to meet end user requirements or needs

  • Demonstrated ability to work independently as well as collaboratively in large projects, and contribute to an active and respectful intellectual environment

  • Creative, positive, and collaborative work style

  • Excellent oral and written communication skills

 

Additional Requirements to be hired at a Level 4

  • Typically, 12+ years of related experience with a Bachelor’s degree; alternatively, 8+ years with a Master’s degree; or equivalent career experience

  • Proven ability to lead troubleshooting and resolution of high-impact incidents in complex, large-scale environments

  • Demonstrated leadership in cross-team collaboration and mentoring

  • Experience in software engineering, Linux systems programming, or complex scripting

  • Experience managing one or more of the following:

    • data center networking (TCP/IP, Ethernet, BGP, ECMP)

    • batch workload managers (such as Slurm), including installation, configuration, routine operations, job lifecycle concepts, and troubleshooting common failure modes

    • Cray/HPE HPC ecosystems (e.g., CSM/COS, Slingshot interconnect, and related components)

  • Ability to lead and coordinate projects with traditional or Agile methodologies (such as Scrum or Kanban)

  • Ability to analyze and resolve significant and unique issues requiring evaluation of multiple intangible factors

  • Ability to exercise independent judgment in methods, techniques and evaluation criteria for obtaining results

 

Additional information:

  • Applications will be accepted until the job posting is removed.

  • Appointment type: This is a full-time, career appointment, exempt (monthly paid) from overtime pay.

  • Salary range: 

    • Level 3: The expected salary for this position is $156,864 - $191,724, which fits into the full salary of $139,440 - $235,308 depending upon the candidate’s skills, knowledge, and abilities. This includes education, certifications, and years of experience.

    • Level 4: The expected salary for this position is $178,644 - $218,364, which fits into the full salary of $158,808 - $267,996 depending upon the candidate’s skills, knowledge, and abilities. This includes education, certifications, and years of experience.

  • Background check: This position is subject to a background check. Any convictions will be evaluated to determine if they directly relate to the responsibilities and requirements of the position. Having a conviction history will not automatically disqualify an applicant from being considered for employment.

  • Work modality: This position requires substantial on-site presence, but is eligible for a flexible work mode, and hybrid schedules may be considered. Hybrid work is a combination of performing work on-site at Lawrence Berkeley National Lab, 1 Cyclotron Road, Berkeley, CA and some telework. Individuals working a hybrid schedule must reside within 150 miles of Berkeley Lab. Work schedules are dependent on business needs. In rare cases, full-time telework or remote work modes may be considered.

  • Multi-level Posting: This position will be hired at a level commensurate with the business needs and the skills, knowledge, and abilities of the successful candidate.

  • Export Control Access: This position will involve access to hardware, commodities, and technical information subject to export control regulations including, but not limited to, the Export Administration Regulations ("EAR") and/or International Traffic in Arms Regulations ("ITAR"). Accordingly, any hiring decision may depend in part on Berkeley Lab’s ability to obtain or rely on federal government authorizations as required, if you are not a U.S. citizen, lawful permanent resident of the U.S. (“green card holder”), asylee, refugee, or other qualifying protected individual as defined by 8 U.S.C. 1324b(a)(3).

 

Want to learn more about working at Berkeley Lab? Please visit: careers.lbl.gov

 

Equal Employment Opportunity Employer: The foundation of Berkeley Lab is our Stewardship Values: Team Science, Service, Trust, Innovation, and Respect; and we strive to build community with these shared values and commitments. Berkeley Lab is an Equal Opportunity Employer. We heartily welcome applications from all who could contribute to the Lab's mission of leading scientific discovery, excellence, and professionalism. In support of our rich global community, all qualified applicants will be considered for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected categories under State and Federal law.

 

Misconduct Disclosure Requirement: As a condition of employment, the finalist will be required to disclose if they are subject to any final administrative or judicial decisions within the last seven years determining that they committed any misconduct, are currently being investigated for misconduct, left a position during an investigation for alleged misconduct, or have filed an appeal with a previous employer.






 






















About Lawrence Berkeley National Laboratory

Conducts multidisciplinary scientific research for the U.S. Department of Energy.

Year founded
1931
Employees
4500
Organization type
Government
Headquarters
US

Similar jobs

Platform Engineer roles near Berkeley, California
4h
Save
Mark Applied
Hide
Software Engineer — Platform
New York City or San Francisco
$220k-$300k/yr HybridFull Time
Snorkel AI
Snorkel AI: Software platform for programmatic AI data development and labeling.
5+ YOERequires 5+ years building platform infrastructure, backend services, or production data systems; strong Python, REST API, distributed systems, cloud, Terraform, orchestration, governance, and technical leadership experience.
Python, REST APIs, AWS, S3, RDS, EKS, EventBridge, IAM, Terraform, Prefect, Airflow, Dagster, dbt, OpenTelemetry, ClickHouse, Ray
1d
Save
Mark Applied
Hide
Senior Platform Engineer
San Francisco or United States
$180k-$250k/yr OnsiteFull Time
Ironsite
Ironsite: AI-powered wearable cameras for construction safety and productivity.
5+ YOE5+ years of full-stack product development. Requires React/TypeScript, Node.js, PostgreSQL, Redis, API design, integrations, AWS, scalable web applications, and end-to-end architecture experience.
React, TypeScript, Node.js, PostgreSQL, Redis, AWS, Procore, Autodesk, Trimble, Tagger Portal, SMS
2d
Save
Mark Applied
Hide
Senior Staff Platform Engineer
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOEBachelor’s degree or equivalent experience; 12+ years in large-scale distributed platforms or infrastructure; expertise in cloud, networking, Linux/Unix, programming, automation, and fault-tolerant systems; architecture leadership across teams.
Python, Go, Kubernetes, Terraform, AWS, Azure, Google Cloud Platform, Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, Akamai, AWS CloudFront, Fastly, Cloudflare
2d
Save
Mark Applied
Hide
Senior Staff Platform Engineer
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOEBachelor’s degree or equivalent experience, 12+ years in industry, and expertise architecting large-scale distributed platforms, cloud infrastructure, networking, automation, reliability, and production systems.
Python, Go, Kubernetes, Terraform, AWS, Azure, Google Cloud Platform, Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, Akamai, AWS CloudFront, Fastly, Cloudflare
2d
Save
Mark Applied
Hide
Staff Platform Engineer
Miami or New York City or San Francisco or United States
RemoteFull Time
Félix Pago
Félix Pago: Remittance platform for money transfers via WhatsApp messaging.
10+ YOE10+ years of software engineering experience with deep expertise in cloud architecture, Kubernetes, Istio, Terraform, networking, security, Golang, system architecture, developer portals, and platform leadership.
Kubernetes, AWS, GCP, Istio, Terraform, Backstage, Port.io, Helm, Terratest, Golang, CI/CD, Internal Developer Platform (IDP), MLOps, LLM
3d
Save
Mark Applied
Hide
​​Platform Engineer, Onboard Compute
San Jose, California, United States
$156k-$186k/yr HybridFull Time
Muon Space
Muon Space: Designs, builds, and operates low Earth orbit satellite constellations.
3+ YOEBachelor's degree in a technical field and 3+ years of engineering experience; FPGA/GPU workflows, build systems, test environments, integration pipelines, scripting, Jira, Confluence, planning, and communication skills.
Jira, Confluence, Vivado, Quartus, NVIDIA Jetson, NVIDIA IGX
4d
Save
Mark Applied
Hide
Platform Engineer
San Francisco or Singapore or Europe
$150k-$250k/yr OnsiteFull Time, Contract
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
2+ YOERequires 2–4 years owning production cloud infrastructure, deep AWS and container experience, CI/CD and observability expertise, backend architecture judgment, and production coding skills.
Amazon Web Services (AWS), Terraform, Kubernetes, Amazon Elastic Kubernetes Service (EKS), Helm, Docker, Amazon Elastic Compute Cloud (EC2), AWS CodeBuild, Amazon Elastic Container Registry (ECR), Amazon Simple Storage Service (S3), AWS Identity and Access Management (IAM), CI/CD, APIs
6d
Save
Mark Applied
Hide
Multi-Cloud CDN Scheduling Platform Engineer Intern (CDN Platform) - 2027 Summer
San Jose, California, United States
OnsiteInternship
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Currently pursuing a master's in a related technical discipline; familiarity with Python, Go, Java, or C/C++; basic knowledge of data structures, operating systems, networks, or databases.
Python, Go, Java, C, C++, Linux, SQL, HTTP, DNS
This job has expired