Google
Posted 1mo ago

Staff Software Engineer, TPU Machine Learning Supercomputer

Google
Sunnyvale, California, United States
$207k-$301k/yrOnsiteFull Time
Responsibilities
  • designing software
  • developing features
  • leading projects
Requirements
  • 8+ years software development experience (C++ or Go)
  • Experience with large-scale infrastructure
  • Distributed systems, OS
  • Data structures
  • Algorithms
  • Active experience designing software architecture
Technical tools mentioned
C++Go

Job description

Minimum qualifications:

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience with software development in C++ or Go.
  • 5 years of experience with large-scale infrastructure, distributed systems, or networks, as well as testing and launching software products.
  • 3 years of experience with software design and architecture.
  • Experience with operating systems, data structures, and algorithms.
  • Experience developing, integrating, and testing system and user-space software (including tools, dashboards, and monitoring) for hardware accelerators or TPU systems.

Preferred qualifications:

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • 3 years of experience in a technical leadership role leading project teams and setting technical direction.
  • 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
  • Experience building backend software for high-performance computing (HPC) and machine learning (ML) applications, including knowledge of data analytics, ML architecture, and how common algorithms map to software/hardware operations.
  • Understanding of highly distributed systems, control plane and management Software, and networking concepts.

About the job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

As a member of the TPU Machine Learning Supercomputer (MLSC) team, you will design and develop features to significantly improve the scalability and reliability of large-scale software across TPUs and other distributed networked hardware machines. Your work will span various layers of the software stack, from host daemons to network routing and distributed control software running across Google's internal and cloud infrastructure. You will also provide leadership to help formulate and drive software development plans for future supercomputer generations.The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're behind Google's groundbreaking innovations, empowering the development of AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $301000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Design, develop, test, deploy, and debug critical system software that enables TPU Machine Learning accelerators to function seamlessly.
  • Develop advanced analytics and health management capabilities to effectively manage and optimize large-scale ML systems.
  • Lead high-impact projects and steer successful delivery while ensuring alignment with broader team strategies.
  • Provide technical guidance and mentorship to software engineers to foster their professional growth and development.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

About Google

Provides online search, advertising, cloud computing, and consumer electronics.

Year founded
1998
Employees
190000
Organization type
Public
Headquarters
US

Similar jobs

Software Engineer roles near Sunnyvale, California
5h
Save
Mark Applied
Hide
Staff Engineer
San Jose or Durham or Mexico City or Bangalore or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
12+ YOE12+ years of software engineering experience, including 3+ years in distributed systems, cloud platforms, or security infrastructure; deep Kubernetes and Golang expertise; storage, Linux, networking, and security knowledge; bachelor's or master's degree required.
Kubernetes, Golang, Kubernetes Custom Resource Definitions (CRDs), Container Storage Interface (CSI), OpenShift Security Context Constraints (SCCs), Multi-Category Security (MCS), Container Security Operators, Advanced Cluster Security, CIS, FIPS, Cluster API (CAPI), Docker, containerd, Linux
8h
Save
Mark Applied
Hide
Senior Software Engineer, Data Platform
San Mateo, California, United States
$130k-$280k/yr OnsiteFull Time
Verkada
Verkada: Sells integrated AI-powered cloud physical security hardware and software.
5+ YOERequires 5+ years of data engineering experience, distributed systems and data processing expertise, and experience with data infrastructure and analytics. Kubernetes and data visualization experience preferred.
Kubernetes
10h
Save
Mark Applied
Hide
Senior Software Engineer(AI/ML), Trust
Bangalore or San Francisco
₹4500k-₹6500k/yr RemoteFull Time
Airbnb
AirbnbNASDAQ: ABNB: Online marketplace for vacation rentals and travel experiences.
7+ YOE7+ years in backend or platform engineering; strong Python or Java, data structures, algorithms, data engineering, machine learning systems, scalable architecture, testing, and deployment experience.
Python, Java, TensorFlow, PyTorch, Kubernetes, Apache Spark, Apache Airflow, Kubeflow, Apache Kafka, Ray, Apache Hive
11h
Save
Mark Applied
Hide
Senior Software Engineer, Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build distributed, large-scale data processing systems, frameworks, and platforms using big data technologies while collaborating with Apple TV and Video teams.
11h
Save
Mark Applied
Hide
Senior Software Engineer, Hyperscale Build Environments and Tools
Santa Clara, California, United States
$180k-$270k/yr OnsiteFull Time
Pure Storage
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOE5+ years of software engineering experience in infrastructure, developer productivity, build systems, or platform engineering, with C/C++, Linux, Docker, CI, and modern build-system expertise.
C, C++, Make, CMake, Linux, Docker
11h
Save
Mark Applied
Hide
Software Engineer, Linux Kernel / Android
San Jose or Los Angeles or Bellevue
$180k-$240k/yr OnsiteFull Time
Rivet Industries
Rivet Industries: Building integrated task systems for frontline industrial and defense operators.
3+ YOERequires 3+ years developing Android/Linux system software, strong Linux kernel and C/C++ skills, AOSP, HALs, drivers, hardware bring-up, embedded builds, debugging, security mechanisms, and U.S. Person status.
Android, Linux, Linux kernel, AOSP, HALs, C, C++, USB, MIPI, I2C, UART, GPIO, PCIe, U-Boot, Android Bootloader, Yocto, Buildroot, Bazel, Soong, OTA, TPM, AR/XR, NPU, DSP
11h
Save
Mark Applied
Hide
Principal Staff Software Engineer, Systems Infrastructure
Mountain View, California, United States
$226k-$369k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
10+ YOEBA/BS or equivalent practical experience, 10+ years in software or reliability engineering, 5+ years in technical leadership, distributed systems expertise, and experience defining reliability standards across teams.
Java, Go, C++, Python, LLM, SLO, SLI
12h
Save
Mark Applied
Hide
Software Engineer
Bengaluru or San Francisco or Seattle or Ireland
HybridFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provides electronic signature and agreement management software solutions.
5+ YOERequires 5+ years building production backend or platform services, service design ownership, backend programming, databases, APIs, testing, version control, CI/CD, and distributed systems knowledge.
Java, C#, Go, TypeScript, Node.js, SQL, NoSQL, Kubernetes, Docker, Spark, Flink, Kafka, RESTful, CI/CD