Google
Posted 2d ago

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google
Mountain View, California, United States
$207k-$300k/yrOnsiteFull Time
Responsibilities
  • optimizing workloads
  • designing techniques
  • investigating bottlenecks
Requirements
  • Bachelor's degree or equivalent experience
  • 8 years of software development
  • Python and C++
  • AI inference optimization expertise, and experience with serving codebases and performance tradeoffs
Technical tools mentioned
PythonC++vLLMTensorRT-LLMSGLangDynamoPyTorch profilerGPUTPU

Job description

Minimum qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Preferred qualifications:

  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.

About the job

At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.


We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

About Google

Provides online search, advertising, cloud computing, and consumer electronics.

Year founded
1998
Employees
190000
Organization type
Public
Headquarters
US

Similar jobs

Software Engineer roles near Mountain View, California
6h
Save
Mark Applied
Hide
Software Engineer – Control Systems
Tempe or San Francisco or Phoenix
$130k-$170k/yr HybridFull Time
Azora
Azora: Building modular optical ground stations for laser space communications.
4+ YOEBS/MS in a relevant field, 4+ years of embedded or hardware-adjacent C/C++ experience, embedded Linux or RTOS expertise, hardware interface knowledge, and lab debugging skills.
C, C++, Rust, Python, Embedded Linux, FreeRTOS, Zephyr, VxWorks, RTEMS, SPI, I2C, UART, CAN, PCIe, JTAG, cFS, F´, Zynq, Versal, Microsoft Excel
16h
Save
Mark Applied
Hide
Staff Software Engineer (Crypto, SV) - US - IC7 - 2026
Palo Alto, California, United States
HybridFull Time
Nubank
NubankNYSE: NU: Digital financial platform offering banking, credit, and investment services.
Staff-level impact designing critical backend/platform systems; deep distributed-systems expertise; production ownership of high-stakes financial or blockchain infrastructure; security architecture, incident leadership, cross-team influence, and mentorship.
22h
Save
Mark Applied
Hide
Senior Software Engineer - Cloud Data Platform
Santa Clara, California, United States
$140k-$160k/yr OnsiteFull Time
Picarro
Picarro: Manufacturer of high-precision gas and isotope analysis instruments.
5+ YOERequires 5+ years of software engineering experience, end-to-end production ownership, backend or full-stack expertise, data-intensive systems experience, cloud infrastructure, CI/CD, observability, and AI coding assistant experience.
Python, Go, TypeScript, AWS, GCP, Azure, CI/CD, Cursor, TimescaleDB, InfluxDB, Kubernetes, Terraform, GitOps, GitHub, FSA, HSA, 401K
23h
Save
Mark Applied
Hide
Software Engineer, Accessibility
Cupertino or North America
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
3+ YOEBachelor's in Computer Science or equivalent; object-oriented software design skills; 3+ years relevant experience preferred; proficiency in a listed programming language and strong communication skills.
Objective-C, Swift, C++, C, Apple Accessibility APIs, Accessibility Keyboard, Switch Control, Head Tracking, Eye Tracking
1d
Save
Mark Applied
Hide
Software Engineer Intern (TikTok AI Search & Visual Search Infra Team) - 2027 Summer
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship
TikTok
TikTok: Global short-form video hosting and social media platform.
Pursuing a bachelor's degree in computer science or related discipline; strong programming, data structures, algorithms, Linux, analytical, communication, and collaboration skills required.
Linux, C++, Java, Python, Go, vLLM, SGLang, TensorRT-LLM, LangGraph, CUDA
1d
Save
Mark Applied
Hide
Software Engineer II
San Francisco, California, United States
$214k-$256k/yr HybridFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
5+ YOEBachelor's degree in a specified technical field plus 5 years of progressive post-baccalaureate experience. Requires programming, version control, SQL, algorithms, distributed systems, debugging, monitoring, and software lifecycle expertise.
C++, Python, Java, C#, Go, Git, Phabricator, SQL, MySQL
1d
Save
Mark Applied
Hide
Principal Engineer Software (DLP)
Santa Clara, California, United States
$147k-$238k/yr HybridFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
9+ YOEBS/MS in Computer Science or Engineering and 9+ years of experience. Requires hands-on Core Java, cloud/distributed systems, Spring, REST APIs, MongoDB, Elasticsearch, Kubernetes, Docker, and AWS, Google Cloud, or Azure experience.
Java, Rust, Go, C, C++, Spring, REST API, MongoDB, Elasticsearch, Kubernetes, Docker, AWS, Google Cloud, Azure, Agile
1d
Save
Mark Applied
Hide
Software Engineer, Cloud Infrastructure
San Francisco, California, United States
OnsiteFull Time
Weave Robotics
Weave Robotics: Developing autonomous personal robots for household tasks.
3+ YOE3+ years building backend or media infrastructure in production, with real-time networking, cloud infrastructure, containerization, security, production databases, and reliability experience.
Terraform, OpenTofu, Ansible, Docker, Kubernetes, UDP, TCP, STUN, TURN, ICE, TLS, mTLS, PKI, KMS, HSM, aiortc, Pion, LiveKit