Google DeepMind
Posted 1w ago

Senior Research Engineer, Agentic Data and Tooling, DeepMind

Google DeepMind
New York City, New York, United States
$174k-$252k/yrOnsiteFull Time
Responsibilities
  • building infrastructure
  • integrating data
  • creating tooling
Requirements
  • Bachelor's degree or equivalent experience
  • 5 years with large language models
  • 2 years developing and training machine learning models, and experience with agentic integrations and Model Context Protocol
Technical tools mentioned
Large Language Models (LLMs)Model Context Protocol

Job description

Minimum qualifications:

  • Bachelor's degree in Computer Science, Information Technology, a related technical field, or equivalent practical experience.
  • 5 years of experience working with large language models (LLMs).
  • 2 years of experience developing and training machine learning models.
  • Experience with Agentic integrations and Model Context Protocol.

Preferred qualifications:

  • Master's degree or PhD in Electrical Engineering, Computer Science, or equivalent practical experience.
  • 2 years of experience with full-stack development.
  • Excellent analytical, problem-solving and communication skills with demonstrated attention to detail.
  • A deep passion for AI technology and all of its possibilities .

About the job

At Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup large-scale tests and deploy promising ideas quickly and broadly. Ideas may come from internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.

From creating experiments and prototyping implementations to designing new architectures, engineers work on real-world problems including artificial intelligence, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. But you stay connected to your research roots as an active contributor to the wider research community by partnering with universities and publishing papers.

The Google Deepmind (GDM) Agent Data and Tooling team within the Human Data Platform organization builds the environments, pipelines, and tooling that power Gemini's frontier agentic capabilities. We operate under an active data ownership model — moving beyond commodity data collection to build high-fidelity interactive worlds, capture complex multi-turn trajectories, and land data into model training and evaluation pipelines to drive measurable hillclimbing on coding, computer control, tool-use benchmarks, and more.

Build the critical infrastructure and interactive environments that directly drive Gemini's agentic and reasoning capabilities.

In this role, you will sit at the intersection of software engineering and model training: writing high-velocity production code to create rich interactive worlds.

We are looking for an engineer who loves to deliver code and build 0-1 systems at lightning pace and under high pressure, making heavy use of AI tools to boost velocity/output.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Build agentic data infrastructure at high velocity: Own and deliver key components across the agentic data stack. Rapidly prototype, iterate, and ship robust production code to generate, capture, and curate complex multi-turn agent trajectories at scale.
  • Bridge research and engineering to drive model hillclimbing: Collaborate closely with Gemini research teams to close the loop between data creation and model quality. Understand how multi-turn trajectory design, environment complexity, and reward signals impact SFT, RL training, and capability hillclimbing. Directly integrate curated data into training pipelines and evaluate downstream model performance on frontier benchmarks.
  • Build quality tooling and support gold-standard benchmarks: Create human-in-the-loop annotation tooling and interactive trajectory review surfaces, working in tandem with automated validation checkers leveraging adversarial LLM judges and programmatic verifiers. Support the creation and curation of gold-standard evaluation sets for flagship benchmarks.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

About Google DeepMind

Google DeepMind is a private AI research laboratory developing safe artificial intelligence systems for Alphabet.

Similar jobs

Research Engineer roles near New York City, New York
1d
Save
Mark Applied
Hide
Research Engineer - Data Infrastructure
United Kingdom or United States or Poland or Bulgaria or London or New York City or San Francisco or Warsaw
RemoteFull Time
ElevenLabs
ElevenLabs: AI audio research and voice synthesis software.
Experience building data-intensive systems and distributed data processing at scale, ideally for machine learning pipelines; ability to evaluate data quality and build measurement tooling. Web crawler experience is a bonus.
Kubernetes, GitHub
1d
Save
Mark Applied
Hide
Research Engineer – Semiconductor Materials
Princeton, New Jersey, United States
$104k-$155k/yr OnsiteFull Time
SRI International
SRI International: Independent nonprofit research institute developing technologies and solutions for government and commercial customers.
M.S. or Ph.D. in Materials Science or related field; hands-on MOCVD or comparable crystal growth and characterization experience; U.S. citizenship and ability to obtain and maintain security clearance required.
MOCVD, InP, GaAs, GaSb, MBE, CBE, XRD, PL, SEM, Hall-effect, ECV, MATLAB, C++, Python, LabVIEW
2d
Save
Mark Applied
Hide
Staff+ Research Engineer, RL Data Platform
San Francisco or New York City
$500k-$850k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Full-stack production experience with TypeScript/React and Python, backend services and data pipelines, end-to-end project ownership, stakeholder collaboration, AI tool use, and a relevant bachelor's degree or equivalent experience.
TypeScript, React, Python
1w
Save
Mark Applied
Hide
Research Engineer
San Jose or New York City
$200k-$300k/yr OnsiteFull Time
Tessera Labs
Tessera Labs: Private enterprise AI software helping large organizations modernize ERP and connected business systems.
Significant language-model training or post-training experience, RL tuning, Python and PyTorch or JAX proficiency, distributed GPU training, empirical experimentation, and strong software engineering and writing skills.
SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft, Python, PyTorch, JAX, vLLM, SGLang, Triton, DeepSpeed, Ray, Megatron, TRL
1w
Save
Mark Applied
Hide
Research Engineer - Optimization and Dynamical Systems
Florham Park, New Jersey, United States
$94k-$198k/yr OnsiteFull Time
CACI
CACINYSE: CACI: Provider of specialized IT and mission-critical government services.
5+ YOEMaster’s degree in computer science, applied mathematics, physics, or related field plus 5+ years’ experience; Python and C++, Linux, optimization, signal processing, dynamical systems, modeling, simulation, and Top Secret clearance required.
Python, C++, Linux
1w
Save
Mark Applied
Hide
Research Engineer - Optimization and Dynamical Systems
Florham Park, New Jersey, United States
$94k-$198k/yr OnsiteFull Time
CACI
CACINYSE: CACI: Provider of specialized IT and mission-critical government services.
5+ YOEMaster’s degree in computer science, applied mathematics, physics, or related field; 5+ years’ experience; Python and C++; optimization, signal processing, dynamical systems, Linux, modeling, simulation, leadership, and US citizenship required.
Python, C++, Linux, machine learning
1w
Save
Mark Applied
Hide
Senior Research Engineer, Threat Intelligence
United States or New York City
$140k-$180k/yr RemoteFull Time
SecurityScorecard
SecurityScorecard: Private cybersecurity ratings and third-party risk platform serving organizations managing supply-chain risk.
5+ YOE5–8 years of hands-on engineering experience with threat intelligence, security research, or detection engineering; production threat-intelligence systems experience; Python and TypeScript/Node; cloud, containers, CI/CD, and security standards.
Python, TypeScript, Node, AWS, STIX 2.1, TAXII 2.1, MISP, MITRE ATT&CK, YARA, Sigma, STIX Patterning, CEL, OPA, Splunk, Kinesis, NetFlow, OpenCTI, FAIR, Golang
1w
Save
Mark Applied
Hide
Research Engineer – Benchmarking
San Francisco or New York City or London
$130k-$500k/yr OnsiteFull Time
Mercor
Mercor: AI-powered talent marketplace connecting contractors with global organizations.
Applied AI research, model evaluation or benchmarking, strong coding and ML experience, data structures and algorithms, backend systems, APIs, SQL or NoSQL, cloud platforms, and model behavior analysis.
SQL, NoSQL, NeurIPS, ICML, ACL