This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

Hotdata
Posted 2mo ago

Research Engineer (PhD) - Database Internals / Special Projects

Hotdata
San Francisco or United States
$10k-$20k/moHybridFull Time
Responsibilities
  • researching architecture
  • building prototypes
  • collaborating engineers
Requirements
  • PhD in CS/Distributed Systems/Databases
  • Strong systems programming in Rust/C++
  • Experience with query engines
  • Distributed systems
  • Storage engines
  • Research and prototyping
Technical tools mentioned
RustC++Apache ArrowDataFusionParquet

Job description

Research Engineer (PhD) - Database Internals / Special Projects

Description

Company: Hotdata, Inc.

Location: Remote or Bay Area

Type: Full-time


About Hotdata

Hotdata is building the next generation of the data layer for AI systems.

For the past decade, data infrastructure has been optimized for dashboards and analytics. But when agents are the creators and consumers of databases, we require new primitives for data systems.


Hotdata is building the infrastructure layer that enables agent-native applications to store, query, and reason over structured data efficiently. Our work sits at the intersection of database internals, distributed systems, query engines, and AI infrastructure.


We are looking for a research-oriented engineer with deep systems curiosity to explore new ideas in data systems and help translate them into working prototypes.


The Role

You will work directly with the founders on special projects exploring the future of database architecture, including:

  • New query engine architectures
  • AI-native data systems
  • Incremental computation and streaming data
  • Query planning and execution optimization
  • Vectorized execution and memory layouts
  • Agent-driven data workflows

Your work will start as research prototypes and often evolve into core components of the platform.


What You Will Work On

Examples of problems we are actively exploring:

  • Next-generation query planning and optimization
  • Systems built on Apache Arrow-style columnar memory models
  • Incremental and reactive query engines
  • New approaches to dataflow execution
  • Efficient state management for agents
  • Hybrid analytical/operational storage engines
  • Distributed query processing
  • Rust-based data infrastructure

This role sits close to database internals, not application development.


Responsibilities

  • Research and prototype new architectures for data systems
  • Build experimental query engines or execution frameworks
  • Work on low-level systems components in Rust or C++
  • Explore novel ideas in query planning, execution, and storage
  • Publish technical insights through internal papers, blogs, or talks
  • Translate research ideas into production-grade infrastructure
  • Collaborate with engineers building the Hotdata platform



What We’re Looking For

Required

  • PhD in Computer Science, Distributed Systems, Databases, or related field
  • Deep understanding of database internals or query engines
  • Strong systems programming experience (Rust, C++, or similar)
  • Experience working with at least one of:
  • query engines
  • compilers
  • distributed systems
  • storage engines
  • Strong research and prototyping ability

Bonus

  • Experience with Apache Arrow, DataFusion, DuckDB, Velox, or Spark
  • Contributions to open source data infrastructure
  • Experience with vectorized query execution
  • Background in query optimization
  • Research publications in data systems conferences (SIGMOD, VLDB, CIDR, etc.)

What Makes This Role Unique

  • Work directly on the core architecture of a new data system
  • Small team where research ideas ship into production
  • Opportunity to influence the future of AI data infrastructure
  • Freedom to pursue deep technical ideas


Ideal Candidates

You might be a:

  • PhD student finishing research in databases or distributed systems
  • Systems engineer who has worked on query engines or storage engines
  • Open source contributor to data infrastructure projects
  • Researcher interested in building real systems


Technologies We Care About

Examples of systems and technologies relevant to this work include:

  • Rust
  • Apache Arrow
  • DataFusion
  • Parquet
  • DuckDB
  • Query optimizers
  • Distributed data systems


Why Join Hotdata

Most database companies optimize existing systems. We believe the rise of AI agents fundamentally changes the data layer. Hotdata is building the infrastructure for that shift. If you enjoy working at the boundary of research and real systems, we'd love to talk.



About the Company


We’re an early-stage, VC-backed startup at the intersection of data systems and agentic AI. Founded by three experienced engineers and product leaders, we’re building an agent-first product at the intersection of data systems and developer platforms.

We combine modern data systems, vector indexing, and RAG techniques to give agents instant, intelligent access to enterprise data - without the complexity of traditional data pipelines.

We’re an early team solving hard distributed systems challenges with cutting-edge technologies (DataFusion, Arrow, Rust, cloud-native infra). If you want to shape the foundation of a new compute layer for AI agents, this is the place.

About Hotdata

Building agent-native data infrastructure for AI applications.

Similar jobs

Research Engineer roles near San Francisco, California
1d
Save
Mark Applied
Hide
Research Engineer, Preference Data
San Francisco, California, United States
$250k-$400k/yr OnsiteFull Time
Vizcom
Vizcom: AI-powered tools for industrial designers to visualize concepts instantly.
Experience building training-data or large-scale data pipelines; experimental mindset. Model training, labeling, human feedback, evaluation operations, and privacy or contractual data constraints are preferred.
LeetCode
1d
Save
Mark Applied
Hide
Research Engineer – Benchmarking
San Francisco or New York City or London
$130k-$500k/yr OnsiteFull Time
Mercor
Mercor: Connecting expert human intelligence with frontier AI model development.
Applied AI research, model evaluation or benchmarking, strong coding and ML experience, data structures and algorithms, backend systems, APIs, SQL or NoSQL, cloud platforms, and model behavior analysis.
SQL, NoSQL, NeurIPS, ICML, ACL
2d
Save
Mark Applied
Hide
Robotics Research Engineer - Robot Simulation and Evaluation
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEPhD in computer science, robotics, or related field; 8+ years in robotics, simulation, or robot learning; software design expertise; Python and deep learning software stack proficiency.
NVIDIA Omniverse, Python, PyTorch, JAX, PhysX, Isaac Gym, Isaac Lab, CUDA, Warp, ROS
4d
Save
Mark Applied
Hide
Research Engineer, Lab Automation
Menlo Park or San Francisco
$200k-$250k/yr OnsiteFull Time
Periodic Labs
Periodic Labs: Builds autonomous laboratories for AI-driven scientific discovery.
PhD or equivalent research experience in materials science, chemistry, chemical engineering, or related field; materials lab hardware expertise; Python proficiency; and ability to translate scientific workflows into automation requirements.
Python, Electronic Lab Notebooks, LIMS
5d
Save
Mark Applied
Hide
Research Engineer - New Grad (2027)
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$140k-$200k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
Recent MSc or PhD graduate in machine learning, computer vision, autonomy, robotics, or related field; experience with Python, PyTorch, computer vision, robotics, and distributed model training.
Python, PyTorch
5d
Save
Mark Applied
Hide
Research Engineer, LangSmith Engine
New York City or San Francisco
OnsiteFull Time
LangChain
LangChain: Tools for building and deploying production-ready AI agents.
4+ YOERequires 4+ years in ML/AI research, a relevant master's or Ph.D., LLM and AI agent experience, benchmark and experiment design, and strong software engineering skills.
LangSmith, LangChain, LangGraph, Deep Agents, LLMs, AI agents, GPU infrastructure, SFT, RLHF, RLAIF
5d
Save
Mark Applied
Hide
Research Engineer, Synthetic Data
San Francisco or Singapore
$150k-$250k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
2+ YOERequires 2–4 years in software, ML engineering, or AI research; Python, Linux, Docker, synthetic data pipelines, evaluation frameworks, structured datasets, and independent project ownership.
Python, Linux, Docker
6d
Save
Mark Applied
Hide
Lead Research Engineer, Search & Retrieval
New York City or Frisco or Toronto or Ann Arbor or Eagan or San Francisco or Los Angeles or Irvine or McLean or Washington
$137k-$255k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
7+ YOEBachelor's or master's in computer science, engineering, or related field; 7+ years building production software; search and retrieval expertise; Python, AWS, OpenSearch or Vespa, distributed systems, and technical leadership.
OpenSearch, Vespa, Elasticsearch, Solr, Lucene, Python, AWS, Kafka, RAG, A/B tests
This job has expired