Stanford Tabular and Relational (STAR) Project

Foundation models for tabular and relational data.

Advancing foundation models for structured data, from a single table to the many linked tables of a relational database.

Open Resources

Models, datasets, and code.

Models

ModelDescriptionParams
rt-j Context-efficient relational foundation model pretrained on THE JOIN. 85M open ↗
rt-plurel Relational Transformer pretrained on the PluRel synthetic databases. 22M open ↗

Datasets

DatasetDescriptionDatabases
the-join The largest open relational corpus, with around 6,000 forecasting tasks. 650 open ↗
relbench Real-world relational databases with diverse predictive tasks. 7 open ↗
relbench-v2-extra Real-world databases extending the RelBench v2 collection. 3 open ↗
plurel Synthetic relational databases for scaling-law pretraining. 2,000 open ↗
redelex Databases ported from the CTU Prague Relational Learning Repository. 71 open ↗
tgb The Temporal Graph Benchmark of dynamic graphs for relational learning. 12 open ↗
dbinfer Databases from the 4DBInfer benchmark for graph-centric predictive modeling. 7 open ↗

Talks

Talks and presentations.

People

Built at Stanford, with collaborators across academia and industry.

Partner institutions
University of Oxford
Kumo AI
SAP
NVIDIA