Stanford Tabular and Relational (STAR) Project

Foundation models for tabular and relational data.

Advancing foundation models for structured data, from a single table to the many linked tables of a relational database.

Open Resources

Models, datasets, and code.

Models

ModelDescriptionParams
rt-j Context-efficient relational foundation model pretrained on THE JOIN. 85M open ↗
rt-plurel Relational Transformer pretrained on PluRel synthetic databases, for in-context prediction; the #1 system on RelArena-α. Original 22M PluRel-paper checkpoints preserved under paper/. 85M open ↗
rt-v1 Original Relational Transformer checkpoints from the ICLR 2026 paper, including per-dataset held-out pretraining and continued-pretraining variants. 22M open ↗

Datasets

DatasetDescriptionDatabases
the-join The largest open relational corpus for pretraining, spanning academic, e-commerce, finance, sports, biomedical, government, and text2sql domains, with forecasting and autocomplete tasks. 639 open ↗
relbench The original RelBench v1 databases and tasks, in the self-describing manifest format. 7 open ↗
relbench-v2-extra Additional real-world databases introduced with RelBench v2. 3 open ↗
plurel Synthetic relational databases generated by PluRel for scaling-law pretraining. 2,000 open ↗
redelex Databases from the CTU Prague Relational Learning Repository, ported via redelex. 71 open ↗
tgb Temporal Graph Benchmark datasets exported to the RelBench format, with official splits and negatives. 12 open ↗
dbinfer The 4DBInfer benchmark databases in RelBench format, with their original labels. 7 open ↗

Talks

Talks and presentations.

People

Built at Stanford, with collaborators across academia and industry.

Partner institutions
University of Oxford
Kumo AI
SAP
NVIDIA
Prior Labs