We're building a verified dataset to train a model that automatically extracts word norm metadata from papers. Every verification you submit helps improve the model.
73
Contributors
1,358
Papers in dataset
1,291
Papers with extracted metadata
0% of first milestone
500 human-verified extractions gives us enough labeled examples to train a supervised model that can reliably predict a paper's language, stimuli type, norms collected, and participant details directly from its title and abstract. Reaching 1,000 unlocks our stretch goal: a higher-accuracy model that can flag uncertain predictions for human review.
5,920
Training samples
1,321
Validation samples
0.799
F1 score
90.6%
Accuracy
These datasets are suitable for training and evaluating models. Each download includes a companion codebook describing all columns.
Title, abstract, authors, and DOI for every reviewed paper, labeled “included” or “excluded”. Useful for training a relevance classifier.
All included papers with full metadata and extracted study information—AI extraction alongside any human-verified corrections.