The dataset comprises structured provenance event data extracted through the application of custom-trained SpaCy NLP models developed specifically for museum provenance texts. These models are available on Zenodo. The source material consists of 11,392 provenance entries published by the Art Institute of Chicago (AIC), reflecting stylistic conventions aligned with The AAM Guide to Provenance Research. These records were downloaded from the AIC Public API on April 7, 2022.
Each provenance record was segmented into discrete events through a two-stage process. First, the sentence boundary detection model was applied to isolate individual event statements within the record. Then, the span categorization model identified and annotated relevant entities within each event, such as involved parties, dates, locations, and methods of transfer. Together, these models produced a total of 35,554 structured provenance events.
The two models achieved F1 scores of 0.99 for sentence boundary detection and 0.94 for span categorization on their respective test sets. The quality of the resulting dataset reflects the performance of these models: no manual correction or human supervision was applied following automatic extraction. The annotation methodology underlying the dataset is detailed in Teaching Provenance to AI: An Annotation Scheme for Museum Data, and the technical development and analytical affordances of the models are discussed in Hidden Value: Provenance as a Source for Economic and Social History. The resulting dataset is published in JSON format and is intended to support provenance research, historical network reconstruction, and the development of linked open data infrastructures.
| Datum zugänglich gemacht | 15.04.2023 |
|---|
| Verlag | ZENODO |
|---|