Big (and Open) Data for Scholarship of All Sizes: A New Release ... | HathiTrust Digital Library

lterrat's bookmarks 2016-12-06

Summary:

"HathiTrust today announces the release of a significantly expanded open dataset, the HathiTrust Research Center (HTRC) Extracted Features (EF) Dataset, Version 1.0. This dataset provides researchers with open access to data extracted from the full text of the HathiTrust Digital Library (HTDL) at an unprecedented scale. 

The Extracted Features Dataset opens the complete HathiTrust collection for investigations into historical and cultural trends, the rise and fall of topics within the corpus, and the evolution of words and writing structures in publications dating from the 16th to the late 20th century. The EF Dataset provides quantitative information about word and line counts, parts of speech, and other details within each page of every volume in the HTDL. In addition to these larger-scale investigations, the EF Dataset also allows researchers to closely analyze the contents of a given volume or subset of volumes."

Link:

https://www.hathitrust.org/extracted-features-announcement

From feeds:

Open Access Tracking Project (OATP) » lterrat's bookmarks

Tags:

oa.journals

Date tagged:

12/06/2016, 14:06

Date published:

12/06/2016, 09:06