Image: Jessica Kourkounis / Stringer via Getty Images
2 min. read
When a scholar visits an archive or a rare book library, they might come into contact with hundreds of pages of manuscripts, letters, and ephemera that were written by hand. Depending on the time and place the historical documents come from, the handwriting may be highly stylized, difficult to read, and challenging to translate. Additionally, making connections among these disparate documents can be tricky and time-consuming.
What if, instead of parsing the highly-stylized handwriting of historical documents in an archive or a rare book library, scholars could read a typed transcript of the manuscript? Or could quickly search the manuscript for particular words or phrases?
Staff and research fellows at the Penn Libraries are working towards that future. With the help of eScriptorium, a platform for the transcription of historical materials, they are using manuscripts from the Libraries’ collections to build machine learning models—a type of AI—that can transcribe handwritten text. This process is called Handwritten Text Recognition, or HTR. By doing so, they are helping the Libraries expand access to Penn’s cultural heritage and transform physical archives into rich data resources that can be explored and analyzed by the global research community.
Leading the project is digitization project coordinator Jessie Dummer, who introduced the idea to the Penn Libraries’ Manuscript Collections as Data Research Group, run by Jajwalya Karajgikar, applied data science librarian, and Dot Porter, curator of digital humanities.
“The Manuscript Collections as Data Research Group has been talking about lots of different technologies that could be applied to collections,” says Dummer. “We started thinking about how HTR might be applied to our work in the library, our collections, and the research that our patrons wanted to do.”
Last fall, after setting up eScriptorium with the help of computing power from the School of Arts and Sciences General Purpose Cluster, the team brought on Eleanor Webb and Priyamvada Nambrath, both historians deeply familiar with handwritten manuscripts, as research fellows through the Schoenberg Institute for Manuscript Studies. Over the past eight months, they have been working on the complex task of “teaching” a computer to read two very different historical manuscripts: a 17th century Italian mathematics manual and an 18th century philosophical work from India, written in Sanskrit.
The project team points out that eScriptorium’s ability to transcribe a historical document, especially a complex one, is only as good as the model. Like any platform built on machine learning or artificial intelligence, it is reliant on the data it has available. In practice, this means that if a lot of scholars have used it to transcribe 17th-century Italian handwriting—and have taken the time to make corrections and ensure that the transcription is accurate—the models produced by the platform will have an easier time transcribing similar manuscripts.
Read more at Penn Libraries.
From Penn Libraries
Image: Jessica Kourkounis / Stringer via Getty Images
(Image: Lance Nelson)
Image: shih-wei via Getty Images
A bioengineered bean gum from the lab of Penn Dental’s Henry Daniell is found to reduce the levels of three microbes associated with head and neck squamous cell cancer to almost zero, without affecting the beneficial bacteria normally found in the mouth.
(Image: Kevin Monko/Penn Dental Medicine)