Text Recognition techniques for premodern Italian and Devanāgarī manuscripts (uniqueatpenn.wordpress.com)

🤖 AI Summary
Graduate Fellows Priyamvada Nambrath and Eleanor Webb are utilizing Handwritten Text Recognition (HTR) technology to transcribe premodern manuscripts from the University of Pennsylvania's collections. Their project highlights the potential of HTR, which leverages machine learning to convert handwritten texts into digital formats. While Nambrath focuses on Devanāgarī script manuscripts from the eighteenth century, Webb works on a seventeenth-century Italian mathematics manual using the open-source platform eScriptorium, which integrates with Kraken, an automatic text recognition system. The duo faces distinct challenges; while Webb can refine existing models, Nambrath must create her own ground truth from scratch due to a lack of available models for Devanāgarī. This project is significant for the AI/ML community as it demonstrates the application of advanced text recognition technologies to underrepresented historical scripts, potentially transforming manuscript research. The ability of eScriptorium to segment manuscripts into different text regions and create interoperable data is another key advancement. However, the varying performance of HTR models raises questions about the representativeness of training data and model accuracy, reminding researchers of the complexities involved in transcribing historical documents. As the digital humanities field continues to evolve, this initiative underscores both the promise and limitations of leveraging AI to engage with traditional manuscripts more effectively.
Loading comments...
loading comments...