
Tuesday, November 1, 2022
11:00am to noon
247 Hesburgh Library, Navari Family Center for Digital Scholarship
Text mining, a process for extracting information from unstructured text, requires everyday files (PDF, Word, HTML, etc.) to be transformed into plain text files. Once your files are in a plain text format (no bold, no italics, no underlining, etc.) they are ready for automated processing and computer analysis.
This hands-on workshop will demonstrate and facilitate the use of a free Java-based program called Tika to do this work. More specifically, this workshop will help attendees install Tika and use it to convert just about any file into plain text, and then participants will be empowered to use a myriad of text mining services available on the 'Net.
Please bring your own laptop.
Related LibGuide: Text Mining and Analysis by Eric Lease Morgan
Open to Undergraduates, Graduate Students, Faculty, Staff, Postdocs
Digital Initiatives Librarian
emorgan@nd.edu

Eric Morgan is the Digital Initiatives Librarian in the Navari Family Center for Digital Scholarship. His current work focuses on assisting faculty and students with text mining and analysis. Though his work involves extensive computing expertise, Eric considers himself to be a librarian first, and a computer user second. His professional goal is to discover new ways to use computers to provide better library service. His research interests have included information retrieval, expert systems, and automated personalization.
MIS, 1987, Drexel University; B.A. Philosophy, 1982, Bethany College
For more information: http://www.nd.edu/~emorgan/