Using Simple Text Mining Tools to Power an Intelligent Learning System for Lengthy, Domain Specific Texts
Abstract
This paper describes the development of a system for tracking the opportunities for vocabulary and conceptual learning within lengthy subject area texts. Specifically, we leveraged text mining tools to observe the cumulative frequency of words throughout a novel and chapters from two college-level textbooks. We describe how cumulative lexical occurrence and frequency information may be made accessible and useful to researchers and educators without the need for them to have deep knowledge of NLP or other computational linguistic skills. We applied these tools to three different types of texts, a narrative novel, and two textbooks and provide some preliminary results. Finally, we describe potential methods of using this information to structure learning activities utilizing topical analytics and automatic item generation techniques. By leveraging cumulative lexical frequency throughout a text, researchers and educators may be able to isolate words that may be both difficult and important in comprehending any given text, as well as building domain-specific knowledge.
Publication Title
Communications in Computer and Information Science
Recommended Citation
Sabatini, J., & Hollander, J. (2023). Using Simple Text Mining Tools to Power an Intelligent Learning System for Lengthy, Domain Specific Texts. Communications in Computer and Information Science, 1831 CCIS, 727-733. https://doi.org/10.1007/978-3-031-36336-8_112
