Vector Space Model of Knowledge Representation Based on Semantic Relatedness

Authors

  • Dmitry V. Bondarchuk Ural State University of Railway Transport

DOI:

https://doi.org/10.14529/cmse170305

Keywords:

text-mining, vector space model, semantic relatedness

Abstract

Most of text mining algorithms uses vector space model of knowledge representation. Vector space model uses the frequency (weight) of term to determine its importance in the document. Terms can be semantically similar but different lexicographically, which in turn will lead to the fact that the classification is based on the frequency of the terms does not give the desired result. Analysis of a low-quality results shows that errors occur due to the characteristics of natural language, which were not taken into account. Neglect of these features, namely, synonymy and polysemy, increases the dimension of semantic space, which determines the performance of the final software product developed based on the algorithm. Furthermore, the results of many complex algorithms perceived domain expert to prepare training sample, which in turn also affects quality issue algorithm. We propose a model that in addition to the weight of a term in a document also uses semantic weight of the term. Semantic weight terms, the higher they are semantically closer to each other. To calculate the semantic similarity of terms we propose to use a adaptation of the extended Lesk algorithm. The method of calculating semantic similarity lies in the fact that for each value of the word in question is counted as the number of words referred to the dictionary definition of this value (assuming that the dictionary definition describes several meanings of the word), and in the immediate context of the word in question. As the most probable meaning of the word is selected such that this intersection was more. Vector model based on semantic proximity of terms solves the problem of the ambiguity of synonyms.

References

Budanitsky A., Hirst G. Evaluating WordNet-based Measures of Lexical Semantic Relatedness. Computational Linguistics. 2006. vol. 32. pp. 13–47.

Hotho A., Staab S., Stumme G. WordNet Improve Text Document Clustering. SIGIR 2003 Semantic Web Workshop (Toronto, Canada, July 28 – August 1, 2003). pp. 541–544. DOI: 10.1145/959258.959263.

Sedding J., Dimitar K. WordNet-based Text Document Clustering. COLING 2004 3rd Workshop on Robust Methods in Analysis of Natural Language Data (Geneva, Switzerland, August 23 – 27, 2004). pp. 104–113. DOI: 10.3115/1220355.1220356.

Lesk M. Automatic Sense Disambiguation Using Machine Readable Dictionaries: How to Tell a Pine Cone from an Ice Cream Cone. SIGDOC ’86: Proceedings of the 5th Annual International Conference on Systems Documentation (Toronto, Canada, June 8 – 11, 1986). pp. 24–26. DOI: 10.1145/318723.318728.

Loupy C., El-Beze M., Marteau P.F. Word Sense Disambiguation Using HMM Tagger. Proceedings of the 1st International Conference on Language Resources and Evaluation (Toronto, Canada, June 8 – 11, 1998). pp. 1255–1258. DOI: 10.3115/974235.974260.

Jeh G., Widom J. SimRank: a Measure of Structural-context Similarity. Proceedings of the 8th Association for Computing Machinery’s Special Interest Group on Knowledge Discovery and Data Mining international conference on Knowledge discovery and data mining (Edmonton, Canada, July 23 – 25, 2002). pp. 271–279. DOI: 10.1145/775047.775049.

Kechedzhy K.E., Usatenko O., Yampolskii V.A. Rank Distributions of Words in Additive Many-step Markov Chains and the Zipf law. Physical Reviews E: Statistical, Nonlinear, Biological, and Soft Matter Physics. 2005. vol. 72. pp. 381–386.

Mihalcea R. Using Wikipedia for Automatic Word Sense Disambiguation. Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (New York, USA, Apri; 22 – 27, 2007). pp. 196–203.

Willett P. The Porter Stemming Algorithm: Then and Now. Program: Electronic Library and Information Systems. 2006. Vol. 4., No. 3. P. 219–223.

Bondarchuk D.V. Choosing the Best Method of Data Mining for the Selection of Vacancies. Informacionnye Tehnologii Modelirovanija i Upravlenija [Information Technology Modeling and Management]. 2013. no. 6(84). pp. 504–513. (in Russian)

Salton G. Improving Retrieval Performance by Relevance Feedback. Readings in Information Retrieval. 1997. Vol. 24. pp. 1–5.

Tan P. N., Steinbach M., Kumar V. Top 10 Algorithms in Data Mining. Knowledge and Information Systems. 2008. vol. 14. no. 1. pp. 1–37. DOI: 10.1007/s10115-007-0114-2.

Banerjee S., Pedersen T. An Adapted Lesk Algorithm for Word Sense Disambiguation Using WordNet. Lecture Notes In Computer Science. 2002. vol. 2276. pp. 136–145.

Tezaurus WordNET [Thesaurus WordNET]. Available at: https://wordnet.princeton.edu/ (accessed: 05.02.2017).

Bondarchuk D.V. Intelligent Method of Selection of Personal Recommendations, Guarantees a Non-empty Result. Informacionnye Tehnologii Modelirovanija i Upravlenija [Information Technology Modeling and Management]. 2015. no. 2(92). pp. 130–138. (in Russian)

Published

2017-09-19

Issue

Section

Informatics, Computers and Control