Method of Finding Related Indicators Based on Analysis of Regulatory Legal Acts by NLP Methods
DOI:
https://doi.org/10.14529/ctcr220107Keywords:
socio-economic indicators, N-gram, VDL activity indicator, data mining, NLPAbstract
Modern methods of forecasting time series allow us to obtain very accurate and high-quality forecasts in the presence of retrospective data. However, the results of these methods are determined by the volume and quality of the training sample. When a time series is missing, has a small number of points, or is not reliable at all, time series forecasting methods are ineffective. In this case, it is customary to use approaches to find other indicators that somehow correlate with the desired one, hereinafter referred to as indirect indicators. As part of the work on forecasting socio-economic indicators, it became necessary to form a list of indirect indicators, however, the available solutions for this task do not provide the required reliability. In most cases, these works use data from social networks, forums and other data sources that cannot be considered objective. Since they are an expression of a subjective point of view and may be subject to deliberate falsifications and distortions. Such risks are unacceptable when developing a system created for making managerial decisions at the state level. Aim. Development of methods for searching for indirect indicators based on objective sources of information. These methods make it possible to form a list of indirect indicators without involving experts and eliminating the risks of inaccuracy of primary data. Materials and methods. The research was conducted on the basis of regulatory legal acts of the Russian Federation and its subjects. This source was chosen because regulatory documents are objective and fundamental documents of the state. They are not a representation of the subjective point of view of the author or a group of persons. For the experiment, a part of the regulatory framework from 2016 to 2021 was collected, related to the categories: agriculture, medicine, social sphere and others. Results. The method of finding indirect indicators is defined, various algorithms for ranking indirect indicators are developed and tested, indirect indicators for several socio-economic indicators are formed. The process of identifying indirect indicators is based on the application of Data Mining and NLP methods to the database of regulatory legal acts of the Russian Federation. Conclusion. The resulting solution allowed us to form a list of N-grams associated with the desired indicator. At this stage, the interpretation of the N-gram into an indicator is carried out with the help of an expert, however, this does not require having competencies in the subject area of the indicator.References
Об утверждении методик расчета показателей для оценки эффективности деятельности высших должностных лиц (руководителей высших исполнительных органов государственной власти) субъектов Российской Федерации и деятельности органов исполнительной власти субъектов Российской Федерации: постановление Правительства Рос. Федерации от 03 апреля 2021 г. № 542. URL: https://docs.cntd.ru/document/560760968 (дата обращения: 21.12.2021).
Yancong Xie, Hongxun Jiang. Stock market forecasting based on text mining technology: A support vector machine method. 2019. URL: https://arxiv.org/abs/1909.12789 (дата обращения: 21.12.2021).
F. Swen Kuh, Grace S. Chiu, Anton H. Westveld. Modeling National Latent Socioeconomic Health and Examination of Policy Effects via Causal Inference. 2019. URL: https://arxiv.org/abs/1911.00512 (дата обращения: 21.12.2021).
Isao Yagi, Yuji Masuda, Takanobu Mizuta. Analysis of the Impact of High-Frequency Trading on Artificial Market Liquidity. 2020. URL: https://arxiv.org/abs/2010.13038 (дата обращения: 21.12.2021).
Qi-Qiao He, Patrick Cheong-Iao Pang, Yain-Whar Si. Multi-source Transfer Learning with Ensemble for Financial Time Series Forecasting. 2021. URL: https://arxiv.org/abs/2103.15593 (дата обращения: 21.12.2021).
Dilusha Weeraddana, Nguyen Lu Dang Khoa, Lachlan O Neil, Weihong Wang, Chen Cai. Energy consumption forecasting using a stacked nonparametric Bayesian approach. 2020. URL: https://arxiv.org/abs/2011.05519 (дата обращения: 21.12.2021).
Rajapaksha D., Bergmeir C., Hyndman R.J. LoMEF: A Framework to Produce Local Explanations for Global Model Time Series Forecasts. 2021. URL: https://arxiv.org/pdf/2111.07001.pdf (дата обращения: 21.12.2021).
Sonja Tilly, Giacomo Livan. Macroeconomic forecasting with statistically validated knowledge graphs. 2021. URL: https://arxiv.org/abs/2104.10457 (дата обращения: 21.12.2021).
Jie Huang, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-mei Hwu. Open relation modeling: Learning to define relations between entities. 2021. URL: https://arxiv.org/abs/2108.09241 (дата обращения: 21.12.2021).
Madhav Nimishakavi, Uday Singh Saini, Partha Talukdar. Relation schema induction using tensor factorization with side information. 2016. URL: https://arxiv.org/abs/1605.04227 (дата обращения: 21.12.2021).
Yihong Yuan. Modeling Inter-country Connection from Geotagged News Reports: A Time-Series Analysis. 2017. URL: https://doi.org/10.1007/978-3-319-61845-6_19 (дата обращения: 21.12.2021).
Badgujar A., Chen S., Wang A., Yu K., Intrevado P., Brizan D.G. Quantum Criticism:
A Tagged News Corpus Analysed for Sentiment and Named Entities. 2020. URL: https://arxiv.org/ abs/2006.05267 (дата обращения: 21.12.2021).
Tosin P. Adewumi, Foteini Liwicki, Marcus Liwicki. Word2Vec: Optimal Hyper-Parameters and Their Impact on NLP Downstream Tasks. 2020. URL: https://arxiv.org/abs/2003.11645 (дата обращения: 21.12.2021).
Shuxiao Chen, Koby Crammer, Hangfeng He, Dan Roth, Weijie J. Su. Weighted Training for Cross-Task Learning. 2021. URL: https://arxiv.org/abs/2105.14095 (дата обращения: 21.12.2021).
Официальный интернет-портал правовой информации. URL: http://pravo.gov.ru/ (дата обращения: 21.12.2021).
Veselov D., Kukushkin A., Zamaraev A.N., Yarantsev D., Tihonov S. Solves basic Russian NLP tasks, API for lower level Natasha projects. 2021. URL: https://github.com/natasha/natasha/ (дата обращения: 21.12.2021).
Korobov M. Morphological Analyzer and Generator for Russian and Ukrainian Languages. 2015. URL: https://link.springer.com/chapter/10.1007%2F978-3-319-26123-2_31 (дата обращения: 21.12.2021).
Juan Ramos. Using TF-IDF to Determine Word Relevance in Document Queries. 2003. URL: https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.121.1424&rep=rep1&type=pdf (дата обращения: 21.12.2021).






