The paper presents a comparative analysis of the effectiveness of various text vectorization methods for the task of Sentiment Analysis of Russian-language reviews. The study covers classical frequency-based approaches (TF IDF, n-grams), statistical models (Word2Vec, FastText), and a contextual method based on the pre-trained BERT language model. The practical part of the research includes the implementation of text processing and classification pipelines using logistic regression and Naive Bayes classifiers. Experiments are conducted on a dataset of Russian-language reviews from a marketplace. Key comparison metrics are classification accuracy (accuracy, F1-score) and model training/inference time. The results show that on small datasets, classical methods with linear models demonstrate competitive quality with significantly lower computational costs. Contextual BERT embeddings show the best quality on the test set; however, their use is justified only with sufficient data volumes and the absence of strict real-time inference constraints. Based on the analysis, recommendations are given for choosing a vectorization method depending on the data volume and system performance requirements in real time.
O. I. Zakharova, S. Bednyak, Yaroslav Dmitrievich Kanunnikov· Infokommunikacionnye tehnolo...· 0 citations
The article examines an adaptive algorithm for automated selection and training of machine learning models, AIONet (Adaptive Intelligent Optimization Network), based on the integration of decision tree methods and large language models (LLMs). A comparative analysis is conducted of the capabilities of classical decision tree algorithms and modern LLMs in the task of identifying suitable models and datasets for automated neural network training. A hypothesis is formulated that combining a decision tree as a structure for primary logical selection with an LLM as a context-dependent intelligent module can improve the accuracy and efficiency of model and dataset selection compared to using each approach independently. The key features of the AIONet algorithm’s operation are presented, as well as its potential for application in automated machine learning systems.
S. A. Maslov, O. I. Zakharova· Infokommunikacionnye tehnolo...· 0 citations
The results show that the proposed methodology is not limited to a single algorithm and permits the combination of direct extraction, specialized models, multi-stage pipelines, and OWL ontologies, which can be applied to the analysis of scientific, technical, and regulatory texts in a secure local environment.
O. I. Zakharova, K. N. Ivanov, S. P. Levashkin et al.· COMPUTATIONAL MATHEMATICS AN...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.