Sentiment Analysis of Multilingual Social Media Texts Using Natural Language Processing and Large Language Models

Main Article Content

Pasiyeva Aygul

Abstract

Sentiment analysis becomes difficult when social-media users mix languages, shorten words, attach emojis, or express evaluation through irony and negation. Large language models offer flexible multilingual interpretation, but their output can vary across languages and prompts. This article presents an auditable workflow that combines transparent natural language processing with selective LLM review. A controlled corpus of 180 posts was created in Azerbaijani, Turkish, English, and Italian: 120 balanced training examples and 60 held-out posts covering clean wording, emojis, code-mixing, and negation. The proposed lexicon-character hybrid reached 0.883 accuracy and 0.885 macro-F1, compared with 0.650/0.647 for word TF-IDF and 0.817/0.813 for character TF-IDF. A prespecified gate accepted 58.3% of posts with 94.3% accuracy and sent 25 uncertain cases to LLM or human review. No LLM output entered the numerical score.

Downloads

Download data is not yet available.

Article Details

Data Availability Statement

All data supporting the findings of this study are presented in the text of the scientific work.

Section

Information and Web technologies

Author Biography

Pasiyeva Aygul, University of Turin

Master’s Student in Language Technologies and Digital Humanities

How to Cite

Pasiyeva, A. (2026). Sentiment Analysis of Multilingual Social Media Texts Using Natural Language Processing and Large Language Models. Scientific Collection «InterConf», 309, 115–123. https://interconf.openpubarchive.com/index.php/proceeding/article/view/100

References

Winata, G. I., Aji, A. F., Cahyawijaya, S., et al. (2023). NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages. Proceedings of EACL 2023, 815 –834. https://doi.org/10.18653/v1/2023.eacl- main.57

Zhang, R., Cahyawijaya, S., Cruz, J. C. B., et al. (2023). Multilingual Large Language Models Are Not (Yet) Code -Switchers. Proceedings of EMNLP 2023, 12567–12582. https://doi.org/10.18653/v1/2023.emnlp-main.774

Sharma, G., Chinmay, R., & Sharma, R. (2023). Late Fusion of Transformers for Sentiment Analysis of Code-Switched Data. Findings of EMNLP 2023, 6485– 6490. https://doi.org/10.18653/v1/2023.findings-emnlp.430

Zhu, X., Gardiner, S., Roldán, T., et al. (2024). The Model Arena for Cross-lingual Sentiment Analysis: A Comparative Study in the Era of Large Language Models. Proceedings of WASSA 2024, 141 –152. https://doi.org/10.18653/v1/2024.wassa-1.12

Koto, F., Beck, T., Talat, Z., et al. (2024). Zero -shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon. Proceedings of EACL 2024, 298 –320. https://doi.org/10.18653/v1/2024.eacl- long.18

Zeng, L. (2024). Leveraging Large Language Models for Code -Mixed Data Augmentation in Sentiment Analysis. Proceedings of SICon 2024, 85 –101. https://doi.org/10.18653/v1/2024.sicon-1.6

Shynkarov, Y., Solopova, V., & Schmitt, V. (2025). Improving Sentiment Analysis for Ukrainian Social Media Code-Switching Data. Proceedings of UNLP 2025, 179–193. https://doi.org/10.18653/v1/2025.unlp-1.18

Šmíd, J., Priban, P., & Kral, P. (2025). LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation. Proceedings of ACL 2025, 839–853. https://doi.org/10.18653/v1/2025.acl-long.41

Balaga, P. S., Karthik, N., Vishwanath, C., et al. (2025). Power Doesn’t Reside in Size: A Low Parameter Hybrid Language Model for Sentiment Analysis in Code -Mixed Data. Proceedings of EMNLP 2025, 14808 –14816. https://doi.org/10.18653/v1/2025.emnlp-main.749

Ibrohim, M. O., Caselli, T., Bosco, C., et al. (2026). Multilingual Structured Sentiment Analysis for Environmental Sustainability. Proceedings of LREC 2026, 7937–7954. https://doi.org/10.63317/4jdnsugtypgu

Stróżyna, M., Lewoniewski, W., & Czumałowska, I. (2026). Evaluating Multilingual Sentiment Classifiers Using an LLM -Annotated Wikipedia Benchmark. Proceedings of GEM 2026, 692 –703. https://doi.org/10.18653/v1/2026.gem-main.63

Dai, S., & Lin, W. (2026). ALPS -Lab at SemEval -2026 Task 3: A Multilingual Generative LLM Approach for Dimensional Aspect Sentiment Analysis. Proceedings of SemEval 2026, 1652 –1658. https://doi.org/10.18653/v1/2026.semeval-1.212