Comparing Prompting Strategies for Low-Resource Indonesian Languages: GPT-4o-mini and IndoBERT on Javanese and Sundanese

Authors

  • Sandi Fadilah Universiti Muhammadiyah Malaysia
  • Ika Safitri Windiarti Universitas Muhammadiyah Palangkaraya
  • Gellysa Urva Institut Teknologi dan Bisnis Riau Pesisir

DOI:

https://doi.org/10.52072/jutekinf.v14i1.1946

Keywords:

Sentiment Analysis, Low-Resource Languages, Few-Shot Prompting, Indonesian Regional Languages, Language Models

Abstract

Javanese and Sundanese are two of Indonesia's most widely spoken regional languages, yet both suffer from severe underrepresentation in NLP research due to limited annotated data and few pre-trained resources. This study compares GPT-4o-mini and IndoBERT-base on three-class sentiment analysis for both languages using NusaX-Senti, evaluating zero-shot and few-shot prompting (nine examples) against an off-the-shelf fine-tuned classifier. Experiments used 100 stratified samples per language. GPT-4o-mini substantially outperforms IndoBERT-base: few-shot achieves Macro-F1 of 0.868 (Javanese) and 0.882 (Sundanese) versus 0.295 and 0.317 for IndoBERT (p<0.001, McNemar test). Few-shot prompting significantly improves Javanese performance (delta F1=+0.122, p=0.015), driven by the Neutral class (F1: 0.560 - 0.800), but shows no significant gain for Sundanese where zero-shot baseline already reaches 0.858. For Indonesian regional languages where domain-specific fine-tuning data is unavailable, prompted large language models offer a practical alternative to pre-trained BERT-based classifiers, requiring only minimal labeled examples rather than full retraining infrastructure.

Downloads

Download data is not yet available.

References

Aji, A. F., Winata, G. I., Koto, F., Cahyawijaya, S., Romadhony, A., Mahendra, R., Kurniawan, K., Moeljadi, D., Prasojo, R. E., Baldwin, T., Lau, J. H., & Ruder, S. (2022). One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7226–7249. https://doi.org/10.18653/v1/2022.acl-long.500

Alahmadi, D. (2025). Human Versus AI: A Comparative Study of Zero-Shot LLMs and Transformer Models Against Human Annotations for Arabic Sentiment Analysis. International Journal of Advanced Computer Science and Applications, 16(8). https://doi.org/10.14569/IJACSA.2025.0160882

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners (Version 4). arXiv. https://doi.org/10.48550/ARXIV.2005.14165

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. https://doi.org/10.18653/v1/N19-1423

Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Chang, B., Sun, X., Li, L., & Sui, Z. (2024). A Survey on In-context Learning. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1107–1128. https://doi.org/10.18653/v1/2024.emnlp-main.64

Jazuli, A., Widowati, & Kusumaningrum, R. (2024). Optimizing Aspect-Based Sentiment Analysis Using BERT for Comprehensive Analysis of Indonesian Student Feedback. Applied Sciences, 15(1), 172. https://doi.org/10.3390/app15010172

Kastrati, M., Imran, A. S., Hashmi, E., Kastrati, Z., Daudpota, S. M., & Biba, M. (2025). Unlocking language barriers: Assessing pre-trained large language models across multilingual tasks and unveiling the black box with Explainable Artificial Intelligence. Engineering Applications of Artificial Intelligence, 149, 110136. https://doi.org/10.1016/j.engappai.2025.110136

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. Proceedings of the 28th International Conference on Computational Linguistics, 757–770. https://doi.org/10.18653/v1/2020.coling-main.66

Lossio-Ventura, J. A., Weger, R., Lee, A. Y., Guinee, E. P., Chung, J., Atlas, L., Linos, E., & Pereira, F. (2024). A Comparison of ChatGPT and Fine-Tuned Open Pre-Trained Transformers (OPT) Against Widely Used Sentiment Analysis Tools: Sentiment Analysis of COVID-19 Survey Data. JMIR Mental Health, 11, e50150. https://doi.org/10.2196/50150

Downloads

Published

2026-06-17

How to Cite

Fadilah, S., Windiarti, I. S., & Urva, G. (2026). Comparing Prompting Strategies for Low-Resource Indonesian Languages: GPT-4o-mini and IndoBERT on Javanese and Sundanese. Jurnal Teknologi Komputer Dan Informasi, 14(1), 10–24. https://doi.org/10.52072/jutekinf.v14i1.1946

Most read articles by the same author(s)

Similar Articles

You may also start an advanced similarity search for this article.