Topic Modeling of Indonesian News Articles on The Free Nutritious Meal Program: A Comparative Study of LDA, LSA, and NMF

Authors

  • Irmma Dwijayanti Universitas Jenderal Achmad Yani Yogyakarta, Indonesia
  • Muhammad Habibi Universitas Jenderal Achmad Yani Yogyakarta, Indonesia

Keywords:

Topic Modeling, Indonesian News, LDA, LSA, NMF

Abstract

The rapid expansion of digital journalism in Indonesia has produced a large volume of long-form news articles covering complex socio-political and economic issues. One of the most widely discussed initiatives is the Free Nutritious Meal Program (Makan Bergizi Gratis, MBG), which has been reported from perspectives of nutrition, education, public policy, and economic sustainability. Due to the scale of coverage, manual analysis is impractical, requiring automated methods to extract latent themes. This study applies and compares three classical topic modeling techniques: Latent Dirichlet Allocation (LDA), Latent Semantic Analysis (LSA), and Non-Negative Matrix Factorization (NMF). The corpus was preprocessed using tokenization, stop-word removal, stemming, and normalization, followed by experiments with different numbers of topics. Model performance was evaluated with topic coherence and qualitative interpretability. The findings show that LDA provides the highest coherence and generates semantically rich topics. NMF yields sharper topic boundaries but sometimes lacks interpretability, while LSA demonstrates computational efficiency with lower semantic depth. Thematic analysis identified five dominant themes: policy implementation, nutritional adequacy, funding challenges, public responses, and governance risks. Temporal analysis revealed that MBG discourse has been present since 2020, but increased sharply in 2024 during Prabowo Subianto’s presidential campaign and grew further in 2025 with nationwide implementation. These results confirm that MBG is not only a nutrition and welfare policy but also a politically salient initiative intersecting with electoral strategies, public health, and socio-economic development.

Downloads

Download data is not yet available.

References

[1] A. Tri Haryanto, “APJII: Jumlah Pengguna Internet Indonesia Tembus 221 Juta Orang,” detikinet. Accessed: Sep. 16, 2025. [Online]. Available: https://inet.detik.com/cyberlife/d-7169749/apjii-jumlah-pengguna-internet-indonesia-tembus-221-juta-orang

[2] D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet Allocation,” Journal of Machine Learning Research, vol. 3, pp. 993–1022, 2003, doi: 10.1162/jmlr.2003.3.4-5.993.

[3] R. Egger and J. Yu, “A Topic Modeling Comparison Between LDA, NMF, Top2Vec, and BERTopic to Demystify Twitter Posts,” Frontiers in Sociology, vol. 7, p. 886498, May 2022, doi: 10.3389/FSOC.2022.886498/BIBTEX.

[4] M. Hankar, M. Kasri, and A. Beni-Hssane, “A comprehensive overview of topic modeling: Techniques, applications and challenges,” Neurocomputing, vol. 628, p. 129638, May 2025, doi: 10.1016/J.NEUCOM.2025.129638.

[5] Y. T. Guo, Q. Q. Li, and C. S. Liang, “The rise of nonnegative matrix factorization: Algorithms and applications,” Inf Syst, vol. 123, p. 102379, Jul. 2024, doi: 10.1016/J.IS.2024.102379.

[6] J. Wang and X. L. Zhang, “Deep NMF topic modeling,” Neurocomputing, vol. 515, pp. 157–173, Jan. 2023, doi: 10.1016/J.NEUCOM.2022.10.002.

[7] M. Rüdiger, D. Antons, A. M. Joshi, and T. O. Salge, “Topic modeling revisited: New evidence on algorithm performance and quality metrics,” PLoS One, vol. 17, no. 4, p. e0266325, Apr. 2022, doi: 10.1371/JOURNAL.PONE.0266325.

[8] A. Rkia, A. Fatima-Azzahrae, A. Mehdi, and L. Lily, “NLP and Topic Modeling with LDA, LSA, and NMF for Monitoring Psychosocial Well-being in Monthly Surveys,” Procedia Comput Sci, vol. 251, pp. 398–405, Jan. 2024, doi: 10.1016/J.PROCS.2024.11.126.

[9] V. Leplat, L. T. K. Hien, A. Onwunta, and N. Gillis, “Deep Nonnegative Matrix Factorization With Beta Divergences,” Neural Comput, vol. 36, no. 11, pp. 2365–2402, Oct. 2024, doi: 10.1162/NECO_A_01679.

[10] M. F. Asnawi, M. Hanafi, N. F. Kurniawan, A. Suwondo, A. Nasrullah, and C. Setyawan, “Topic Modelling Analysis on Indonesian News Using BERT Topic Model,” 2024 6th International Conference on Cybernetics and Intelligent System, ICORIS 2024, 2024, doi: 10.1109/ICORIS63540.2024.10903779.

[11] D. Medvecki, B. Bašaragin, A. Ljajić, and N. Milošević, “Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian,” pp. 161–173, Feb. 2024, doi: 10.1007/978-3-031-50755-7_16.

[12] C. P. Chai, “Comparison of text preprocessing methods,” Nat Lang Eng, vol. 29, no. 3, pp. 509–553, May 2023, doi: 10.1017/S1351324922000213.

[13] M. Nesca, A. Katz, C. K. Leung, and L. M. Lix, “A scoping review of preprocessing methods for unstructured text data to assess data quality,” Int J Popul Data Sci, vol. 7, no. 1, p. 1757, 2022, doi: 10.23889/IJPDS.V6I1.1757.

[14] M. Risky and A. Aprillia, “Komparasi Algoritma Klasifikasi Dan Penerapan Ner pada Analisis Sentimen Bencana Alam Banjir,” InComTech : Jurnal Telekomunikasi dan Komputer, vol. 14, no. 2, pp. 130–139, Aug. 2024, doi: 10.22441/INCOMTECH.V14I2.22962.

[15] R. Ariftiarno and M. Dwijayanti, “Analisis Sentimen Netizen Twitter Terhadap Pelayanan Provider Telkomsel: Komparasi Naive-Bayes dan K-Nearest Neighbors,” InComTech : Jurnal Telekomunikasi dan Komputer, vol. 14, no. 1, pp. 60–77, May 2024, doi: 10.22441/INCOMTECH.V14I1.22597.

[16] C. P. Chai, “Comparison of text preprocessing methods,” Nat Lang Eng, vol. 29, no. 3, pp. 509–553, May 2023, doi: 10.1017/S1351324922000213.

[17] M. Nesca, A. Katz, C. K. Leung, and L. M. Lix, “A scoping review of preprocessing methods for unstructured text data to assess data quality,” Int J Popul Data Sci, vol. 7, no. 1, p. 1757, 2022, doi: 10.23889/IJPDS.V6I1.1757.

[18] S. Sarica and J. Luo, “Stopwords in technical language processing,” PLoS One, vol. 16, no. 8, p. e0254937, Aug. 2021, doi: 10.1371/JOURNAL.PONE.0254937.

[19] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Inf Syst, vol. 121, p. 102342, Mar. 2024, doi: 10.1016/J.IS.2023.102342.

[20] R. N. Rathi and A. Mustafi, “The importance of Term Weighting in semantic understanding of text: A review of techniques,” Multimed Tools Appl, vol. 82, no. 7, pp. 9761–9783, Mar. 2023, doi: 10.1007/S11042-022-12538-3/METRICS.

[21] Y. Setiawan, N. U. Maulidevi, and K. Surendro, “The Optimization of n-Gram Feature Extraction Based on Term Occurrence for Cyberbullying Classification,” Data Sci J, vol. 23, no. 1, 2024, doi: 10.5334/DSJ-2024-031.

[22] M. Liang and T. Niu, “Research on Text Classification Techniques Based on Improved TF-IDF Algorithm and LSTM Inputs,” Procedia Comput Sci, vol. 208, pp. 460–470, Jan. 2022, doi: 10.1016/J.PROCS.2022.10.064.

[23] A. Abdelrazek, Y. Eid, E. Gawish, W. Medhat, and A. Hassan, “Topic modeling algorithms and applications: A survey,” Inf Syst, vol. 112, p. 102131, Feb. 2023, doi: 10.1016/J.IS.2022.102131.

[24] A. Simonetti, A. Albano, A. Plaia, and M. Tumminello, “Ranking coherence in topic models using statistically validated networks,” J Inf Sci, vol. 51, no. 3, pp. 744–765, Jun. 2025, doi: 10.1177/01655515221148369.

[25] Y. Wu, C. H. Tseng, J. Shang, S. Mao, G. Nenadic, and X. J. Zeng, “EDU-level Extractive Summarization with Varying Summary Lengths,” EACL 2023 - 17th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2023, pp. 1655–1667, 2023, doi: 10.18653/V1/2023.FINDINGS-EACL.123.

[26] R. Li, F. González-Pizarro, L. Xing, G. Murray, and G. Carenini, “Diversity-Aware Coherence Loss for Improving Neural Topic Models,” Proceedings of the Annual Meeting of the Association for Computational Linguistics, vol. 2, pp. 1710–1722, May 2023, doi: 10.18653/v1/2023.acl-short.145.

[27] S. Syahrial, R. Perucha, and F. Afidh, “Fine-Tuning Topic Modelling: A Coherence-Focused Analysis of Correlated Topic Models,” Infolitika Journal of Data Science, vol. 2, no. 2, pp. 82–87, Nov. 2024, doi: 10.60084/IJDS.V2I2.236.

[28] D. Xu et al., “Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding,” EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024, pp. 7744–7757, 2024, doi: 10.18653/V1/2024.FINDINGS-EMNLP.456.

[29] H. Roris and P. Sianturi, “Politics on a plate: equivocal communication in Indonesian presidential nutrition policy,” Front Commun (Lausanne), vol. 10, p. 1612652, Sep. 2025, doi: 10.3389/FCOMM.2025.1612652.

[30] G. Agarwal, S. K. Dinkar, and A. Agarwal, “Binarized spiking neural networks optimized with Nomadic People Optimization-based sentiment analysis for social product recommendation,” Knowl Inf Syst, vol. 66, no. 2, pp. 933–958, Feb. 2024, doi: 10.1007/S10115-023-01956-W.

[31] V. N. N. Tsara and Y. Purbokusumo, “Framing Analysis of Free Nutritious Meal Program (MBG) News on Detik.com: Implications for Public Perception and Government Policy,” Universitas Gadjah Mada, Yogyakarta, 2025.

Published

2026-08-31

How to Cite

[1]
I. Dwijayanti and M. Habibi, “Topic Modeling of Indonesian News Articles on The Free Nutritious Meal Program: A Comparative Study of LDA, LSA, and NMF”, InComTech, vol. 16, no. 2, pp. 128–142, Aug. 2026.

Issue

Section

Articles