Public Approval Rating Updates for President Prabowo Using a Bayesian Dirichlet-Multinomial Hybrid Model

  • anwar anwar anwar unismuh
  • Andries Riesfandhy Faculty of Politics and Law, Political Science Program, Universitas Andi Sudirman
Keywords: Bayesian Dirichlet-Multinomial; Public Opinion Survey; Social Media Monitoring; Sentiment Analysis; IndoBERT; Indonesia.

Abstract

Conventional estimation of presidential approval ratings relies on face-to-face surveys that excel in population representativeness (probability sampling) but suffer from significant time lag between data collection and publication. Conversely, social media monitoring (SMM) provides real-time data but is often biased toward active digital platform demographics. This study proposes a hybrid model based on Bayesian Dirichlet-Multinomial (BDM) to integrate survey data as prior belief with SMM data as likelihood to generate more dynamic and accurate posterior estimates. Prior data were drawn from the Indikator Politik Indonesia national survey of October 2025 (n=1,220, MoE ±2.9%), indicating a public approval rate of 77.7%. Update data were collected from the X (Twitter) platform during November 2025 through crawling and scraping techniques, yielding 2.5 million raw data points processed through Named Entity Recognition (NER), text normalization, and Transformer-based classification (IndoBERT v2.0). After deduplication, 310,000 unique accounts were retained with 120 million total impressions. Sentiment distribution showed 61% positive, 15% negative, and 24% neutral. Bayesian updating produced a corrected public approval estimate of 68.4% (a 9.3 percentage point decline from the survey prior), detecting pockets of dissatisfaction on food price issues underrepresented in conventional surveys. Model validation using Leave-One-Out Cross-Validation (LOO-CV) produced an Expected Log-Pointwise Predictive Density (ELPD) of -127.4, superior to the single-survey baseline model (-148.9). This model offers a new framework as a public opinion early warning system that is responsive to current issue dynamics without sacrificing the statistical validity of traditional survey methods.

References

Association of Internet Researchers. (2019). Ethical decision-making and Internet research: Recommendations from the AoIR ethics working committee (Version 3.0). https://aoir.org/reports/ethics3.pdf

Badan Pusat Statistik. (2025). Profil kependudukan Indonesia 2025. Badan Pusat Statistik.

Barberá, P. (2015). Birds of the same feather tweet together: Bayesian ideal point estimation using Twitter data. Political Analysis, 23(1), 76–91. https://doi.org/10.1093/pan/mpu011

Brosius, H. B., & Weimann, G. (1996). Who sets the agenda? Agenda-setting as a two-step flow. Communication Research, 23(5), 561–580.

Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Kuncoro, A., Mahendra, R., & Fung, P. (2021). IndoNLI: A natural language inference dataset for Indonesian. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 10511–10527).

Ceron, A., Curini, L., Iacus, S. M., & Porro, G. (2014). Every tweet counts? How sentiment analysis of social media can improve our knowledge of citizens' political preferences. New Media & Society, 16(2), 340–358.

Colleoni, E., Rozza, A., & Arvidsson, A. (2014). Echo chamber or public sphere? Predicting political orientation and measuring political homophily in Twitter using big data. Journal of Communication, 64(2), 317–332.

Dahl, R. A. (2015). On democracy (2nd ed.). Yale University Press.

Fathony, R., Winata, G. I., & Purwarianti, A. (2023). Indonesian social media language normalization corpus: A crowdsourced approach. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 4521–4530).

Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian data analysis (3rd ed.). CRC Press.

Holbrook, A. L., & Krosnick, J. A. (2010). Social desirability bias in voter turnout reports: Tests using the item count technique. Public Opinion Quarterly, 74(1), 37–67.

Indikator Politik Indonesia. (2025, October). Tren kepuasan publik terhadap kinerja presiden: Survei Oktober 2025.

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770).

Kurniawan, F., & Louvan, S. (2022). IndoBERT: A state-of-the-art language model for Indonesian. In Proceedings of the 13th International Conference on Language Resources and Evaluation (pp. 3541–3549).

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.

Mellon, J., & Prosser, C. (2017). Twitter and Facebook are not representative of the general population: Political attitudes and demographics of British social media users. Research & Politics, 4(3), 1–10.

Neuman, W. R., Guggenheim, L., Jang, S. M., & Bae, S. Y. (2014). The dynamics of public attention: Agenda-setting theory meets big data. Journal of Communication, 64(2), 193–214.

O'Connor, B., Balasubramanyan, R., Routledge, B. R., & Smith, N. A. (2010). From tweets to polls: Linking text sentiment to public opinion time series. In Proceedings of the Fourth International AAAI Conference on Web and Social Media (pp. 122–129).

Pasek, J. (2018). It's not my consensus: Motivated reasoning and the sources of scientific illiteracy. Public Understanding of Science, 27(7), 787–806.

Röder, M., Both, A., & Hinneburg, A. (2015). Exploring the space of topic coherence measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining (pp. 399–408).

We Are Social, & Meltwater. (2025). Digital 2025 Indonesia report. https://wearesocial.com/id/blog/2025/01/digital-2025

Published
2026-07-30
Abstract viewed = 55 times
PDF downloaded = 14 times