LLM Token Consumption Optimization Using BERT Pre-filtering and Retrieval-Augmented Generation (RAG)

Authors

  • Nanang Kurnia Alfian Master of Informatics Engineering, Universitas Amikom Yogyakarta
  • Robert Marco Master of Informatics Engineering, Universitas Amikom Yogyakarta

Keywords:

Token Optimization, BERT, RAG, Bayesian Optimization, Text Classification

Abstract

The adoption of Large Language Models (LLM) for software issue resolution causes token consumption inefficiency due to noisy and redundant context. This research proposes a hybrid system placing a BERT classification model as a context relevance pre-filter upstream of the Retrieval-Augmented Generation (RAG) pipeline, with hyperparameters optimized using Bayesian Optimization. A Design Science Research approach is applied through an ablation study of five scenarios (S0-S4) on GitHub Issues and SWE-bench datasets. The classification model achieves an F1-macro of 0.887, and BERT-based filtering (S2) reduces token consumption by 8.35% versus baseline, statistically significant (Wilcoxon, p<0.05). Bayesian Optimization yields a validation F1-macro of 0.889, higher than default hyperparameters. The main finding: token efficiency comes from the BERT pre-filter component, not from RAG which adds context.

References

P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," Adv. Neural Inf. Process. Syst., vol. 33, pp. 9459-9474, 2020.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional trans-formers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171-4186.

T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, "Optuna: A next-generation hyperparame-ter optimization framework," in Proc. 25th ACM SIGKDD, 2019, pp. 2623-2631.

C. E. Jimenez et al., "SWE-bench: Can language models resolve real-world GitHub issues?," in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.

N. F. Liu et al., "Lost in the middle: How language models use long contexts," Trans. Assoc. Com-put. Linguist., vol. 12, pp. 157-173, 2024.

A. Sabbatella, A. Ponti, I. Giordani, A. Candelieri, and F. Archetti, "Prompt optimization in large lan-guage models," Mathematics, vol. 12, no. 6, art. 929, 2024.

A. Vaswani et al., "Attention is all you need," Adv. Neural Inf. Process. Syst., vol. 30, pp. 5998-6008, 2017.

B. Lester, R. Al-Rfou, and N. Constant, "The power of scale for parameter-efficient prompt tuning," arXiv:2104.08691, 2021.

E. J. Hu et al., "LoRA: Low-rank adaptation of large language models," arXiv:2106.09685, 2021.

A. Archetti and A. Candelieri, Bayesian Optimization and Data Science. Berlin: Springer, 2019.

M. Balandat et al., "BoTorch: A framework for efficient Monte-Carlo Bayesian optimization," Adv. Neural Inf. Process. Syst., vol. 33, pp. 21524-21538, 2020.

N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddings using Siamese BERT-networks," in Proc. EMNLP-IJCNLP, 2019, pp. 3982-3992.

A. R. Hevner, S. T. March, J. Park, and S. Ram, "Design science in information systems research," MIS Quarterly, vol. 28, no. 1, pp. 75-105, 2004.

Y. Gao et al., "Retrieval-augmented generation for large language models: A survey," IEEE Trans. Knowl. Data Eng., 2025.

J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," Adv. Neural Inf. Process. Syst., vol. 35, pp. 24824-24837, 2022.

Downloads

Published

2026-09-28

How to Cite

Nanang Kurnia Alfian, & Robert Marco. (2026). LLM Token Consumption Optimization Using BERT Pre-filtering and Retrieval-Augmented Generation (RAG). Jurnal Ilmiah Multidisiplin Indonesia (JIM-ID), 5(09), 2342–2347. Retrieved from https://ejournal.seaninstitute.or.id/index.php/esaprom/article/view/8886