Sequran: Composite Scoring and Reranking Techniques for Refining Quran Search Engine Results

Authors

  • Ray Ramadita UIN Sunan Gunung Djati Bandung
  • Wisnu Uriawan UIN Sunan Gunung Djati Bandung
  • Wildan Budiawan Zulfikar UIN Sunan Gunung Djati Bandung

Keywords:

Composite Scoring, Reranking Techniques, Information Retrieval, Search Engine, Quran

Abstract

Information Retrieval (IR) systems are crucial in the development of accurate and relevant search engines, especially in domain-specific applications such as Sequran. While fine-tuning is a common optimization approach, this method often requires significant computational resources, diverse datasets, and time-consuming hyperparameter tuning, with the risk of performance degradation. To address these challenges, this research introduces and evaluates a novel architecture that enhances search relevance by integrating lexical and semantic signals, offering a practical alternative to resourceintensive fine-tuning. The proposed architecture involves a two-stage process. The first stage, composite scoring, enhances term frequency (BM25S) with a semantic intent booster. The second stage utilizes a Cross-Encoder (jina-reranker-v2-base-multilingual) to refine the ranking based on contextual relevance. Evaluation is conducted on a domain-specific Islamic query-answer dataset consisting of approximately 20 queries and around 18,000 text entries (over 6,000 verses, each paired with two tafsir types). The proposed method achieves Precision@10 = 0.230 and Recall@10 = 0.564 over the fine-tuned baseline (0.167/0.423), representing absolute gains of 6.3 and 14.1 percentage points, equivalent to relative improvements of 37.7% and 33.3%. Gains are consistent across both metrics, with the larger recall gain indicating that reranking surfaces relevant passages ranked below the cut-off by lexical matching alone. This improvement demonstrates a clear trade-off, as the total execution time increased to approximately 18.5 seconds, which may constrain realtime use. The main implication of this research is the validation of a practical architecture for improving IR systems, offering a viable alternative for domain-specific contexts such as Sequran.

Dimensions

Author Biographies

Ray Ramadita, UIN Sunan Gunung Djati Bandung

Informatics Department, Faculty of Science and Technology

Wisnu Uriawan, UIN Sunan Gunung Djati Bandung

Informatics Department, Faculty of Science and Technology

Wildan Budiawan Zulfikar, UIN Sunan Gunung Djati Bandung

Informatics Department, Faculty of Science and Technology

References

[1] B. K. Hussan, “Comparative study of semantic and keyword based search engines,” Advances in Science, Technology and Engineering Systems Journal, vol. 5, no. 1, pp. 106–111, 2020.

[2] K. K. Dukhnovska and I. L. Myshko, “Analysis and comparison of full-text search algorithms,” Control Systems & Computers, no. 3, pp. 45–52, 2024.

[3] M. M. Al Haromainy, A. P. Sari, D. A. Prasetya, M. Subhan, A. Lisdiyanto, and T. Septianto, “Enhancing thematic holy Quran verse retrieval through vector space model and query expansion for effective query answering,” in 2023 IEEE 9th Information Technology International Seminar (ITIS). Batu Malang, Indonesia: IEEE, Oct. 18–20, 2023, pp. 1–6.

[4] S. E. Pratama, W. Darmalaksana, D. S. Maylawati, H. Sugilar, T. Mantoro, and M. A. Ramdhani, “Weighted inverse document frequency and vector space model for hadith search engine,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 18, no. 2, pp. 1004–1014, 2020.

[5] J. Chamorro-Padial, F.-J. Rodrigo-Gin´es, and R. Rodr´ıguez-S´anchez, “Finding answers to COVID-19-specific questions: An information retrieval system based on latent keywords and adapted TF-IDF,” Journal of Information Science, vol. 50, no. 4, pp. 935–951, 2024.

[6] F. Beirade, H. Azzoune, and D. E. Zegour, “Semantic query for Quranic ontology,” Journal of King Saud University Computer and Information Sciences, vol. 33, no. 6, pp. 753–760, 2021.

[7] S. Hakak, G. A. Gilkar, and W. Z. Khan, “Performance comparison of Qur’anic search engines,” in 2020 International Conference on Computing and Information Technology (ICCIT-1441). Tabuk, Saudi Arabia: IEEE, Sept. 9–10, 2020, pp. 1–4.

[8] I. K. Fitriani, M. A. Bijaksana, and K. M. Lhaksmana, “Qur’an search system for handling cross verse based on phonetic similarity,” Jurnal Sisfokom (Sistem Informasi dan Komputer), vol. 10, no. 1, pp. 46–51, 2021.

[9] N. I. Purwita, M. A. Bijaksana, K. M. Lhaksmana, and M. Z. Naf’an, “Typo handling in searching of Quran verse based on phonetic similarities,” Register: Jurnal Ilmiah Teknologi Sistem Informasi, vol. 6, no. 2, pp. 130–140, 2020.

[10] A. Octavia, M. A. Bijaksana, and K. M. Lhaksmana, “Verse search system for sound differences in the Qur’an based on the text of phonetic similarities,” Jurnal Sisfokom (Sistem Informasi dan Komputer), vol. 9, no. 3, pp. 317–322, 2020.

[11] J. Wang, M. Pan, T. He, X. Huang, X. Wang, and X. Tu, “A pseudo-relevance feedback framework combining relevance matching and semantic matching for information retrieval,” Information Processing & Management, vol. 57, no. 6, 2020.

[12] M. Wang, J. Liu, J. Wang, Y. Wang, and X. Chu, “A topicality relevance-aware intent model for web search,” IEEE Access, vol. 11, pp. 65 739–65 748, 2023.

[13] M. C. Avornicului, V. P. Bresfelean, S. C. Popa, N. Forman, and C. A. Comes, “Designing a prototype platform for real-time event extraction: A scalable natural language processing and data mining approach,” Electronics, vol. 13, no. 24, pp. 1–27, 2024.

[14] K. A. Hambarde and H. Proenca, “Information retrieval: Recent advances and beyond,” IEEE Access, vol. 11, pp. 76 581–76 604, 2023.

[15] S. Bruch, S. Gai, and A. Ingber, “An analysis of fusion functions for hybrid retrieval,” ACM Transactions on Information Systems, vol. 42, no. 1, pp. 1–35, 2023.

[16] A. Esteva et al., “COVID-19 information retrieval with deep-learning based semantic search, question answering, and abstractive summarization,” npj Digital Medicine, vol. 4, pp. 1–9, 2021.

[17] A. Kouadria, O. Nouali, and M. Y. H. Al-Shamri, “A multi-criteria collaborative filtering recommender system using learning-to-rank and rank aggregation,” Arabian Journal for Science and Engineering, vol. 45, no. 4, pp. 2835–2845, 2020.

[18] M. Esposito, E. Damiano, A. Minutolo, G. De Pietro, and H. Fujita, “Hybrid query expansion using lexical resources and word embeddings for sentence retrieval in question answering,” Information Sciences, vol. 514, pp. 88–105, 2020.

[19] J. Zhan, J. Mao, Y. Liu, J. Guo, M. Zhang, and S. Ma, “Optimizing dense retrieval model training with hard negatives,” in Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, July 11–15, 2021, pp. 1503–1512.

[20] R. Upadhyay, A. Askari, G. Pasi, and M. Viviani, “Beyond topicality: Including multidimensional relevance in cross-encoder re-ranking: The health misinformation case study,” in European Conference on Information Retrieval. Glasgow, UK: Springer, March 24–28, 2024, pp. 262–277.

[21] A. Brandsen, S. Verberne, K. Lambers, and M. Wansleeben, “Can BERT dig it? Named entity recognition for information retrieval in the archaeology domain,” Journal on Computing and Cultural Heritage (JOCCH), vol. 15, no. 3, pp. 1–18, 2022.

[22] P. Su and K. Vijay-Shanker, “Investigation of improving the pre-training and fine-tuning of BERT model for biomedical relation extraction,” BMC Bioinformatics, vol. 23, pp. 1–20, 2022.

[23] H. Soudani, E. Kanoulas, and F. Hasibi, “Fine tuning vs. retrieval augmented generation for less popular knowledge,” in SIGIR-AP 2024: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region. Tokyo, Japan: Association for Computing Machinery, Dec. 9–12, 2024, pp. 12–22.

[24] X. Ma, J. Guo, R. Zhang, Y. Fan, and X. Cheng, “Scattered or connected? An optimized parameter-efficient tuning approach for information retrieval,” in Proceedings of the 31st ACM International Conference on Information & Knowledge Management. Atlanta, GA, USA: Association for Computing Machinery, Oct. 17–21, 2022, pp. 1471–1480.

[25] H. Chen and P. N. Garner, “Bayesian parameterefficient fine-tuning for overcoming catastrophic forgetting,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 4253–4262, 2024.

[26] Y. Zhai et al., “Investigating the catastrophic forgetting in multimodal large language model fine-tuning,” in Conference on Parsimony and Learning. Hongkong, China: PMLR, Jan. 3–6, 2024, pp. 202–227.

[27] C. H. Tu et al., “Holistic transfer: Towards nondisruptive fine-tuning with partial target data,” Advances in Neural Information Processing Systems, vol. 36, pp. 29 149–29 173, 2023.

[28] M. Wortsman et al., “Robust fine-tuning of zeroshot models,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA: IEEE, June 18–24, 2022, pp. 7949–7961.

[29] M. Y. Kim, J. Rabelo, K. Okeke, and R. Goebel, “Legal information retrieval and entailment based on BM25, transformer and semantic thesaurus methods,” The Review of Socionetwork Strategies, vol. 16, pp. 157–174, 2022.

[30] R. Ramadita et al., “Sequran: Sentence-BERT for specific domain of Quran search engine,” 2024, presented in International Invention Competition for Young Moslem Scientists 2024.

[31] X. Zhang et al., “MIRACL: A multilingual retrieval dataset covering 18 diverse languages,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 1114–1131, 2023.

[32] X. Zhang, X. Ma, P. Shi, and J. Lin, “Mr. TyDi: A multi-lingual benchmark for dense retrieval,” in Proceedings of the 1st Workshop on Multilingual Representation Learning. Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 127–137.

[33] J. H. Clark et al., “Tydi QA: A benchmark for information-seeking question answering in ty pologically di verse languages,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 454–470, 2020.

[34] Sequran, “Indo-Islamic queries and answers dataset,” 2025. [Online]. Available: https://huggingface.co/datasets/ramadita/Indo-Islamic-QA

[35] A. Pauli, L. Derczynski, and I. Assent, “Anchoring fine-tuning of sentence transformer with semantic label information for efficient truly fewshot classification,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 11 254–11 264.

[36] T. Sultana, A. K. Mandal, H. Saha, M. N. Sultan, and M. D. Hossain, “Intent identification by semantically analyzing the search query,” Modelling, vol. 5, no. 1, pp. 292–314, 2024.

[37] N. L. P. I. Candrawengi and M. W. P. Dananjaya, “Machine learning approaches for search intentdriven website optimization,” in 2024 10th International Conference on Smart Computing and Communication (ICSCC). Bali, Indonesia: IEEE, July 25–27, 2024, pp. 197–202.

[38] C. D. Schultz, “Informational, transactional, and navigational need of information: Relevance of search intention in search engine advertising,” Information Retrieval Journal, vol. 23, no. 2, pp. 117–135, 2020.

[39] F. Rahman, Tema-tema pokok Al-Quran. Al Mizan, 2018.

[40] C. P. Harshitha and N. R. Sunitha, “Topic identification for semantic grouping based on hidden Markov model,” in 2020 5th International Conference on Communication and Electronics Systems (ICCES). Coimbatore, India: IEEE, June 10–12, 2020, pp. 932–937.

[41] T. D. Jayasiriwardene and G. U. Ganegoda, “Keyword extraction from Tweets using NLP tools for collecting relevant news,” in 2020 International Research Conference on Smart Computing and Systems Engineering (SCSE). Colombo, Sri Lanka: IEEE, Sep. 24, 2020, pp. 129–135.

[42] R. Khatun and A. Sarkar, “Deep-KeywordNet: Automated English keyword extraction in documents using deep keyword network based ranking,” Multimedia Tools and Applications, pp. 68 959–68 991, 2024.

[43] R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, and A. Jatowt, “YAKE! Keyword extraction from single documents using multiple local features,” Information Sciences, vol. 509, pp. 257–289, 2020.

[44] A. Gupta, A. Chadha, and V. Tewari, “A natural language processing model on BERT and YAKE technique for keyword extraction on sustainability reports,” IEEE Access, vol. 12, pp. 7942–7951, 2024.

[45] R. A. Yunmar, A. Setiawan, and H. Tantriawan, “The combination of YAKE and language processing for unsupervised term extraction ontology learning,” in IOP Conference Series: Earth and Environmental Science, vol. 537. IOP Publishing, 2020.

[46] M. Grootendorst et al., “MaartenGr/KeyBERT: v0.9,” 2025. [Online]. Available: https://zenodo.org/records/14831372

[47] J. Sammet and R. Krestel, “Domain-specific keyword extraction using BERT,” in Proceedings of the 4th Conference on Language, Data and Knowledge. Vienna, Austria: NOVA CLUNL, 2023, pp. 659–665.

[48] B. Issa, M. B. Jasser, H. N. Chua, and M. Hamzah, “A comparative study on embedding models for keyword extraction using KeyBERT method,” in 2023 IEEE 13th International Conference on System Engineering and Technology (ICSET). Shah Alam, Malaysia: IEEE, Oct. 2, 2023, pp. 40–45.

[49] M. Liu, X. Luo, G. Wang, and W. Z. Lu, “Intelligent information extraction from government on-site inspection reports of construction projects: A graph-based text mining approach,” Advanced Engineering Informatics, vol. 58, 2023.

[50] K. Liu, Y. Li, Y. Qi, N. Qi, and M. Zhai, “Text information mining in cyberspace: An information extraction method based on T5 and KeyBERT,” in 2024 IEEE 9th International Conference on Data Science in Cyberspace (DSC). Jinan, China: IEEE, Aug. 23–26, 2024, pp. 621–628.

[51] E. T. Luthfi, Z. I. M. Yusoh, and B. M. Aboobaider, “BERT based named entity recognition for automated Hadith narrator identification,” International Journal of Advanced Computer Science and Applications, vol. 13, no. 1, pp. 604–611, 2022.

[52] S. L. Chen, P. J. Burns, T. J. Bolt, P. Chaudhuri, and J. P. Dexter, “Leveraging part-of-speech tagging for enhanced stylometry of Latin literature,” in Proceedings of the 1st Workshop on Machine Learning for Ancient Languages (ML4AL 2024). Hybrid in Bangkok, Thailand and Online: Association for Computational Linguistics, Aug. 2024, pp. 251–259.

[53] S. Jin, S. Chen, and X. Xie, “Property-based test for part-of-speech tagging tool,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE). Melbourne, Australia: IEEE, Nov. 15–19, 2021, pp. 1306–1311.

[54] M. Nadim, D. Akopian, and A. Matamoros, “A comparative assessment of unsupervised keyword extraction tools,” IEEE Access, vol. 11, pp. 144 778–144 798, 2023.

[55] T. Taipalus, “Vector database management systems: Fundamental concepts, use-cases, and current challenges,” Cognitive Systems Research, vol. 85, pp. 1–8, 2024.

[56] J. J. Pan, J. Wang, and G. Li, “Survey of vector database management systems,” The VLDB Journal, vol. 33, pp. 1591–1615, 2024.

[57] X. H. Lu, “BM25S: Orders of magnitude faster lexical search via eager sparse scoring,” 2024. [Online]. Available: https://arxiv.org/abs/2407.03618

[58] D. Teodoro et al., “Information retrieval in an infodemic: The case of COVID-19 publications,” Journal of Medical Internet Research, vol. 23, no. 9, 2021.

[59] K. Pfleger and B. Larson, “System and method for determining a composite score for categorized search results,” 2010.

[60] N. Reimers and I. Gurevych, “The curse of dense low-dimensional information retrieval for large index sizes,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Online: Association for Computational Linguistics, Aug. 2021, pp. 605–611.

[61] L. C. Chen, “A study of optimizing search engine results through user interaction,” IEEE Access, vol. 8, pp. 79 024–79 045, 2020.

[62] M. D. Almadhoun and N. H. A. H. Malim, “Effects of using multi-category web pages on rank estimation of Google search engine results page,” Web Intelligence, vol. 23, no. 1, pp. 39–55, 2025.

[63] A. Urman and M. Makhortykh, “You are how (and where) you search? Comparative analysis of web search behavior using web tracking data,” Journal of Computational Social Science, vol. 6, pp. 741–756, 2023.

[64] N. Craswell, B. Mitra, E. Yilmaz, D. Campos, E. M. Voorhees, and I. Soboroff, “TREC deep learning track: Reusable test collections in the large data regime,” in Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, July 11–15, 2021, pp. 2369–2375.

[65] E. Bassani, “ranx: A blazing-fast python library for ranking evaluation and comparison,” in European Conference on Information Retrieval. Stavanger, Norway: Springer, April 10–11, 2022, pp. 259–264.

[66] Z. H. Amur, Y. K. Hooi, G. M. Soomro, H. Bhanbhro, S. Karyem, and N. Sohu, “Unlocking the potential of keyword extraction: The need for access to high-quality datasets,” Applied Sciences, vol. 13, no. 12, pp. 1–19, 2023.

[67] M. Wrzalik and D. Krechel, “CoRT: Complementary rankings from transformers,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Association for Computational Linguistics, June 2021, pp. 4194–4204.

[68] S. Chang, G. J. Ahn, and S. Park, “Improving performance of neural IR models by using a keyword-extraction-based weak-supervision method,” IEEE Access, vol. 12, pp. 46 851–46 863, 2024.

[69] G. Deena and K. Raja, “Keyword extraction using latent semantic analysis for question generation,” Journal of Applied Science and Engineering, vol. 26, no. 4, pp. 501–510, 2022.

[70] B. Plank, “Sliced at SemEval-2022 Task 11: Bigger, better? Massively multilingual LMs for multilingual complex NER on an academic GPU budget,” in Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022). Seattle, United States: Association for Computational Linguistics, July 2022, pp. 1494–1500.

[71] L. Gao, Z. Dai, T. Chen, Z. Fan, B. Van Durme, and J. Callan, “Complement lexical retrieval model with semantic residual embeddings,” in European Conference on Information Retrieval. Virtual: Springer, March 28–April 1, 2021, pp. 146–160.

Downloads

Published

2026-09-08

How to Cite

[1]
R. Ramadita, W. Uriawan, and W. Budiawan Zulfikar, “Sequran: Composite Scoring and Reranking Techniques for Refining Quran Search Engine Results”, CommIT (Communication and Information Technology) Journal, vol. 20, no. 2, pp. 401–413, Sep. 2026.
Abstract 44  .
PDF downloaded 32  .