Comparative Analysis of Knowledge Internalization through Supervised Fine-Tuning and Contextual Reasoning through Retrieval-Augmented Fine-Tuning in Large Language Models for Indonesian Village Regulations

Syahrial Jeremia Sinaga, Vivaldi Adventus Simangunsong, Oppir Hutapea, Yohanssen Pratama4, Adli Abdillah Nababan

Abstract


Large language models (LLMs) are prone to hallucination, making their reliability in Indonesian-language legal question answering uncertain, particularly for Village Regulations, which are predominantly available as scanned PDF documents. This study provides a controlled comparison between knowledge internalization through Supervised Fine-Tuning (SFT) and contextual reasoning through Retrieval-Augmented Fine-Tuning (RAFT), using both LoRA and full fine-tuning from the same Llama-3.1-8B-Instruct checkpoint. The study also develops a traceable document-processing pipeline and a question-answer dataset derived from 55 Village Regulations from Badung Regency, consisting of 50 scanned and five born-digital documents. Model performance was evaluated using ROUGE-L, METEOR, BERTScore, and RAGAS. At the fourth epoch, RAFT-LoRA achieved the highest aggregate score (0.883), followed by RAFT-Full (0.881), while the best-performing SFT configuration reached 0.749. RAFT improved faithfulness from 0.640 to 0.872 and context recall from 0.637 to 0.907. The results indicate that training data structure had a greater impact on performance than the parameter-update strategy, while LoRA maintained RAFT performance with lower adaptation costs. These findings suggest that legal knowledge can be maintained in updatable external documents without requiring model retraining, while the model can be trained to generate evidence-grounded responses. However, the findings are limited to a single regency and automated evaluation; therefore, validation by legal experts is required before deployment.

Keywords


hallucination; large language model; legal question answering; retrieval-augmented fine-tuning; supervised fine-tuning; village regulations

Full Text:

PDF

References


A. Matarazzo and R. Torlone, "A Survey on Large Language Models with Some Insights on Their Capabilities and Limitations," arXiv preprint arXiv:2501.04040, 2025, DOI: 10.48550/arXiv.2501.04040.

A. Vaswani et al., "Attention is All You Need," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 30, 2017, pp. 5998–6008, DOI: 10.48550/arXiv.1706.03762.

Z. Ji et al., "Survey of Hallucination in Natural Language Generation," ACM Comput. Surv., Vol. 55, No. 12, pp. 1–38, Mar. 2023, DOI: 10.1145/3571730.

L. Huang et al., "A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions," ACM Trans. Inf. Syst., Vol. 43, No. 2, Jan. 2025, DOI: 10.1145/3703155.

S. Cahyawijaya et al., "Cendol: Open Instruction-Tuned Generative Large Language Models for Indonesian Languages," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (Volume 1: Long Papers), Bangkok, Thailand, Aug. 2024, pp. 14899–14914, DOI: 10.18653/v1/2024.acl-long.796.

Republik Indonesia, Undang-Undang Nomor 6 Tahun 2014 tentang Desa, Lembaran Negara Republik Indonesia Tahun 2014 Nomor 7, 2014.

Republik Indonesia, Undang-Undang Nomor 12 Tahun 2011 tentang Pembentukan Peraturan Perundang-undangan, 2011.

Republik Indonesia, Peraturan Presiden Nomor 33 Tahun 2012 tentang Jaringan Dokumentasi dan Informasi Hukum Nasional, 2012.

Z. Gekhman et al., "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?," in Proc. 2024 Conf. Empirical Methods Natural Lang. Process. (EMNLP), Miami, FL, USA, Nov. 2024, pp. 7765–7784, DOI: 10.18653/v1/2024.emnlp-main.444.

O. Ovadia, M. Brief, M. Mishaeli, and O. Elisha, "Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs," arXiv preprint arXiv:2312.05934, Dec. 2023, DOI: 10.48550/arXiv.2312.05934.

T. Zhang et al., "RAFT: Adapting Language Model to Domain Specific RAG," arXiv preprint arXiv:2403.10131, 2024, DOI: 10.48550/arXiv.2403.10131.

S. Gupta, V. Agrawal, Y. Negi, S. Karunanithi, and A. Balakrishnan, "Reducing Hallucinations in Legal AI: A Retrieval Augmented Generation-based Model for Accurate Legal Guidance," IEEE Access, Vol. 14, pp. 44451–44463, 2026, DOI: 10.1109/ACCESS.2026.3675624.

P. Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP tasks," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 33, 2020, pp. 9459–9474, DOI: 10.48550/arXiv.2005.11401.

K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, "Retrieval Augmentation Reduces Hallucination in Conversation," in Findings Assoc. Comput. Linguistics: EMNLP 2021, Punta Cana, Dominican Republic, Nov. 2021, pp. 3784–3803, DOI: 10.18653/v1/2021.findings-emnlp.320.

A. Grattafiori et al., "The Llama 3 Herd of Models," arXiv preprint arXiv:2407.21783, 2024, DOI: 10.48550/arXiv.2407.21783.

M. Turski, T. Stanisławek, K. Kaczmarek, P. Dyda, and F. Graliński, "CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data," in Document Analysis and Recognition – ICDAR 2023, Lecture Notes in Computer Science, Vol. 14189. Cham, Switzerland: Springer, 2023, pp. 348–365, DOI: 10.48550/arXiv.2304.14953.

R. Sinha and R. B S, "Digitization of Document and Information Extraction using OCR," arXiv preprint arXiv:2506.11156, 2025, DOI: 10.48550/arXiv.2506.11156.

R. Smith, "An Overview of the Tesseract OCR Engine," in Proc. 9th Int. Conf. Document Anal. Recognit. (ICDAR), Vol. 2, Curitiba, Brazil, 2007, pp. 629–633, DOI: 10.1109/ICDAR.2007.4376991.

A. Upadhye, “A Comprehensive Survey of Text Data Cleaning Techniques: Challenges, Methods, and Best Practices,” Journal of Scientific and Engineering Research, Vol. 7, No. 8, pp. 205–210, 2020.

F. Josi, C. Wartena, and U. Heid, "Preparing Legal Documents for NLP Analysis: Improving the Classification of Text Elements by using Page Features," in Comput. SCI. Inf. Technol. (CS & IT), 2022, DOI: 10.5121/csit.2022.120102.

C. Merola and J. Singh, "Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation," in Knowledge-Enhanced Information Retrieval: 2nd Int. Workshop, KEIR 2025, Revised Selected Papers, Lucca, Italy, Apr. 2025. Berlin, Germany: Springer, 2025, pp. 3–18, DOI: 10.1007/978-3-032-02899-0_1.

P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, "SQuAD: 100,000+ Questions for Machine Comprehension of Text," in Proc. 2016 Conf. Empirical Methods Natural Lang. Process. (EMNLP), Austin, TX, USA, Nov. 2016, pp. 2383–2392, DOI: 10.18653/v1/D16-1264.

Y. Wan et al., "SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-Grained Evaluation," arXiv preprint arXiv:2405.09939, May 2024, DOI: 10.48550/arXiv.2405.09939.

S. Yuen, T. Su, Z. Wang, Y. Du, and A. Sobey, "Automatic Dataset Generation for Knowledge Intensive Question Answering Tasks," arXiv preprint arXiv:2505.14212, 2025, DOI: 10.48550/arXiv.2505.14212.

X. Yu, Z. Zhang, F. Niu, X. Hu, X. Xia, and J. Grundy, "What Makes a High-Quality Training Dataset For Large Language Models: A Practitioners' Perspective," in Proc. 39th IEEE/ACM Int. Conf. Automated Softw. Eng. (ASE '24), New York, NY, USA: ACM, 2024, pp. 656–668, DOI: 10.1145/3691620.3695061.

E. J. Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models," arXiv preprint arXiv:2106.09685, 2021, DOI: 10.48550/arXiv.2106.09685.

L. Wang et al., "Parameter-Efficient Fine-Tuning in Large Models: A Survey of Methodologies," arXiv preprint arXiv:2410.19878, Oct. 2024, DOI: 10.48550/arXiv.2410.19878.

Y. Chang et al., "A Survey on Evaluation of Large Language Models," ACM Trans. Intell. Syst. Technol., Vol. 15, No. 3, Mar. 2024, DOI: 10.1145/3641289.

C.-Y. Lin, "ROUGE: A Package for Automatic Evaluation of Summaries," in Text Summarization Branches Out, Barcelona, Spain, Jul. 2004, pp. 74–81. [Online]. Available: https://aclanthology.org/W04-1013/

S. Banerjee and A. Lavie, "METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments," in Proc. ACL Workshop Intrinsic Extrinsic Eval. Measures Mach. Transl. Summarization, Ann Arbor, MI, USA, Jun. 2005, pp. 65–72.

T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, "BERTScore: Evaluating Text Generation with BERT," in Proc. Int. Conf. Learn. Represent. (ICLR), 2020, DOI: 10.48550/arXiv.1904.09675.

S. Es, J. James, L. Espinosa-Anke, and S. Schockaert, "RAGAs: Automated Evaluation of Retrieval Augmented Generation," in Proc. 18th Conf. Eur. Chapter Assoc. Comput. Linguistics: Syst. Demonstrations, St. Julian's, Malta, Mar. 2024, pp. 150–158, DOI: 10.18653/v1/2024.eacl-demo.16.




DOI: https://doi.org/10.32520/stmsi.v15i9.6913

Article Metrics

Abstract view : 0 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.