Does Linguistic Register Affect Faithfulness in Spanish Cache-Augmented Generation? An Exploratory Study with Open-Source LLMs
Keywords:
- Cache-augmented generation,
- CAG,
- Large language models,
- Linguistic register,
- Faithfulness,
- Spanish,
- Linguistic robustness,
- RAGAS
Abstract
Cache-Augmented Generation (CAG) provides an alternative to retrieval-based knowledge augmentation by preloading external knowledge into a reusable key-value (KV) cache. However, the extent to which CAG remains stable when semantically equivalent queries are expressed through different linguistic registers remains insufficiently understood, particularly in Spanish. This exploratory study investigates whether formal-colloquial register variation affects the Faithfulness of CAG responses. A controlled paired design was constructed from 10 Spanish concepts and two complementary tasks definition generation and concept identification yielding 20 semantic cases and 40 formal-colloquial queries per model. Three open-source instruction-tuned LLMs (Qwen2.5-1.5B-Instruct, SmolLM2-1.7B-Instruct and Ministral-3B-Instruct) were evaluated under deterministic generation conditions using a fixed preloaded knowledge context. Faithfulness was assessed with RAGAS using Granite 3.2 as the LLM judge. Of 120 planned generations, 113 produced valid Faithfulness evaluations, yielding 53 complete formal-colloquial pairs. Neither paired parametric nor non-parametric analyses identified statistically significant register effects for any individual model and the between-model analysis did not detect systematic differences in register sensitivity. Descriptively, however, SmolLM2 tended to favor formal formulations, Ministral showed the opposite tendency and Qwen2.5 exhibited the most stable profile, combining comparatively high Faithfulness, limited formal-colloquial separation and relatively low within-condition variability. These findings provide initial evidence that formal-colloquial variation does not produce a detectable systematic effect on CAG Faithfulness under the conditions examined, while suggesting that descriptive sensitivity to linguistic register may depend on the underlying model and task configuration.
Downloads
References
Oche AJ, Folashade AG, Ghosal T, Biswas A (2025) A systematic review of key Retrieval-Augmented Generation (RAG) systems: Progress, gaps and future directions. arXiv. https://arxiv.org/abs/2507.18910
Carrasco-Sáez JL, Contreras-Saavedra C, San-Martín-Quiroga S, Contreras-Saavedraand CE, Viveros-Muñoz R (2025) Analyzing higher education students’ prompting techniques and their impact on ChatGPT’s performance: An exploratory study in Spanish. Appl Sci 15: 7651. https://www.mdpi.com/2076-3417/15/14/7651
Viveros-Muñoz R, Carrasco-Sáez J, Contreras-Saavedra C, San-Martín-Quiroga S, Contreras-Saavedra CE (2025) Does the grammatical structure of prompts influence the responses of generative Artificial Intelligence? An exploratory analysis in Spanish. Appl Sci 15: 3882. https://www.mdpi.com/2076-3417/15/7/3882
BJ Chan, CT Chen, JH Cheng, HH Huang (2025) Don’t do RAG: When cache-augmented generation is all you need for knowledge tasks. ACM Digital Library 893-897. https://dl.acm.org/doi/abs/10.1145/3701716.3715490
Lu S, Wang H, Rong Y, Chen Z, Tang Y (2025) TurboRAG: Accelerating retrieval-augmented generation with precomputed KV caches for chunked text. Assoc Comput Linguistics 6588-6601. https://aclanthology.org/2025.emnlp-main.334/
Agarwal S, Sundaresan S, Mitra S, Mahapatra D, Gupta A, et al. (2025) Cache-craft: Managing chunk-caches for efficient retrieval-augmented generation. ACM Digital Library 3: 1-28. https://dl.acm.org/doi/abs/10.1145/3725273
Corallo G, Weller O, Petroni F, Papotti P (2026) CacheNotes: Task-aware key-value cache compression for reasoning-intensive knowledge tasks Assoc Comput Linguistics 1: 6571-6590. https://aclanthology.org/2026.eacl-long.309/
Rashkin H, Reitter D, Tomar GS, Das D (2021) Increasing faithfulness in knowledge-grounded dialogue with controllable features. Assoc Comput Linguistics 1: 704-718. https://aclanthology.org/2021.acl-long.58/
Es S, James J, Anke LE, Schockaert S (2024) Ragas: Automated evaluation of retrieval augmented generation. Assoc Comput Linguistics 150-158. https://aclanthology.org/2024.eacl-demo.16/
Wallat J, Heuss M, Rijke M, Anand A (2025) Correctness is not faithfulness in retrieval augmented generation attributions. ACM Digital Library 22-32. https://dl.acm.org/doi/abs/10.1145/3731120.3744592
SG Kwak, JH Kim (2017) Central limit theorem: The cornerstone of modern statistics. Korean J Anesthesiol 70: 144-156. https://ekja.org/journal/view.php?doi=10.4097/kjae.2017.70.2.144

