The Combination of RAG and RL: A Technological Breakthrough and Future Directions in Mathematical Reasoning with LLMs
Yuchen Wang
Source record
Source: Crossref
Published: Sep 13, 2026
DOI: 10.70267/aitia.20261925
Open original source ↗Source abstract
Large language models (LLMs) face challenges in mathematical reasoning tasks, including logical disjunction, process hallucination, and a loss of creativity, which severely restricts the reliability and application of LLMs for this purpose. To tackle these issues, this paper applies secondary research, systematically explores the mathematical logic of fusing retrieval-augmented generation and reinforcement learning, and constructs a three-dimensional fusion framework of ‘fusion object –optimization granulari ty–application scenario’. We identify three core technical paths: RL-optimized RAG retrieval strategies, RL-enhanced RAG reasoning chains, and RL-adapted RAG knowledge fusion. According to the results, the precision, sturdiness, and comprehension of mathem atical reasoning can be improved remarkably with the help of this ‘Knowledge Supply-Strategy Optimization’ closed-loop framework. It gets 10% -20% accuracy improvement on the benchmark datasets like MATH and AIME that also controls process hallucination bel ow 5%. This paper further analyzes the limitation of current technologies in terms of cross-domain theorem association, multimodal adaptations, and low-resource scenarios. The article also proposes future research ideas, including agent-based fusion framework and lightweight process supervision, which can efficiently and reliably support mathematics education, scientific research assistance, and engineering computation.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.