PRISM-LoRA: Principal-Direction-Guided Single-LoRA Merging Framework for Continual Learning of Large Language Models
Taehyeong Kwon, Okran Jeong
Source abstract
Continual learning of large language models requires incorporating new knowledge while limiting catastrophic forgetting; however, many low-rank adaptation (LoRA)-based methods retain task-specific modules or separate learning spaces that grow with the task sequence. We propose PRISM-LoRA, a replay-free framework that repeatedly trains a single LoRA module, merges its update into the backbone, and reinitializes it for the next task. Unlike LoRA-based continual-learning approaches that retain task-specific modules or allocate separate task-specific learning spaces, PRISM-LoRA uses the directional structure of the accumulated backbone weight change to guide both subsequent training and consolidation while reusing a single LoRA module. Before each task, PRISM-LoRA decomposes the accumulated backbone weight change using singular value decomposition and uses the squared singular values as directional spectral energies. These energies are treated as empirical weight-space indicators of directions associated with retained performance, rather than as direct evidence of stored knowledge. Principal-direction regularization penalizes overlap with high-energy principal directions during training, whereas dynamic merge applies direction-wise scaling to the learned weight change before merging it into the backbone. On T5-Large, PRISM-LoRA achieved Avg OP scores of 79.4±0.2% on Standard CL and 73.3±0.3% on the 15-task Long Sequence benchmark across three random seeds. The Standard CL result was comparable to CLoRA, whereas the Long Sequence result was 1.6%p higher than CSF. On LLaMA2-7B, PRISM-LoRA achieved Avg OP scores of 80.4 ± 0.1% and 75.9 ± 0.1% on the Standard CL and Long Sequence benchmarks, respectively. Ablation and performance-recovery analyses further showed that both directional protection stages contributed to retention and that a small number of high-energy directions were associated with a substantial portion of the retained performance.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.