Hallucination Mitigation in Large Language Model-Based Tool Recommendation: A Cross-Provider Architectural Ablation Study Across Two Model Generations
Lavdim Menxhiqi, Galia Marinova
Source record
Source: Crossref
Published: Jun 1, 2026
DOI: 10.20944/preprints202606.0008.v1
Open original source ↗Source abstract
In a closed-inventory Large Language Model (LLM) system such as Online-CADCOM, which recommends engineering tools from a verified inventory, mention-level hallucination occurs when the model recommends tools not present in the inventory. We evaluate a three-mechanism mitigation stack consisting of database-grounded context injection, fixed vocabulary constraints, and enforced JavaScript Object Notation (JSON) output across three commercial LLM providers (OpenAI, Anthropic, Google), two model generations, and two output modes (standard and reasoning), totaling 6,912 Application Programming Interface (API) calls over 12 configurations. The hallucination rate (HR) decreases from 59–74% to 3.3–14.9% under the full architecture, with cross-provider averages stable across generations (6.8% Generation 1 (Gen1), 7.5% Generation 2 (Gen2)). A key finding is the C3 anomaly, observed under the configuration where only JSON output enforcement is active without any grounding mechanisms: JSON enforcement alone increases hallucination above the unconstrained baseline for all providers (+10.1 percentage points (pp) Gen1, +15.1 pp Gen2). We hypothesize that mandatory entity fields in the JSON schema exert slot-filling pressure, forcing models to populate tool-name slots from training-data priors when no grounding vocabulary is provided. This hypothesis requires further experimental validation. Reasoning-mode models provide no statistically significant improvement under architectural constraints. A frequency-weighted audit shows that the majority of remaining out-of-inventory mentions correspond to real engineering tools absent from the platform's inventory. Under the full architecture, 35–63% of responses still contain at least one such mention, indicating that handling unseen tools remains an open challenge for closed-inventory recommendation systems.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.