MetaPCR-LLM: Universal Proof-Carrying Reasoning Across Heterogeneous Logics for Large Language Model Validation
Boris Galitsky, Vladimir Solodkin, Aleksandr Beznosikov
Source record
Source: Crossref
Published: Sep 10, 2026
DOI: 10.20944/preprints202609.0849.v1
Open original source ↗Source abstract
Large language models (LLMs) increasingly produce reasoning that combines deductive, abductive, probabilistic, defeasible, temporal and constraint-based inference, yet most symbolic validation approaches assume a single target logic or translate heterogeneous reasoning into one common formalism. We introduce MetaPCR-LLM, a proof-carrying framework for validating LLM reasoning across heterogeneous native logics. MetaPCR-LLM represents a reasoning trace as a typed proof graph in which each step is assigned to an appropriate native validator and carries a proof certificate or machine-checkable witness together with provenance, uncertainty, temporal, and defeasibility metadata. Transitions between logical systems are treated as explicit proof-carrying bridges, allowing the framework to distinguish locally valid inference from globally invalid cross-logic composition. Counter-abduction further tests accepted explanations against independently validated alternatives, while the same mechanism supports certified recomposition of reasoning fragments produced by multiple LLMs. We evaluate the framework on Truthful-Halluc, Med-Halluc, eSNLI-Halluc and Autoimmune-narrate-halluc across single- and mixed-logic reasoning regimes. MetaPCR-LLM achieves 0.86 overall accuracy, 0.88 step-validity accuracy and 0.86 chain-certification accuracy, with a cross-dataset macro-F1 of 0.838. Its pooled hallucination F1 reaches 0.83, compared with 0.77 for the ValidLLP-style baseline, an absolute improvement of 0.06 at the reported precision. The advantage remains especially pronounced on heterogeneous reasoning: mixed-logic F1 reaches 0.80, compared with 0.74 for ValidLLP-style validation and 0.70 for MetaPCR without bridge checking. Bridge validation reaches 0.88 accuracy and explicit proof diagnostics improve failure-type and validator localization. In multi-LLM recomposition, the complete certificate--bridge--label configuration reaches 0.86 accuracy and reduces invalid compositions from 0.29 to 0.18. Human evaluation further shows an increase in reasoning-assessment accuracy from 0.68 to 0.88. These results show that proof-carrying heterogeneous validation improves the aggregate answer-level comparison while also providing reliable composition, localization and auditing of reasoning that spans multiple formal systems.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.