Research on Deep Learning-Based Text and Symbol Recognition
Li Zeng, Tongguang Ni, Cheng Tsuang Koh
Source abstract
Handwritten mathematical expression recognition converts two-dimensional formula images into structured LaTeX sequences. However, existing vision-to-sequence methods struggle to explicitly model complex spatial structures and symbol-level ambiguities. To address this issue, the authors propose the Structure-Role Refined Network (SR2-Net). The model incorporates explicit structural modeling in the encoder. The Explicit Structural Parsing Module (ESP) aggregates spatial features into symbol-level representations and captures geometric relationships among symbols, producing structure-enhanced features. The Semantic–Role Decoupling Module (SRD) further separates visual semantics from structural roles and integrates them through structure-aware fusion, improving consistency under complex layouts. Finally, a structure-aware decoder generates LaTeX sequences by jointly exploiting structural context and semantic constraints. Experiments show that SR2-Net improves recognition accuracy and structural consistency, especially for formulas with complex nesting and superscript–subscript structures.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.