CAA-CLIP: Decoupled Pair and Cultural-Term Adaptation for Chinese Ancient Architecture Image–Text Retrieval
Xiaoyu Zhang, Hao Wang, Wangyu Wu, Qingqing Hu
Source abstract
Image–text retrieval for Chinese ancient architecture requires fine-grained discrimination among visually similar sites and alignment with specialist terms that are sparse in general vision–language corpora. We propose CAA-CLIP, a parameter-efficient framework for exact image–annotation retrieval and annotation-derived cultural-term retrieval. CAA-CLIP adapts five frozen CLIP-family encoders through separate pair and term branches and fuses row-normalized scores with validation-selected weights. Evaluation uses 581 image–text pairs from 58 sites, five site-grouped outer folds, and three optimization seeds per fold. After seed averaging, CAA-CLIP obtains 29.06±3.05% Avg R@1, 63.19±6.34% Avg R@5, 78.65±6.64% Avg R@10, and 43.82±6.50% term mAP. Against matched validation-selected plain fusion, the pair-retrieval differences are 0.43–0.98 points, and the term-mAP difference is 9.05 points. No primary comparison reaches significance after Holm correction over the five folds, although the term difference is positive in every fold. Equal CAA fusion attains 44.39% term mAP, indicating that validation-selected term weights do not improve this endpoint. These findings distinguish expert fusion for exact-pair retrieval from direct term supervision for annotation-defined term retrieval in small, site-structured heritage collections.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.