Abstract
Reproducibility must validate architectural robustness, not just numerical accuracy. We evaluate ColBERT-v2 and ConstBERT across five dimensions, finding that while ConstBERT reproduces within 0.05% MRR@10 on MS-MARCO, both models show a drop of 86-97% on long, narrative queries (TREC ToT 2025). Ablations prove this failure is architectural: performance plateaus at 20 words because the MaxSim operator's uniform token weighting cannot distinguish signal from filler noise. Furthermore, undocumented backend parameters create an 8-point gap due to ConstBERT's sparse centroid coverage, and fine-tuning with 3x more data actually degrades performance by up to 29%. We conclude that architectural constraints in multi-vector retrieval cannot be overcome by adaptation alone. Code: https://github.com/utshabkg/multi-vector-reproducibility.
Recommended Citation
U. K. Ghosh et al., "Reproduction Beyond Benchmarks: ConstBERT And ColBERT-v2 Across Backends And Query Distributions," SIGIR 2026 Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2921 - 2930, Association for Computing Machinery, Jul 2026.
The definitive version is available at https://doi.org/10.1145/3805712.3808561
Department(s)
Computer Science
Publication Status
Open Access
Keywords and Phrases
colbert; constbert; late interaction; multi-vector retrieval; out-of-domain generalization; plaid; trec tip-of-the-tongue
Document Type
Article - Conference proceedings
Document Version
Citation
File Type
text
Language(s)
English
Rights
© 2026 The Author(s), All rights reserved.
Creative Commons Licensing

This work is licensed under a Creative Commons Attribution 4.0 License.
Publication Date
19 Jul 2026
