SyncAI.news, a Varaisys broadcasting
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
JP

Jiwon Park, Seohyun Pyeon, Jinwoo Kim, Rina Carines Cabral, Zhenuyan He, Yihao Ding, Soyeon Caren Han

· 1 min read

ResearcharXiv cs.CL

DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections

arXiv:2508.15851v4 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information seeking: synthesizing multimodal evidence scattered across multiple documents and structural formats. Existing QAs remain narrow in scope, relying on unimodal text and short-span reasoning that fail to capture the complexity of real information-seeking. We introduce DocHop-QA, a benchmark of 11,379 instances for evaluating multimodal, multi-document, multi-hop scientific QA. Built from publicly available PubMed articles, DocHop-QA incorporates textual passages, tables, and layout cues, enabling cross-document inference without explicit hyperlinks. To scale realistic QA construction, we develop an LLM-driven generation pipeline grounded in 11 scientific reasoning concepts, producing diverse and coherent question-answer pairs. To highlight the utility and versatility of the dataset, we propose a task-driven evaluation framework spanning four settings, including generative answering, multimodal evidence integration and structured index prediction. Experiments show that current models struggle with DocHop-QA's long-context, multi-evidence demands, establishing it as a rigorous testbed for advancing next-generation scientific QA systems.

Original source

This story was published by arXiv cs.CL and written by Jiwon Park, Seohyun Pyeon, Jinwoo Kim, Rina Carines Cabral, Zhenuyan He, Yihao Ding, Soyeon Caren Han. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News