MS MARCO Web Search: A Large-scale Information-rich Web Dataset with Millions of Real Click Labels
WWW 2024(2024)
摘要
Recent breakthroughs in large models have highlighted the critical
significance of data scale, labels and modals. In this paper, we introduce MS
MARCO Web Search, the first large-scale information-rich web dataset, featuring
millions of real clicked query-document labels. This dataset closely mimics
real-world web document and query distribution, provides rich information for
various kinds of downstream tasks and encourages research in various areas,
such as generic end-to-end neural indexer models, generic embedding models, and
next generation information access system with large language models. MS MARCO
Web Search offers a retrieval benchmark with three web retrieval challenge
tasks that demand innovations in both machine learning and information
retrieval system research domains. As the first dataset that meets large, real
and rich data requirements, MS MARCO Web Search paves the way for future
advancements in AI and system research. MS MARCO Web Search dataset is
available at: https://github.com/microsoft/MS-MARCO-Web-Search.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要