MessIRve: A Large-Scale Spanish Information Retrieval Dataset

Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, current IR benchmarks lack Spanish data, hindering the development of information access tools for Spanish speakers. We introduce MessIRve, a large-scale Spanish IR dataset with around 730 thousand queries from Google’s autocomplete API and relevant documents sourced from Wikipedia. MessIRve’s queries reflect diverse Spanishspeaking regions, unlike other datasets that are translated from English or do not consider dialectal variations. The large size of the dataset allows it to cover a wide variety of topics, unlike smaller datasets. We provide a comprehensive description of the dataset, comparisons with existing datasets, and baseline evaluations of prominent IR models. Our contributions aim to advance Spanish IR research and improve information access for Spanish speakers.

Palabras clave: INFORMATION RETRIEVAL , RESOURCES AND EVALUATION , NATURAL LANGUAGE PROCESSING , NLP DATASETS

Ver el registro completo

Archivos asociados

Tamaño: 283.9Kb

Formato: PDF

Descargar

Licencia

Excepto donde se diga explícitamente, este item se publica bajo la siguiente descripción: Creative Commons Attribution 2.5 Unported (CC BY 2.5)

Identificadores

URI: http://hdl.handle.net/11336/258637

URL: https://arxiv.org/abs/2409.05994

DOI: https://doi.org/10.48550/arXiv.2409.05994

Colecciones

Articulos(ICC)
Articulos de INSTITUTO DE INVESTIGACION EN CIENCIAS DE LA COMPUTACION

Citación

Valentini, Francisco Tomás; Cotik, Viviana Erica; Furman, Damián Ariel; Bercovich, Ivan; Altszyler Lemcovich, Edgar Jaim; et al.; MessIRve: A Large-Scale Spanish Information Retrieval Dataset; Cornell University; arXiv; 9-2024; 1-13

Altmétricas