An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking

Koustav Rudra; Zeon Trevor Fernando; Avishek Anand

Details

Original language	English
Publication status	E-pub ahead of print - 30 Mar 2021

Abstract

Recently introduced pre-trained contextualized autoregressive models like BERT have shown improvements in document retrieval tasks. One of the major limitations of the current approaches can be attributed to the manner they deal with variable-size document lengths using a fixed input BERT model. Common approaches either truncate or split longer documents into small sentences/passages and subsequently label them - using the original document label or from another externally trained model. In this paper, we conduct a detailed study of the design decisions about splitting and label transfer on retrieval effectiveness and efficiency. We find that direct transfer of relevance labels from documents to passages introduces label noise that strongly affects retrieval effectiveness for large training datasets. We also find that query processing times are adversely affected by fine-grained splitting schemes. As a remedy, we propose a careful passage level labelling scheme using weak supervision that delivers improved performance (3-14% in terms of nDCG score) over most of the recently proposed models for ad-hoc retrieval while maintaining manageable computational complexity on four diverse document retrieval datasets.

Keywords

cs.IR, H.3.3

Cite this

An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking. / Rudra, Koustav; Fernando, Zeon Trevor; Anand, Avishek.
2021.

Research output: Working paper/Preprint › Preprint

Rudra, K, Fernando, ZT & Anand, A 2021 'An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking'. <https://arxiv.org/abs/2103.16669>

Rudra, K., Fernando, Z. T., & Anand, A. (2021). An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking. Advance online publication. https://arxiv.org/abs/2103.16669

Rudra K, Fernando ZT, Anand A. An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking. 2021 Mar 30. Epub 2021 Mar 30.

Rudra, Koustav ; Fernando, Zeon Trevor ; Anand, Avishek. / An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking. 2021.

Download

@techreport{eba3f829cef144ba9a2123b1334bdf5d,

title = "An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking",

abstract = " Recently introduced pre-trained contextualized autoregressive models like BERT have shown improvements in document retrieval tasks. One of the major limitations of the current approaches can be attributed to the manner they deal with variable-size document lengths using a fixed input BERT model. Common approaches either truncate or split longer documents into small sentences/passages and subsequently label them - using the original document label or from another externally trained model. In this paper, we conduct a detailed study of the design decisions about splitting and label transfer on retrieval effectiveness and efficiency. We find that direct transfer of relevance labels from documents to passages introduces label noise that strongly affects retrieval effectiveness for large training datasets. We also find that query processing times are adversely affected by fine-grained splitting schemes. As a remedy, we propose a careful passage level labelling scheme using weak supervision that delivers improved performance (3-14% in terms of nDCG score) over most of the recently proposed models for ad-hoc retrieval while maintaining manageable computational complexity on four diverse document retrieval datasets. ",

keywords = "cs.IR, H.3.3",

author = "Koustav Rudra and Fernando, {Zeon Trevor} and Avishek Anand",

note = "Paper is about the performance analysis of contextual ranking strategies in an ad-hoc document retrieval",

year = "2021",

month = mar,

day = "30",

language = "English",

type = "WorkingPaper",

}

Download

TY - UNPB

T1 - An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking

AU - Rudra, Koustav

AU - Fernando, Zeon Trevor

AU - Anand, Avishek

N1 - Paper is about the performance analysis of contextual ranking strategies in an ad-hoc document retrieval

PY - 2021/3/30

Y1 - 2021/3/30

N2 - Recently introduced pre-trained contextualized autoregressive models like BERT have shown improvements in document retrieval tasks. One of the major limitations of the current approaches can be attributed to the manner they deal with variable-size document lengths using a fixed input BERT model. Common approaches either truncate or split longer documents into small sentences/passages and subsequently label them - using the original document label or from another externally trained model. In this paper, we conduct a detailed study of the design decisions about splitting and label transfer on retrieval effectiveness and efficiency. We find that direct transfer of relevance labels from documents to passages introduces label noise that strongly affects retrieval effectiveness for large training datasets. We also find that query processing times are adversely affected by fine-grained splitting schemes. As a remedy, we propose a careful passage level labelling scheme using weak supervision that delivers improved performance (3-14% in terms of nDCG score) over most of the recently proposed models for ad-hoc retrieval while maintaining manageable computational complexity on four diverse document retrieval datasets.

AB - Recently introduced pre-trained contextualized autoregressive models like BERT have shown improvements in document retrieval tasks. One of the major limitations of the current approaches can be attributed to the manner they deal with variable-size document lengths using a fixed input BERT model. Common approaches either truncate or split longer documents into small sentences/passages and subsequently label them - using the original document label or from another externally trained model. In this paper, we conduct a detailed study of the design decisions about splitting and label transfer on retrieval effectiveness and efficiency. We find that direct transfer of relevance labels from documents to passages introduces label noise that strongly affects retrieval effectiveness for large training datasets. We also find that query processing times are adversely affected by fine-grained splitting schemes. As a remedy, we propose a careful passage level labelling scheme using weak supervision that delivers improved performance (3-14% in terms of nDCG score) over most of the recently proposed models for ad-hoc retrieval while maintaining manageable computational complexity on four diverse document retrieval datasets.

KW - cs.IR

KW - H.3.3

M3 - Preprint

BT - An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking

ER -

Research@Leibniz University

An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking

Authors

Research Organisations

External Research Organisations

Details

Abstract

Keywords

Cite this