Language:

When Does Confidence-Based Cascade Deferral Suffice?

arXiv.org, 2024-01

2024. This work is published under http://creativecommons.org/licenses/by/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License. ;http://creativecommons.org/licenses/by/4.0 ;EISSN: 2331-8422 ;DOI: 10.48550/arxiv.2307.02764

Full text available

Citations Cited by

Actions
1. Add to My Research
2. Remove from My Research
3. E-mail
4. Print
5. Permalink
6. Citation
7. EasyBib
8. EndNote
9. RefWorks
10. Delicious
11. Export RIS
12. Export BibTeX

Title:
When Does Confidence-Based Cascade Deferral Suffice?
Author: Jitkrittum, Wittawat ; Gupta, Neha ; Menon, Aditya Krishna ; Narasimhan, Harikrishna ; Ankit Singh Rawat ; Kumar, Sanjiv
Subjects: Classifiers ; Computer Science - Learning ; Statistics - Machine Learning
Is Part Of: arXiv.org, 2024-01
Description: Cascades are a classical strategy to enable inference cost to vary adaptively across samples, wherein a sequence of classifiers are invoked in turn. A deferral rule determines whether to invoke the next classifier in the sequence, or to terminate prediction. One simple deferral rule employs the confidence of the current classifier, e.g., based on the maximum predicted softmax probability. Despite being oblivious to the structure of the cascade -- e.g., not modelling the errors of downstream models -- such confidence-based deferral often works remarkably well in practice. In this paper, we seek to better understand the conditions under which confidence-based deferral may fail, and when alternate deferral strategies can perform better. We first present a theoretical characterisation of the optimal deferral rule, which precisely characterises settings under which confidence-based deferral may suffer. We then study post-hoc deferral mechanisms, and demonstrate they can significantly improve upon confidence-based deferral in settings where (i) downstream models are specialists that only work well on a subset of inputs, (ii) samples are subject to label noise, and (iii) there is distribution shift between the train and test set.
Publisher: Ithaca: Cornell University Library, arXiv.org
Language: English
Identifier: EISSN: 2331-8422
DOI: 10.48550/arxiv.2307.02764
Source: Freely Accessible Journals
arXiv.org
ROAD: Directory of Open Access Scholarly Resources
ProQuest Central

Back to results list


INSPIRE LIBRARY - TON DUC THANG UNIVERSITY	(84-028) 37 755 057	Feedback
19 Nguyen Huu Tho St. Dist.7, HCM	thuvien@tdtu.edu.vn	Feedback

When Does Confidence-Based Cascade Deferral Suffice?

Searching Remote Databases, Please Wait