Article,

The Ladder: A Reliable Leaderboard for Machine Learning Competitions

A. Blum, and M. Hardt.
(2015)cite arxiv:1502.04585.

Abstract

The organizer of a machine learning competition faces the problem of maintaining an accurate leaderboard that faithfully represents the quality of the best submission of each competing team. What makes this estimation problem particularly challenging is its sequential and adaptive nature. As participants are allowed to repeatedly evaluate their submissions on the leaderboard, they may begin to overfit to the holdout data that supports the leaderboard. Few theoretical results give actionable advice on how to design a reliable leaderboard. Existing approaches therefore often resort to poorly understood heuristics such as limiting the bit precision of answers and the rate of re-submission. In this work, we introduce a notion of "leaderboard accuracy" tailored to the format of a competition. We introduce a natural algorithm called "the Ladder" and demonstrate that it simultaneously supports strong theoretical guarantees in a fully adaptive model of estimation, withstands practical adversarial attacks, and achieves high utility on real submission files from an actual competition hosted by Kaggle. Notably, we are able to sidestep a powerful recent hardness result for adaptive risk estimation that rules out algorithms such as ours under a seemingly very similar notion of accuracy. On a practical note, we provide a completely parameter-free variant of our algorithm that can be deployed in a real competition with no tuning required whatsoever.

BibTeX key: blum2015ladder
entry type: article
year: 2015
url: http://arxiv.org/abs/1502.04585
note: cite arxiv:1502.04585

Users

Comments and Reviewsshow / hide

Please log in to take part in the discussion (add own reviews or comments).

Cite this publication

@article{blum2015ladder, abstract = {The organizer of a machine learning competition faces the problem of maintaining an accurate leaderboard that faithfully represents the quality of the best submission of each competing team. What makes this estimation problem particularly challenging is its sequential and adaptive nature. As participants are allowed to repeatedly evaluate their submissions on the leaderboard, they may begin to overfit to the holdout data that supports the leaderboard. Few theoretical results give actionable advice on how to design a reliable leaderboard. Existing approaches therefore often resort to poorly understood heuristics such as limiting the bit precision of answers and the rate of re-submission. In this work, we introduce a notion of "leaderboard accuracy" tailored to the format of a competition. We introduce a natural algorithm called "the Ladder" and demonstrate that it simultaneously supports strong theoretical guarantees in a fully adaptive model of estimation, withstands practical adversarial attacks, and achieves high utility on real submission files from an actual competition hosted by Kaggle. Notably, we are able to sidestep a powerful recent hardness result for adaptive risk estimation that rules out algorithms such as ours under a seemingly very similar notion of accuracy. On a practical note, we provide a completely parameter-free variant of our algorithm that can be deployed in a real competition with no tuning required whatsoever.}, added-at = {2019-05-31T16:17:22.000+0200}, author = {Blum, Avrim and Hardt, Moritz}, biburl = {https://www.bibsonomy.org/bibtex/2273718926231ac4e12cec948f20b7c4c/kirk86}, description = {[1502.04585] The Ladder: A Reliable Leaderboard for Machine Learning Competitions}, interhash = {d644f18f21b8f1751eaa414d4d5cc594}, intrahash = {273718926231ac4e12cec948f20b7c4c}, keywords = {bounds deep-learning generalization machine-learning mathematics probability stable stats theory}, note = {cite arxiv:1502.04585}, timestamp = {2019-05-31T16:19:43.000+0200}, title = {The Ladder: A Reliable Leaderboard for Machine Learning Competitions}, url = {http://arxiv.org/abs/1502.04585}, year = 2015 }

BibSonomy

The Ladder: A Reliable Leaderboard for Machine Learning Competitions

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on