GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding is a 2018 paper by Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy and Samuel R. Bowman that introduces the GLUE benchmark.
The idea is that a useful natural language understanding model should handle many different tasks, not just one. So GLUE collects nine existing English sentence and sentence-pair tasks (like sentiment, paraphrase and natural language inference) into a single benchmark with one average score and a public leaderboard. It also adds a hand-crafted diagnostic dataset to analyse which linguistic phenomena models get right or wrong.