GLUE
GLUE (General Language Understanding Evaluation) is a benchmark for evaluating natural language understanding models across a collection of different tasks, rather than just one.
Dataset introduced in GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.
It bundles nine English sentence and sentence-pair tasks, covering things like sentiment analysis, paraphrase detection, sentence similarity and natural language inference, and reports a single average score across them. It also includes a hand-crafted diagnostic dataset for analysing which linguistic phenomena a model handles well or poorly. It was a standard way to compare pre-trained language models like GPT and BERT when they first came out.