Howard & Ruder, 2018: Universal Language Model Fine-tuning for Text Classification
Universal Language Model Fine-tuning for Text Classification is a 2018 paper by Jeremy Howard and Sebastian Ruder (arXiv), which introduced ULMFiT, a method for Transfer Learning in NLP.
ULMFiT has three stages:
- Pre-train an LSTM Language Model (AWD-LSTM) on a large general corpus (Wikitext-103).
- Fine-Tuning the language model on text from the target task.
- Fine-tune a classifier on top of the language model.
It introduced techniques like discriminative fine-tuning, slanted triangular learning rates and gradual unfreezing, to avoid forgetting what was learned during pre-training.
It showed that transfer learning, which was already standard in computer vision, works well for NLP too. ULMFiT is covered in Deep Learning for Coders with Fastai and Pytorch: AI Applications Without a PhD.