Home /permanent

BERT

BERT (Bidirectional Encoder Representations from Transformers) is an encoder-only Transformer language model released by Google in 2018. It was introduced in the paper BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

BERT is pre-trained on unlabelled text with Masked Language Modelling, where some tokens are hidden and the model predicts them from the context on both sides, plus a next sentence prediction task. The pre-trained model can then be fine-tuned for downstream tasks like classification or question answering.

It inspired many follow-ups, including RoBERTa, and the masked prediction idea was carried over to speech in models like HuBERT and w2v-BERT.