Home /permanent

BERTScore

BERTScore scores a candidate against a reference by matching their contextual BERT token embeddings, greedily pairing each token with its most similar counterpart and aggregating the cosine similarities into precision, recall, and F1 (Zhang et al., 2019). Unlike ROUGE: A Package for Automatic Evaluation of Summaries's exact word overlap, this captures semantic and paraphrase similarity, so "a quick fox" and "a fast fox" score highly, but the result is a single opaque similarity number that says nothing about factuality or which content is missing.

References

Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. BERTScore: Evaluating Text Generation with BERT. 2019. doi:10.48550/ARXIV.1904.09675.