MoverScore
MoverScore scores a candidate against a reference over contextual embeddings, but instead of BERTScore's greedy one-to-one token matching it finds the minimal Earth Mover's Distance (Word Mover's Distance) needed to transform one text's embeddings into the other's (Zhao et al., 2019). This soft, many-to-one alignment lets partial and distributed matches count, which tends to track human judgment better than hard matching, though it is still a single opaque similarity number that says nothing about factuality.
References
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. 2019. doi:10.48550/ARXIV.1909.02622. ↩