Do text embeddings perfectly encode text?

A paper titled Text Embeddings Reveal As Much as Text (EMNLP 2023) addresses whether input text can be recovered from the output embeddings produced by neural embedding models. The authors, including Jack Morris, a PhD student at Cornell Tech, developed a method called vec2text that inverts text embeddings back into text.
The work is set against the rise of Retrieval Augmented Generation (RAG) systems and vector databases, which store embedding vectors rather than the underlying text. The paper considers scenarios in which a database of embeddings is compromised or sold, asking whether the original text could be reconstructed from those vectors. It notes that embeddings are typically sequences of numbers with no constraints beyond a requirement of semantic similarity, making individual values unreadable.
Text embeddings are the output of neural networks. The paper cites the data processing inequality, stating that functions cannot add information to their input, and notes that nonlinear layers such as ReLU destroy some information, making perfect retention impossible. Similar inversion work in computer vision, including Dosovitskiy (2016) and later work on ImageNet classifier outputs, is referenced as motivation.
For a toy setting, the authors restricted inputs to 32 tokens (about 25 words) embedded into 768 floating-point vectors, roughly 3 kilobytes at 32-bit precision. Their first approach, a transformer trained on embedding-text pairs, reached a BLEU score of around 30 out of 100 and near-zero exact match. Re-embedding generated hypotheses showed cosine similarity around 0.97 to the ground-truth embedding. This led to an iterative corrector model, vec2text, which trains on a ground-truth embedding, a hypothesis text, and its embedding to predict closer text.
A single correction pass raised BLEU from 30 to 50. With 50 steps and additional techniques, the method recovered 92% of 32-token sequences exactly and reached a BLEU score of 97. The authors leave to future work the relationship between text length, embedding size, and invertibility, defenses against inversion, and applications to other modalities.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
'Vec2text' can serve as a solution for accurately reverting embeddings back into text, thus highlighting the urgent need for revisiting security protocols around embedded data.