Language Models for Text Classification: From Bag-of-Words to Jev
Ahead of AI has published a technical article titled "Language Models for Text Classification: From Bag-of-Words to Jev," examining the history of language models for decision-making to contextualize the recently released Jev AI model, which the article says has been a cultural phenomenon in technical communities over the past two weeks.
The article's author states they are not affiliated with Jev, were not offered free access to it, and that the piece is not a product endorsement. The author describes Jev as aiming to classify things, and writes that their view evolved from "classifiers used to be my bread & butter; I can easily build this myself" to "wow, this actually works better than I thought."
According to the article, state-of-the-art GPT and open-weight LLMs can perform the same kinds of classification tasks as Jev while also handling more general decision-making, but Jev's stated advantage is handling those classification tasks faster and more cheaply. For a narrow, well-defined problem, the article says, Jev probably won't classify anything better, faster, or cheaper than a special-purpose classifier, but its selling point is being far more general than task-specific models. The author summarizes this as "Jev is essentially a text classifier," but also "not 'just' a text classifier."
The article begins with a chronological tour of applied text classification: bag-of-words representations used with naive Bayes, logistic regression, SVMs, Random Forest and XGBoost; then deep neural networks including CNNs and RNNs; then transformer-based encoder, decoder and encoder-decoder architectures. It cites reported accuracies on the IMDb movie review dataset: about 89.9% for bag-of-words with logistic regression, 85.66% for an LSTM RNN trained from scratch, 95.4% test accuracy for ULMFiT, about 90.07% for a text CNN in the author's experiments, approximately 95% for ModernBERT with little fine-tuning, and approximately 92% for a GPT-2 124M model. The article also notes BERT's 2018 classification token, T5's 2019 text-to-text approach, and refers to a DeepSeek V4.1 Flash encoder-decoder variant, along with the 2017 "Attention Is All You Need" paper.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency