HuggingFace released JEV-27B-VL on September 30, 2026, saying it makes zero-shot image recommendations with 59% top-5 hit rate, outperforming collaborative filtering on 200 users' data. It is the company's first vision update since the release of JEV-27B, a text-only decision model.
HuggingFace reported an AUC of 0.727 for JEV-27B-VL, measured on MicroLens, a public dataset of 200 users. That compares with 0.728 for collaborative filtering, which uses 59,045 users' watch histories.
JEV-27B-VL is built on the JEV-27B architecture and targets video recommendation and image classification. Availability begins immediately, initially for developers and researchers.
"Show it images and text, ask a question, and it answers in one forward pass with a calibrated probability for every option," said Hai Yu, cloudyu. The model can look at images while thinking through questions step by step.
The announcement follows the release of JEV-27B, a text-only decision model. HuggingFace said the new model extends the capabilities of its previous release, enabling image-based decision-making without training data.
HuggingFace did not say how the model handles complex visual tasks, and raised the open question of how it scales to larger datasets. The company said it will continue to refine the model's capabilities and expand its use cases.
Source: huggingface