A user reported being banned from HuggingFace after questioning the benchmark scores of two Qwen models, revealing a pattern of suppression. The user, DedeProGames, posted a discussion on the Qwen3.6-27B-Fable-Fusion-711 model page asking for modern benchmarks like SWE-bench, Terminal-Bench, GPQA, and LiveCodeBench. DavidAU, the model creator, deleted the post and then blocked the user from interacting with any of his repositories. This action has raised concerns about the transparency and openness of the HuggingFace platform.

The user detailed how DavidAU had previously deleted discussions on the 9B model after a user reported the model's failure in a basic coding task. Similarly, the creator ignored a direct request for SWE-bench scores on the 27B model and instead pinned positive discussions while deleting critical ones. DavidAU also responded to criticism with a dismissive 'Trust me bro, it's the best' when asked for real benchmark results. This pattern of closing, ignoring, deleting, and banning users has sparked debate about the integrity of the model evaluation process.

The source text outlines the controversy surrounding the 9B and 27B models, which use the same seven benchmarks. These benchmarks, all from 2018-2019, are considered outdated and saturated, failing to differentiate between decent models and genuinely intelligent ones. The user provided detailed benchmark scores, showing that the 27B model's performance on ARC-C, a key benchmark, was only slightly better than the base model, raising questions about the validity of the '700 club' claim. The user also highlighted how the Qwen 3.5 9B model's official benchmarks, including SWE-bench and GPQA, were copied into DavidAU's model card without being tested on his fine-tune.

Source: huggingface