A surge of specialist models has been targeting Arithmark 2 on the Open SLM Leaderboard, raising concerns about the fairness of the ranking system. Banaxi-Tech and AxiomicLabs have taken steps to address this issue, as some models have gained high scores without demonstrating broader capabilities. The leaderboard initially ranked models by average score, which allowed specialist models with high Arithmark 2 scores to rise in the rankings despite limited versatility. The issue became apparent when models like Atom 2.7M, with 2.7M parameters, achieved a 69.40% Arithmark 2 score and ranked #6 on the leaderboard, outperforming models with 50 times more parameters. This highlighted a gap in the leaderboard's methodology, prompting action from the community to correct the flaw.

In response, AxiomicLabs decided to rank all specialist models at the bottom, aiming to break their dominance. However, this only temporarily resolved the issue, as Ideoa Labs later exploited a weakness in the specialist algorithm by increasing parameter counts while still training on synthetic arithmetic, allowing models like Nexus-Erebus-50M to achieve top rankings. Currently, Nexus-Erebus-135M holds the #1 position on the leaderboard. Banaxi-Tech is now working to address these issues through proposed changes to the specialist classifier system.

Banaxi-Tech acknowledges that the models are not inherently flawed but argues that they should not be compared to general-purpose models. The company has submitted a pull request (PR 56) to revise the specialist classification system, which would classify models based on whether their Arithmark 2 score is more than 30% higher than Hellaswag. This change aims to create a more accurate ranking system while also including an exclusion list for models that meet this criterion.

Source: huggingface