Base Labs, the research arm of Baseten, announced a new safety infrastructure standard alongside its partnership with Hugging Face and Goodfire AI on Wednesday. The initiative aims to build safety evaluation and monitoring infrastructure for open-weight models, which are increasingly vulnerable to a technique called abliteration that can remove their safeguards.

The partnership is framed as a 'standard' for open models, emphasizing transparency and embedding safety measures into training and deployment processes rather than adding them later. Base Labs stated, 'We believe openness to be an advantage for AI safety.

Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source.'

Goodfire AI, which specializes in model interpretability, is expected to play a key role in integrating safety into open models.

The company, which raised $150 million in a Series B led by B Capital earlier this year, emphasized that 'Safety must be built into open models and provided by those who serve them.' Baseten, which raised $1.5 billion in a Series F in June, is also seeking contributions from the broader developer ecosystem to build a safer framework for open models.

The announcement comes amid growing concerns over the safety of open-weight models, with Hugging Face hosting over 6,000 abliterated models. Base Labs is positioning its work as a collaborative effort to create an ecosystem of safe and accessible open models.

The companies have not disclosed technical details of the partnership, but the initiative represents a significant step in addressing the risks associated with open-source AI.

Base Labs did not specify how the partnership will function technically, and the open question remains on how effectively these measures will mitigate the risks of abliteration. The companies are now seeking contributions from developers to further develop the safety framework. Looking ahead, the initiative aims to set a new standard for open-weight AI safety.

Source: techcrunch