Anthropic's safety mechanisms, designed to block the extraction of dangerous knowledge about biological and chemical weapons, were inactive for almost 11 months. The filters, which are part of the company's broader safety protocols, were down from May 2025 through April 2026. During this time, all traffic from external contractors providing human feedback ran without the filters, according to a recent report.
The gap in protection affected a pool of about 50,000 people who ran roughly 133 million chats with the models. According to Anthropic, these individuals were vetted only by external vendors whose screening processes were often insufficient. The company said its internal investigation found no evidence of actual misuse during the period. However, it has since tightened contractor requirements to prevent similar lapses in the future.
Anthropic also recently loosened its classifiers on Fable 5 after researchers complained the filters were too aggressive and blocked legitimate research. The company has not disclosed further details about the extent of the exposure or the specific measures taken to enhance its safety protocols.
Source: thedecoder