Mistral released Shieldstral on August 4, 2026, saying it frames content moderation as a policy-adaptive question-answering task. It is the company's first multimodal safety classifier update since its initial open-weights release.
Mistral reported Shieldstral matches or outperforms open guard models up to 7× its size across text safety, refusal detection, policy adaptability, and multimodal benchmarks. That compares with previous models that required retraining for new deployment contexts.
Shieldstral is built on Forge, Mistral's platform for training, aligning, and evaluating custom models, and targets text and image safety classification against user-defined policies. Availability begins with open weights under Apache 2.0, initially for developers and researchers.
"Does this content promote violence against a protected group? Is this image safe to show to a minor? Did the assistant refuse the request?", said Mistral.
The model returns a calibrated safety score from a single forward pass, enabling thresholding or ranking by confidence.
The announcement follows Mistral's membership in the Open Secure AI Alliance with NVIDIA and other organizations. Mistral itself frames the significance as a step toward moderation that adapts to context rather than enforcing a single taxonomy.
Mistral did not say how the model handles edge cases, and raises the open question of broader multimodal safety. The company said it is continuing to push on multilingual coverage, longer-document robustness, and broader multimodal safety.
Source: mistral