PingWest reported on August 5 that Mistral has recently open-sourced Shieldstral, a 3-billion-parameter multimodal security classification model, under the Apache 2.0 license. This model supports dual-modality input of text and images and can run on a single Nvidia GPU with 16GB of VRAM, delivering performance comparable to open-source models seven times its size and achieving SOTA in multimodal content moderation. Its core innovation lies in transforming moderation policies into a natural language Q&A format, allowing adaptation to different scenarios without retraining. It is now available on the Hugging Face platform and supports 12 languages.
