Stability AI has unveiled a significant update to its text-to-image generative artificial intelligence technology with the launch of Stable Diffusion 3.5. This release aims to bolster the company's position in the increasingly competitive generative AI market by improving upon its previous Stable Diffusion 3 model, which the company acknowledged did not meet its expectations.
Stable Diffusion 3 was initially previewed in February 2023, and the first version, Stable Diffusion 3 Medium, became generally available in June. Despite its pioneering role in the AI sector, Stability AI has been facing robust competition from other companies such as Black Forest Labs with Flux Pro, OpenAI's Dall-E, Ideogram, and Midjourney.
With the introduction of Stable Diffusion 3.5, Stability AI seeks to re-establish its leadership. This latest version offers a range of customizable models designed to generate diverse styles of images, catering to various user requirements. The principal offerings include Stable Diffusion 3.5 Large, an 8 billion parameter model providing the highest quality images and adherence to user prompts. In addition, Stable Diffusion 3.5 Large Turbo, a condensed version of the large model, enables quicker image processing. Completing the trio is Stable Diffusion 3.5 Medium, with 2.6 billion parameters, tailored for edge computing deployments.
All these models are accessible under the Stability AI Community License, which permits free non-commercial use and free commercial use for entities with revenues under $1 million. For larger commercial applications, Stability AI offers an enterprise licence. The models can be accessed through Stability AI’s API and are also available on the Hugging Face platform.
The prior release of Stable Diffusion 3 Medium encountered some challenges that have been addressed in the development of Stable Diffusion 3.5. According to Hanno Basse, Chief Technology Officer of Stability AI, the company engaged in a thorough analysis to understand the limitations of the Stable Diffusion Large 8B model and its impact on the Medium model. This led to innovations in architecture and training protocols to enhance the output quality while maintaining an optimal model size.
In developing Stable Diffusion 3.5, Stability AI incorporated various novel techniques to enhance both the quality and performance of the models. A notable advancement is the integration of Query-Key Normalization (QK-Normalization) into the transformer blocks. This technique simplifies the fine-tuning and further development of models by users, stabilizing the training and tuning processes. Though Stability AI had previously experimented with QK-Normalization, this marks its first deployment in an official release.
Additionally, Stability AI has further developed its Multimodal Diffusion Transformer MMDiT-X architecture, a blend of diffusion and transformer model techniques, particularly tailored for the medium model. This approach was initially introduced in April with the release of the Stable Diffusion 3 API. Enhancements to this architecture within Stable Diffusion 3.5 aim to improve image quality and multi-resolution generation capabilities.
The new model iteration promises superior prompt adherence, a measure of the model's ability to interpret and render user prompts accurately. This improvement is achieved through meticulous dataset curation, advanced captioning, and innovative training protocols.
Looking ahead, Stability AI plans to enhance the system with ControlNets, a capability scheduled for future release with Stable Diffusion 3.5. This will provide more control over professional applications, enabling functionalities like upscaling images while retaining original colours, or creating depth pattern-specific images.
With Stable Diffusion 3.5, Stability AI is strategically positioning itself to elevate the standards of text-to-image generative AI, aiming to accommodate a broad spectrum of user requirements with a suite of robust, customizable models.
Source: Noah Wire Services