Stability AI Unveils Stable Diffusion 3.5 Series Models
Stability AI has announced the launch of its latest text-to-image generation models, Stable Diffusion 3.5 Large and Stable Diffusion 3.5 Large Turbo. These models represent a significant advancement in the company's technology, focusing on deliverables such as increased customizability, efficiency, and flexibility. Both models come with a licensing model that permits free use for non-commercial purposes and limited commercial use.
The Stable Diffusion 3.5 Large model boasts an impressive 8 billion parameters and is capable of generating professional-grade images at a resolution of 1 megapixel. Stability AI claims that this model offers superior results in terms of prompt adherence and overall image quality. Its counterpart, Stable Diffusion 3.5 Large Turbo, is a distilled version that enhances speed by condensing the number of operational steps to four.
A key feature of these models is their emphasis on customizability. Users are given the opportunity to fine-tune the model or construct personalised workflows to suit their specific needs. The models are also designed to be compatible with standard consumer hardware, making them accessible to a wider audience. The diversity of output is another focal point, as the models can produce a variety of styles including different skin tones, 3D visuals, photography, and painting.
This release comes after the previous iteration, Stable Diffusion 3 Medium, which was introduced last June and faced criticism particularly in its depiction of human anatomy, notably hands. In acknowledging community feedback, Stability AI describes Stable Diffusion 3.5 not as an immediate solution, but as a progressive step in the model's overall development. While certain known issues have been addressed, ongoing challenges remain with basic prompt execution.
The architecture of Stable Diffusion 3.5 retains similarities with its predecessor but introduces two significant changes: the incorporation of QK normalization and double attention layers. These modifications aim at enhancing the model's performance and usability.
The release of these models is complemented by a licensing framework that supports free use for projects with non-commercial intent, and for commercial ventures with an annual revenue not exceeding $1 million. However, the license contains specific restrictions against developing competing foundational models. This is balanced by allowances for custom models developed through techniques like LoRAs and hypernetworks.
Looking forward, Stability AI has scheduled the release of Stable Diffusion 3.5 Medium later this month. This version will use 2.5 billion parameters and is also tailored to run on consumer hardware. However, users may experience a minimal reduction in output quality compared to the Large models.
The Stable Diffusion 3.5 inference code is currently accessible on GitHub, while the model itself is hosted on huggingface. Additionally, users can interact with the model via platforms such as Replicate, ComfyUI, and DeepInfra, or through Stability AI’s own API, providing multiple avenues to utilise the new technology.
Source: Noah Wire Services