Stable Diffusion 3.5 Optimized with TensorRT Delivers 2x Performance and 40% Memory Savings on NVIDIA RTX GPUs
Key point
Stability AI partnered with NVIDIA to release TensorRT-optimized Stable Diffusion 3.5 models, boosting performance by up to 2.3x.
Details
Stability AI partnered with NVIDIA to release optimized Stable Diffusion 3.5 (SD3.5) models applying TensorRT and FP8 quantization technology. This optimization increases image generation speed and significantly reduces VRAM requirements on supported NVIDIA RTX GPUs.
The SD3.5 Large model achieves generation speeds up to 2.3x faster compared to the original PyTorch model, with VRAM usage reduced from 19GB to 11GB, a decrease of about 40%. The SD3.5 Medium model shows a 1.7x improvement in generation speed, enabling efficient workflows.
The optimized models run smoothly on the following hardware:
- NVIDIA GeForce RTX 40 and 50 series
- RTX PRO GPUs based on NVIDIA Blackwell and Ada Lovelace architectures
These models are available for both commercial and non-commercial use under the Stability AI Community License. Model weights can be found on Hugging Face, and the related code is available on NVIDIA GitHub.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.