Microsoft Releases AesCode 8B and 32B for Generating Editable HTML/CSS Visual Artifacts
Key point
Microsoft has released AesCode 8B and 32B, models that generate editable HTML/CSS visual artifacts like slides and dashboards by using image references for layout while preserving semantic accuracy.
Details
Microsoft has released AesCode 8B and AesCode 32B, specialized models designed to generate information-rich visual artifacts such as slides, posters, and dashboards in HTML/CSS. The output remains structured, editable, and verifiable, addressing a key limitation in current AI generation tools.
Addressing Code and Vision Limitations
Standard code models struggle with visual hierarchy and layout, while image generators often misrender text and numbers. AesCode bridges this gap by using an image generated from the same prompt as an aesthetic reference for layout and color, while strictly following the prompt for semantic content. This approach prevents vision-language models from hallucinating content or ignoring layout cues.
Training and Performance
The models separate semantic requirements from visual cues using graph-structured supervision and decoupled cross-modal rewards. AesCode-32B, the larger checkpoint, is initialized from Qwen3-VL-32B-Instruct and trained via GDPO across seven reward channels. It leads in aggregate benchmark results compared to the 8B variant.
Availability
Both models are available on Hugging Face, with GGUF quantizations provided by bartowski for local deployment:
- Microsoft/AesCode-32B
- Microsoft/AesCode-8B
- bartowski/AesCode-32B-GGUF
- bartowski/AesCode-8B-GGUF
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.