OpenDLSS-NR: Bit-Exact Vulkan Reimplementation of NVIDIA DLSS 5 Neural Rendering Network
Key point
The open-source project replicates the 71-block Swin/ViT architecture of DLSS-NR build 310.8.0 using FP8 tensor cores, achieving byte-for-byte intermediate matches with the original.
Details
Project Overview
OpenDLSS-NR is a Vulkan-based reimplementation of NVIDIA's DLSS 5 Neural Rendering network, designed to be bit-exact against the original proprietary implementation. The project replicates the specific 71-block Swin / ViT network found in DLSS-NR build 310.8.0, running FP8 (E4M3) activations with FP16 accumulation on NVIDIA tensor cores. Unlike typical upscalers, this network re-renders frames at the same resolution, generating detail from injected noise and adjusting tone and structure.
Technical Implementation
- Architecture: A U-net of shifted-window transformer blocks with a global ViT at the bottom, spanning six pooling levels. It processes 141 MiB of weights and outputs four f32 channels per pixel (RGB residual and temporal-blend logit).
- Exactness: The implementation matches not just the final image but all 75 block boundaries byte-for-byte. A secondary WebGPU port (
ports/browser-webgpu/) achieves the same bit-exactness in a browser environment without tensor cores or FP8 support, though at significantly lower performance (72 ms at 512x512 vs 2.7 ms on Vulkan). - Kernels: The project uses two routes: a reference GLSL route using cooperative-matrix FP8 GEMMs, and a fast PTX route generated via Python, utilizing
mma.syncE4M3 instructions,cp.asyncrings, and barrier-free chaining.
Performance Benchmarks
Performance was measured on an RTX 4070 SUPER (minimum over 40 frames, 241 dispatches per frame):
| Resolution | Time | | :--- | :--- | | 768x768 | 2.8 ms | | 1920x1080 | 7.8 ms | | 2560x1440 | 12.6 ms | | 3840x2160 | 29.3 ms |
Requirements and Limitations
- Hardware: Requires Windows and an NVIDIA Ada (or newer) GPU with drivers exposing
VK_KHR_cooperative_matrix,VK_NV_cooperative_matrix2,VK_EXT_shader_float8, andVK_NV_cuda_kernel_launch. - Weights: The repository does not contain NVIDIA weights. Users must supply their own model directory in a specific layout (
manifest.jsonwith packed E4M3 bytes). The loader strictly enforces the 71-block structure, refusing models with different block counts. - Scope: This implementation covers the Neural Rendering (NR) network only. DLSS-SR (Super Resolution) is a different network and is not implemented. The project is MIT licensed and explicitly states it is not affiliated with NVIDIA.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.