AI Briefing
KO

Redwood Defines 'NLS Depth' for AI Monitoring; Distinguishes Kimi-K3's High Depth from DeepSeek's Overestimation

·2026.09.11 03:40

Key point

Redwood Research defined 'NLS depth' and distinguished the depth increase in Kimi-K3 using Gated DeltaNet from the overestimation in DeepSeek-V4-Pro, which does not use DeltaNet, due to Sinkhorn-Knopp operations.

Details

Redwood Research defined the 'NLS depth' metric based on Natural Language-rooted nodes (NL-rooted nodes) to measure 'opaque serial depth' within AI models. This metric is based on the depth of computational paths that do not pass through NL-rooted bottlenecks. According to the source, the opaque serial depth of standard Transformers using CoT (Chain-of-Thought) tends to be proportional to the number of layers. Kimi-K3, which uses Gated DeltaNet, was analyzed to have a higher NLS depth per layer compared to the standard Transformers used for comparison. DeepSeek-V4-Pro, which does not use DeltaNet, also showed high values, but this is distinguished as a case where depth was overestimated due to Sinkhorn-Knopp operations. The researchers pointed out that this operation, while serial, has low FLOPs and does not use learned parameters, identifying the high metric values as a flaw in the current NLS depth definition. Additionally, they revealed that most post-training methods, such as prompt engineering or general CoT training, keep nodes in an NL-rooted state and therefore do not increase NLS depth.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.