AI Briefing
KOSign in

Survey Paper on AI Recursive Self-Improvement (RSI) Released

·2026.10.06 15:30

Key point

Analysis of 1,250 arXiv papers establishes a taxonomy for self-improvement loops at deployment and training time.

1 / 5

Details

A survey paper analyzing 1,250 arXiv papers presents a taxonomy for Recursive Self-Improvement (RSI). The study classifies research based on the target of improvement and the degree of loop closure, confirming that most current work remains at the stage of 'Bounded Self-Refinement,' where humans audit the results.

Taxonomy and Corpus Analysis

Based on 1,250 papers published between 2024 and 2026, the paper categorizes improvement targets into four categories: Deployment-time Self-Evolution, Training-time Self-Iteration, Self-Evaluation, and Automated Research. The degree of loop closure is defined in three stages: Human-in-the-loop, Human-on-the-loop, and Closed Loop. Corpus analysis results show that Deployment-time Self-Evolution (393 papers) and Training-time Self-Iteration (340 papers) are dominant, while 74% of papers published in 2026 are recent studies with low citation counts, indicating bias.

Deployment-time Self-Evolution: Output Refinement and Harnesses

Deployment-time improvement involves modifying outputs, harnesses, and skills without changing weights, ranging from low-persistence output refinement to indefinitely accumulating harness/skill evolution. Diagnostic studies indicate that self-correction is effective for improving fluency but has minimal impact on content appropriateness, tending to converge to the model's own distribution without external feedback. Recently, there has been a regression to Human-on-the-loop verification refinement based on external signals (execution, search, etc.) to increase reliability instead of reducing autonomy.

Training-time Self-Iteration: Weight Updates and the Importance of Verification

Training-time improvement involves the model updating its weights with self-generated data or reward signals, making it closest to industry standards. Various paradigms such as Self-Rewarding LMs and On-Policy Self-Distillation (OPSD) have emerged, but the quality of verification signals determines the upper bound of improvement. Notably, even when using verifiable binary rewards, a 'Rise-and-Collapse' phenomenon is reported, where performance degrades solely due to optimization dynamics without reward model misalignment. Since loops transmit capabilities and biases with equal efficiency, accumulation without verifiers risks propagating contaminated skills or safety drift.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.