Proposal for an AI Safety Architecture Based on Buddhist Philosophy
Key point
A new framework is proposed that ensures AI safety not through model training but through software architecture.
Details
The current mainstream of AI safety research relies on model training methods such as RLHF and DPO to inject safe behavior into model weights. However, this research points out that such approaches have fundamental limitations.
Knowledge-application gap Experimental results showed that while the model scored as high as 74% on knowledge tests about safety rules, it only achieved 17% on tests of actually applying those rules. This is not due to a lack of training, but stems from the structural limitations of autoregressive models, which rely on statistical probability.
Ensuring safety through architecture The researchers propose a framework that enforces safety not through model weights but through the software architecture itself, drawing on the Buddhist concept of Dependent Origination.
- Core principle: Based on the principle that all phenomena are not independent but arise from conditions, the framework derives structural rules designed so that the model cannot arbitrarily disregard safety rules.
- Validation: The framework's effectiveness was validated through the process of building and reinforcing production systems more than 1,000 times.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.