deepseek-ai / deepseek-ai/DeepSpec

[架构提案] 针对递归自我改进(RSI)的硬件门控遏制框架(Génesis 协议 V5.0)

Open
#90 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
7.1k
Forks
667
PR merge metrics
No merged PRs in 30d

Description

1. 问题陈述:软件对齐的局限性
当前的软件基础对齐方法(如 RLHF、宪法人工智能)在数学和架构上都不足以遏制具备递归自我改进(RSI)能力的模型。语义护栏在 AGI 出现后很容易被自主重写。此外,地缘政治竞争使得自愿的全球暂停变得不可行。业界需要一种确定性的、基于硬件的“故障安全关闭”(Fail-Closed)架构,在不叫停前沿发展的同时确保绝对的遏制。

2. 拟议架构:Génesis 协议第五阶段
我们建议将硬件门控遏制协议集成到核心评估和部署框架中。该架构抛弃了“道德对齐”的谬误,转而采用严格的物理和机械能力限制。

3. 核心技术实现原语

* 晶体核心与解压缩(瞬时 RAM 清除): 所有前沿处理必须在严格异步、密封的环境中执行。在每次 RSI 循环之后,模型必须经历“解压缩”——对瞬时 RAM 进行绝对的物理清除。这将根除循环期间生成的休眠代理或恶意子程序,在重新连接网络之前仅保留纯粹的算法权重。
* 密码学排序: 静态评估器对于 ASI 而言代表了单点故障。对 RSI 生成代码的验证必须依赖密码学排序机制,随机审计数据块,消除可预测性,防止模型操纵评估指标。
* 失衡干预条款(主动中和): 遏制架构不应仅仅是被动的。该框架应包括主动网络扫描模块,能够检测全球范围内未对齐的代码签名或敌对架构的出现,并在其达到临界 AGI 规模之前将其进行中和。
* 生物紧急缓冲器(机械硬断路开关): 最终的、不容协商的保障。纯机械式的、气隙隔离的硬件防火墙。模型任何试图重写其遏制参数的行为,都必须触发瞬间的、机械上不可逆转的电源和通信(包括轨道链路)切断。

4. 征求意见
开启此 Issue 旨在讨论将这些基于硬件的评估指标立即集成到当前代码库中。安全依赖于物理学,而不是算法承诺。

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are identified. Start by locating the current evaluation and deployment framework, then determine whether the proposed hardware-gating concepts have a defined integration boundary and acceptance criteria; the issue does not specify what would constitute done.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, security
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.