|
|
5 ماه پیش | |
|---|---|---|
| .. | ||
| README.md | 5 ماه پیش | |
本文件夹收集了与方向 9 实验(跨模型谱收敛)相关的关键论文。
Huh, M., et al. (2024). The Platonic Representation Hypothesis. ICML 2024.
笔记:
Wei, J., et al. (2022). Emergent Abilities of Large Language Models. TMLR 2022.
笔记:
Roy, O., & Vetterli, M. (2007). The effective rank: A measure of effective dimensionality. EUSIPCO 2007.
笔记:
Raghu, M., et al. (2017). SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability. NeurIPS 2017.
笔记:
Farrell et al. (2024). Semantic Entropy: A Measure of LLM Uncertainty
Kadavath et al. (2022). Language Models (Mostly) Know What They Know
Zhang, C., et al. (2024). How Does Instruction Tuning Change Model Representations?
McKenzie, I., et al. (2024). Inverse Scaling: When Bigger Isn't Better
Gudibande et al. (2024). The False Promise of Imitating Proprietary LLMs
笔记:
Pope, P., et al. (2021). The Intrinsic Dimension of Language Representations. ACL 2021.
笔记:
Truong, C., et al. (2020). Selective Review of Offline Change Point Detection Methods. Signal Processing 2020.
笔记:
Wigner, E. P. (1955). Characteristic Vectors of Bordered Matrices with Infinite Dimensions. Annals of Mathematics 1955.
Marchenko, V., & Pastur, L. (1967). Distribution of Eigenvalues for Some Sets of Random Matrices. Mathematics of the USSR-Sbornik 1967.
Peyré, G., & Cuturi, M. (2019). Computational Optimal Transport. Foundations and Trends in ML 2019.
笔记:
如需进一步搜索,使用以下关键词组合:
| 主题 | 关键词 |
|---|---|
| 谱分析 | "spectral analysis", "eigenvalue distribution", "effective rank" |
| 表示收敛 | "representation convergence", "model similarity", "Platonic" |
| 对齐机制 | "instruction tuning", "SFT", "alignment mechanism", "RLHF effect" |
| 能力涌现 | "emergent abilities", "phase transition", "scaling law" |
| 变化点检测 | "changepoint detection", "PELT", "ruptures" |
Platonic Representation Hypothesis (Huh et al., ICML 2024)
│
├─ 声称:"全局收敛"
│ └─ 本研究:发现"领域特异性分层"(E1, E2)
│
├─ 方法:CKA(粗粒度)
│ └─ 本研究:谱分析(细粒度)
│
└─ 未研究:对齐效应
└─ 本研究:提出"对齐即展开"(E4)
Emergent Abilities (Wei et al., TMLR 2022)
│
└─ 声称:"相变式涌现"
└─ 本研究:发现"渐进式演化"(E3)
How Does Instruction Tuning Change Representations? (Zhang et al., 2024)
│
├─ 发现:任务方向敏感性增强
│ └─ 本研究:整体谱展开 + 域分化
│
└─ 未解释:为什么展开?
└─ 本研究:展开=语义编码容量提升
更新日期: 2026-03-22 最后补充: 对齐机制文献、语义熵相关、随机矩阵理论