|
You are here |
www.alignmentforum.org | ||
| | | | |
www.lesswrong.com
|
|
| | | | | It's currently possible to (mostly or fully) cheaply reproduce the performance of a model by training another (initially weaker) model to imitate the... | |
| | | | |
www.lesswrong.com
|
|
| | | | | This is Section 6 of "Scheming AIs." | |
| | | | |
joecarlsmith.com
|
|
| | | | | From a talk at Anthropic in April 2025. | |
| | | | |
scottaaronson.blog
|
|
| | | Update (Nov. 22): Theoretical computer scientist and longtime friend-of-the-blog Boaz Barak writes to tell me that, coincidentally, he and Ben Edelman just released a big essay advocating a version of "Reform AI Alignment" on Boaz's Windows on Theory blog, as well as on LessWrong. (I warned Boaz that, having taken the momentous step of posting... | ||