|
You are here |
www.alignmentforum.org | ||
| | | | |
vkrakovna.wordpress.com
|
|
| | | | | (This post is based on an overview talk I gave at UCL EA and Oxford AI society (recording here). Cross-posted to the Alignment Forum. Thanks to Janos Kramar for detailed feedback on this post and to Rohin Shah for feedback on the talk.) This is my high-level view of the AI alignment research landscape and... | |
| | | | |
www.greaterwrong.com
|
|
| | | | | The DeepMind mech interp team has pivoted from chasing the ambitious goal of complete reverse-engineering of neural networks, to a focus on pragmatically making as much progress as we can on the critical path to preparing for AGI to go well, and choosing the most important problems according to our comparative advantage. We believe that this pragmatic approach has already shown itself to be more promising. We don't claim that these ideas are unique, indeed we've been helped to these conclusions by the thoughts of many others both in academia (1 2 3) and the safety community (1 2 3). But we have found this framework helpful for accelerating our progress, and hope to distill and communicate it to help other have more impact. We close with recommendations for h... | |
| | | | |
thezvi.wordpress.com
|
|
| | | | | The cycle of language model releases is, one at least hopes, now complete. OpenAI gave us GPT-5.1 and GPT-5.1-Codex-Max. xAI gave us Grok 4.1. Google DeepMind gave us Gemini 3 Pro and Nana Banana Pro. Anthropic gave us Claude Opus 4.5. It is the best model, sir. Use it whenever you can. One way Opus... | |
| | | | |
cset.georgetown.edu
|
|
| | | Place to find CSET's publications, reports, and people | ||