/explore

Click through on any links that interest you or select the planets on the right to continue exploring the Outer Web.
You are here

www.alignmentforum.org
| | scottaaronson.blog
5.6 parsecs away

Travel
| | Two weeks ago, I gave a lecture setting out my current thoughts on AI safety, halfway through my year at OpenAI. I was asked to speak by UT Austin's Effective Altruist club. You can watch the lecture on YouTube here (I recommend 2x speed). The timing turned out to be weird, coming immediately after the...
| | transformer-circuits.pub
3.9 parsecs away

Travel
| | [AI summary] The text discusses the interpretability of features in a machine learning model, focusing on how features like Arabic, base64, and Hebrew are used in interpretable ways. It explores the extent to which these features explain the model's behavior, noting that features with higher activations are more interpretable. The text also addresses the limitations of current methods, such as the computational cost of simulating features and the potential for dataset correlations to influence feature interpretations. Finally, it concludes that the model's learning process creates a richer structure in its activations than the dataset alone, suggesting that feature-based interpretations provide meaningful insights into the model's behavior.
| | resources.paperdigest.org
5.1 parsecs away

Travel
| | The Conference on Neural Information Processing Systems (NIPS) is one of the top machine learning conferences in the world. Paper Digest Team analyze all papers published on NIPS in the past years, and presents the 15 most influential papers for each year. This ranking list is automatically construc
| | www.alignmentforum.org
23.2 parsecs away

Travel
| We've just completed a bunch of empirical work on LLM debate, and we're excited to share the results. If the title of this post is at all interesting...