/explore

Click through on any links that interest you or select the planets on the right to continue exploring the Outer Web.
You are here

simonwillison.net
| | deepmind.google
1.2 parsecs away

Travel
| | We ask the question: "What is the optimal model size and number of training tokens for a given compute budget?" To answer this question, we train models of various sizes and with various numbers...
| | blog.moonglow.ai
1.8 parsecs away

Travel
| | Parameters and data. These are the two ingredients of training ML models. The total amount of computation ("compute") you need to do to train a model is proportional to the number of parameters multiplied by the amount of data (measured in "tokens"). Four years ago, it was well-known that if
| | www.danieldemmel.me
1.8 parsecs away

Travel
| | Part two of the series Building applications using embeddings vector search and Large Language Models
| | bryanbrattlof.com
11.2 parsecs away

Travel
| Developing our first neural network with Keras