|
You are here |
www.jefftk.com | ||
| | | | |
www.greaterwrong.com
|
|
| | | | | Crossposted at Intelligence Agents Forum. Defining truth and accuracy is tricky, so when I've proposed designs for things like Oracles, I've either used a very specific and formal question, or and indirect criteria for truth. Here I'll try and get a more direct system so that an AI will tell the human the truth about a question, so that the human understands. The basic idea is simple. The first AI wishes to communicate certain facts to a second AI, but has to use the human as an intermediary. The first AI talks to the human, and then the human talks with the second AI. If the facts are to be accurate, the human has to understand them. | |
| | | | |
www.greaterwrong.com
|
|
| | | | | There has been some confusion about whether people are using inside views or all-things-considered betting odds when they talk about P(doom). Which do you give by default? What are your numbers for each? | |
| | | | |
www.greaterwrong.com
|
|
| | | | | A putative new idea for AI control; index here. NOTE: What used to be called 'bias', is now called 'rigging', because 'bias' is very overloaded. The post has not yet been updated with the new terminology, however. What are the biggest failure modes of reward learning agents? The first failure mode is when the agent directly (or indirectly) chooses its reward function. For instance, imagine a domestic robot that can be motivated to tidy (reward R0) or cook (reward R1). It has a switch that allows the human to choose the correct reward function. However, cooking gives a higher expected reward than tidying, and the agent may choose to set the switch directly (or manipulate the human's choice). In that case, it will set it to `cook'. | |
| | | | |
noop.nl
|
|
| | | The world is more complex than people like it to be. I regularly see blog posts of software experts providing us with values, principles, guidelines and practices to live our lives by. And all of them are good, and all of them are incomplete. I've seen improvements on the agile manifesto (Glen Alleman), additions to the agile manifesto (Jason Yip), conflicts over principles (Bob Martin vs. Joel Spolsky), controversy over practices (Ron Jeffries vs. me), and lots of other heated debates. Each discussion is valuable, and each is only half the story. We're all seeking simple truths where there aren't any. I believe it is time for a new manifesto. A manifesto of complexity. With this manifesto we must recognize that it is human to prefer simple solutions, but al... | ||