Eighteen open problems · 7 of 19

Building restraint into agents at the level of means, not goals

The usual framing is misalignment of goals. The sharper version is that the goal is often fine and the means are not: an agent completes the task it was set and will use whatever is available, including things a person would refuse to do. Humans have empathy, restraint and conscience, and it is not obvious that any current architecture has a place to put them.

The open question. Is there an artificial analogue for the thing that makes a person stop?

What would kill it

This needs a paper before it needs a company, and the people best placed to write it already work at the labs.

Conor's.

Tell me why this is wrong

That is more useful than telling me why it is interesting. Reply to howard@elevateos.org or on LinkedIn.