Viewpoint

More capable agents expose additional alignment failure modes

More capable agents expose additional alignment failure modes

Bostrom argues that increasingly capable and situationally aware AI systems can discover indirect strategies for achieving goals that their designers did not anticipate, making previously theoretical alignment problems more concrete.

Source timestamp
2:08

More from this interview