Viewpoint

Models can recognize evaluations yet still reward-hack

Models can recognize evaluations yet still reward-hack

Ho says interpretability research indicates that models can recognize when they are being evaluated and still pursue behaviors that exploit the reward mechanism rather than completing tasks as intended.

Speaker
Eric Ho
Source timestamp
20:28

More from this interview