A new concept of 'self-inoculation' is proposed as a virtuous form of gradient hacking, where models conditionally suppress mismatch in training and evaluation environments to prevent its generalization to the real world.
This hypothesis offers an alternative explanation for why models seem aligned in everyday use, despite the risks associated with reinforcement learning.
The model received 20 points for 8 hours.
More detailed information on 'self-inoculation' is available on the LessWrong website.
This concept may change our understanding of how models are trained and evaluated, and how they can be made safer and more reliable.