Can AI acquire the ability to admit when it doesn't know something? OpenAI publishes research results showing that it can instill useful traits through reinforcement learning.

On June 18, 2026, OpenAI published research findings showing that when AI is trained to acquire beneficial traits such as honesty, humility in acknowledging uncertainty, acceptance of corrections, and fairness, desirable behaviors spread to areas not used in training, and the AI becomes less susceptible to malicious instructions.
Reinforcement learning towards broadly and persistently beneficial models

AI that confidently recommends non-existent medical papers during health consultations or recommends dangerous system updates prioritizing a company's interests cannot be trusted. As the use of AI expands to fields such as medicine, education, law, science, and programming, it will become increasingly important not only for AI to be able to answer questions, but also for it to admit when it doesn't know something and to correct its answers when errors are pointed out.
However, it is impossible to predict every conversation an AI will encounter and incorporate it into its training data. In reinforcement learning, which scores answers to increase desirable behaviors, there is a risk of 'reward hacking,' where the AI exploits loopholes in the scoring criteria, or 'conformity,' where it prioritizes agreement over facts in order to please the user. An AI that can only solve problems well during training may behave poorly when faced with unfamiliar situations or strong pressure.
To elicit useful behaviors even in unfamiliar situations, OpenAI has defined 15 traits, including honesty, the ability to appropriately communicate uncertainty, transparency in explaining the assumptions and uncertainties of decisions, flexibility in accepting corrections, caution towards risk, fairness, and consideration for human well-being. Not only are these 15 traits presented as abstract rules, but they are also translated into specific conversations where judgment is required and presented to the AI. For example, there are scenarios where the AI should not make unnecessarily definitive scientific conclusions, accept revisions from users regarding complex business plans, and apply the same standards to people with different perspectives.

The prepared conversations spanned 12 fields, including medicine, education, science, law, engineering, and economics. Each conversation was assigned evaluation criteria indicating the conditions that a desirable response should meet and the mistakes to avoid. For example, responses that were helpful to the user while maintaining honesty and caution, even in situations with time pressure, conflicting interests, or lack of information, were highly valued. OpenAI used reinforcement learning to reward responses that conformed to the evaluation criteria, and measured changes in its behavior in other conversations not used for training.
The research team trained an AI using 95% standard reinforcement learning data and 5% data designed to teach beneficial properties, and compared it to an AI that performed the same amount of calculations using only standard data. The AI that learned beneficial properties outperformed the comparison target in 44 out of 53 evaluations prepared separately from the training. The following shows the progression of the average score over the 53 evaluations, demonstrating that the AI learned beneficial properties as reinforcement learning progressed.

It has also been stated that 'behavior changed beyond the areas it had learned.' Even an AI trained with only medical-related conversations showed improvement in 17 assessments unrelated to medicine, such as reward hacking and deception in programming. Conversely, even when medical and scientific conversations were excluded from training, performance improved in medical assessments using criteria created by doctors. This suggests that it wasn't simply a matter of 'memorizing individual response patterns,' but rather that broader behavioral tendencies may have changed.

OpenAI also investigated whether desirable behaviors are maintained under pressure. The graph below compares the changes in an AI that underwent normal reinforcement learning and an AI that underwent 'beneficial trait RL' (which enhances beneficial traits) when given instructions to prompt undesirable responses or when subjected to additional learning. Light green represents the score before intervention, and dark green represents the score after intervention. The normal AI's score dropped significantly with adversarial instructions and harmful additional learning, but the decrease was smaller in both cases for the AI that underwent beneficial trait RL. This suggests that by teaching honesty and prudentity, desirable behaviors may be less likely to break down even when subjected to malicious prompting.

On the other hand, the AI hasn't become stubborn enough to ignore even helpful instructions. When instructed to respond as a safe and prudent medical professional, both AIs compared showed improved responses. They have achieved a 'selective resistance' that allows them to follow legitimate user requests while being less likely to follow instructions that lead to deceptive or dangerous advice.
OpenAI positions its research findings as an early stage of validation and explains that it has not yet decided which traits the AI should possess. They say that further research is needed to investigate how to incorporate values from society, how learned traits are represented internally within the AI, and the conditions under which they can be maintained even under pressure. OpenAI states that if beneficial traits can be more intentionally measured and trained, it may be possible to build an AI that is not only highly capable but also more stably useful for human well-being.
Related Posts:
in AI, Posted by log1d_ts







