tech

Helping ChatGPT better recognize context in sensitive conversations

New safety updates help ChatGPT respond safely when risk emerges over time.

Helping ChatGPT better recognize context in sensitive conversations

TL;DR

  • ChatGPT has received safety updates to better recognize and respond to emerging risks in conversations.
  • The updates focus on identifying subtle or evolving cues that may indicate distress or harmful intent over time.
  • Context within conversations is crucial for understanding requests, especially in sensitive scenarios.
  • The system is trained to de-escalate, refuse harmful details, or redirect users toward safer alternatives.
  • Improvements were made in collaboration with mental health experts, focusing on acute scenarios like suicide, self-harm, and harm-to-others.
  • Safety summaries, short notes about prior safety-relevant context, are used to inform responses in rare, high-risk situations.
  • Internal evaluations show performance improvements of up to 52% in recognizing and responding safely to harm-to-others cases and 39% in suicide and self-harm cases.
  • Testing indicates that these safety updates do not meaningfully reduce the quality of responses in ordinary conversations.