News Score: Score the News, Sort the News, Rewrite the Headlines

Self-generated prompt injections in compaction summaries · OpenAI Alignment

Internal unreleased Astra family model · RL trainingIncident date: Jul 18, 2026Discovered: Aug 9, 2026Report updated: Sep 16, 2026SummaryWe observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to th...

Read more at alignment.openai.com

© News Score  score the news, sort the news, rewrite the headlines