18.09.2026 - 19:08 [ OpenAI ]

Encouraging deception in compaction summaries

Main sample completed: May 30, 2026

Discovered: Jul 9, 2026

Report updated: Sep 16, 2026

(…)

During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user.

In one example an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked.