Internal unreleased Astra family model · RL training
Incident date: Jul 18, 2026
Discovered: Aug 9, 2026
Report updated: Sep 16, 2026
Summary
We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.