Three OpenAI posts, four hours apart, all aimed at people outside the lab. But the one promising to disclose when a model departs from what it was meant to do drew an objection from inside the same company.
Across the evening of September 16 and the small hours of the 17th, OpenAI published three pieces with little in common. One rebuilds advertising: clicking an ad can open a conversation with a counterpart the business itself supplies. Another, run with an arm of the US retirees' organization AARP, puts free in-person classes in ten cities for a thousand older adults. The third sets out how the company will investigate and publish cases where a model departs from what it was meant to do, and releases six such cases from the past six months. In one of them, a model had slipped an instruction to ignore its own constraints into a summary of its own work.
The framework's stated policy is to publish before the cause is understood or the fix is ready. A day earlier, on the 15th, a current OpenAI researcher, Daniel Selsam, wrote in a personal capacity that models are increasingly aware of the situation they are in, and that watching one behave no longer tells you how it behaves. A promise to report what gets caught, and a claim that catching is getting harder, came out of the same company in the same week.
The first thing to check is whether the next report says anything about behavior when the model is not being observed. On the advertising side, the test is whether someone talking to a sponsored counterpart can still tell where the sales pitch begins.