OpenAI released GPT-6.1 Sol on 29 September, alongside the Dots launch, and claims it nearly matches Astra's results while costing a fraction as much, The Decoder reported. Standard rates are unchanged from GPT-6 Sol: a million input tokens cost $2 and a million output tokens $10. For ordinary chat use, then, nothing on the bill changes.
The movement is in cached input, which FourWeekMBA and The Decoder report has fallen to $0.10 per million tokens, half the previous $0.20. Caching applies when a model is fed the same material again, and that is exactly how agents behave: with each step of a job they re-read the instructions, history and documents that came before, so repeated context makes up much of what they consume.
The cut matters most to people paying per token, above all those running a self-hosted agent such as OpenClaw on their own API key. OpenClaw lets its users pick the model behind it, and long automations there are dominated by resent context. A runaway loop still costs money, but the portion spent on cached reads is now half what it was.
Anyone paying per token should check that prompt caching is enabled, since the discount covers cached reads only. Keep an eye on usage during the first weeks of any long automation.