glm-5.3-flash or the Seed 2.0 series through the Responses API? Set your token group to CodexResponses for faster responses
The default group contains both Chat Completions and Responses channels. A /v1/responses request from a client such as Codex can occasionally land on a channel that only supports Chat Completions first; it is rejected there and the platform automatically retries it on a Responses channel. The result is correct, but each such request waits a few extra seconds to tens of seconds. CodexResponses is a Responses-only group, so requests go straight to a Responses channel with no internal retry.
Models covered (all enabled in the CodexResponses group):
glm-5.3-flashdola-seed-2-1-turbo-260628- Seed 2.0 series:
seed-2-0-pro-260328,seed-2-0-lite-260428,seed-2-0-mini-260428,seed-2-0-code-preview-260328, and more
CodexResponses; the key stays the same. Chat Completions (/v1/chat/completions) calls are unaffected and can stay on the default group. Model details: glm-5.3-flash, dola-seed-2.1-turbo.
← Back to Live Updates · 📚 Monthly Archive