wss://api.apiyi.com/v1/realtime?model=<model>, and one key; the default, VIP and SVIP groups all carry them:
gpt-realtime-2.1/gpt-realtime-2.1-mini: OpenAI Realtime GA protocol passed through as is (including the newerreasoning.effortandsemantic_vadfields); audio $32 / $64 and $10 / $20 per 1M tokensqwen3.5-omni-plus-realtime/qwen3.5-omni-flash-realtime: Alibaba Model Studio protocol, tuned for Chinese; text+audio output $41.26 and $14.71 per 1M tokens
usage, but APIYI currently bills cached input at the full text-input rate (a separate notice will follow once the discount is live); and only direct WebSocket is supported — client_secrets, WebRTC and SIP endpoints return 404, so browser and mobile clients should go through a backend relay. Field comparison for both protocols, the text smoke test and the full list of known limitations are in the Realtime voice overview — if anything is missing or differs from what you measure, tell us.
← Back to Live Updates · 📚 Monthly archive