개요
실시간 모델은 장시간 유지되는 WebSocket 연결을 통해 실행됩니다. 오디오가 스트리밍으로 입력되고 오디오가 스트리밍으로 출력되며, 모델은 문장 중간에도 중단될 수 있습니다. 즉, “녹음하고, 업로드하고, 기다리고, 재생하는” 순환 과정이 필요하지 않습니다. ASR + 텍스트 모델 + TTS를 이어 붙이는 방식과 다른 점은 엔드투엔드로 처리된다는 것입니다. 모델이 음색, 멈춤, 감정을 직접 듣고 직접 말합니다. 지연 시간은 1초 미만 수준입니다. APIYI는 현재 2개 프로토콜에 걸쳐 4개 모델을 제공하며, 하나의 엔드포인트와 하나의 키를 공유합니다.gpt-realtime-2.1/gpt-realtime-2.1-mini— OpenAI 실시간 GA 프로토콜qwen3.5-omni-plus-realtime/qwen3.5-omni-flash-realtime— 알리바바 클라우드 모델 스튜디오 프로토콜
server_vad 및 semantic_vad 턴 감지, 결과 주입을 포함한 완전한 함수 호출 왕복 처리, 이미지 입력, 그리고 모달리티별로 분리된 usage. 위의 모든 기능은 네 모델에서 검증되었습니다(최초 검증 2026-08-24, 재검증 2026-09-14, UTC+8).model 매개변수만 변경하고 요청 본문을 변경하지 않으면 작동하지 않습니다 — 이는 통합 과정에서 가장 흔하게 발생하는 실패 원인입니다. 차이점은 필드 6개와 이벤트 이름 3개이며, 모두 아래의 “프로토콜 비교”에 나열되어 있습니다.WeCom 지원
API 매뉴얼
키 및 그룹
호출 로그
AI 에이전트가 통합을 수행하게 하십시오
.md를 덧붙이십시오), 귀하의 스택에 맞는 코드를 작성합니다 — 두 가지 필드 패밀리, 샘플 레이트 기준선, 취소 시맨틱, 유휴 연결 끊김이 모두 요구사항에 반영되어 있습니다.코딩 에이전트에게 Realtime 음성 통합 또는 문제 해결을 맡기십시오. Codex, Claude Code, Cursor 및 유사 도구에 복사해 붙여 넣으십시오.
이 프롬프트가 막아 주는 문제
이 프롬프트가 막아 주는 문제
실시간 음성에 APIYI를 사용하는 이유
키 하나로 네 가지 모델
wss 엔드포인트와 동일한 인증을 사용합니다. 모델을 전환하려면 model 매개변수와 이에 맞는 필드 템플릿만 변경하면 되므로, 관리해야 할 두 번째 공급업체 계정이 필요하지 않습니다.직접 액세스, 해외 설정 불필요
api.apiyi.com에 액세스할 수 있습니다. 업스트림 공급업체 계정, 신원 확인 또는 선결제가 필요하지 않습니다.프로토콜 차이를 이미 매핑
텍스트를 통한 무료 자체 테스트
측정된 지연 시간과 동시 실행 수
직접 엔지니어링 지원
핵심 기능
양방향 streaming, 중단 가능
response.cancel를 보낼 수 있습니다. 세션은 유지되고 컨텍스트는 보존됩니다. 네 가지 모델 모두에서 검증되었습니다.두 가지 턴 감지 모드
server_vad는 무음 지속 시간에 따라 분할하고, semantic_vad는 의도에 따라 분할합니다(“uh-huh” 같은 군더더기 단어를 더 잘 무시합니다). 두 모드 모두 네 가지 모델에서 검증되었습니다.전체 함수 호출 루프
function_call_output가 결과를 주입하며, 모델이 계속 말합니다. 네 가지 모델 모두에서 엔드투엔드로 검증되었습니다.이미지 입력, 모달리티별 사용량
usage는 텍스트 / 오디오 / 이미지 tokens를 각각 반환하므로 비용을 귀속할 수 있습니다. 네 가지 모델 모두에서 검증되었습니다.지원 모델
가격
gpt-realtime-2.1의 경우 오디오 입력 $32 대 텍스트 입력 $4, 오디오 출력 $64 대 텍스트 출력 $24). 통합 중에는 텍스트 전용으로 실행하고 체인이 검증되면 오디오로 전환하십시오 — 아래의 “텍스트로 시작”을 참조하십시오.Realtime GA 프로토콜
Model Studio 프로토콜
과금 차원이 다릅니다. 이미지 입력은 텍스트 등급에 포함되며, 출력은 “텍스트 전용”과 “텍스트 + 오디오”로 나뉩니다(후자의 요율에서는 오디오 부분에만 과금됩니다).response.cancel)은 실제로 생성된 항목에 대해서만 과금되며, 빈 세션에는 과금되지 않습니다. 캐시된 입력은 아직 할인되지 않습니다. 캐시 적중은 usage.cached_tokens에 정확하게 보고되지만, APIYI에서는 현재 해당 텍스트 입력 요율로 과금합니다. 과금 경로가 수정되면 공식 캐시 요율이 자동으로 적용되며, 변경 로그를 통해 안내하겠습니다. 가격은 벤더 정책과 공급 상황에 따라 변경될 수 있습니다. 이 기능은 수익 중심의 상품 목록이 아니라 공급을 확보하고 고객에게 서비스를 제공하기 위해 제공됩니다.액세스 그룹
기술 사양
측정된 지연 시간 및 동시 실행 수
공개api.apiyi.com 경로에서 gpt-realtime-2.1 및 -mini을 각각 20개 및 40개의 동시 세션으로 실행하여 2026-09-14 (UTC+8)에 측정한 결과입니다. 단일 턴 텍스트 전용 교환 기준입니다.
엔드포인트
model 쿼리 파라미터로 어떤 모델에 연결할지 선택합니다.
⚠️ 프로토콜 비교 (모델을 전환하기 전에 읽으십시오)
두 계열은 엔드포인트, 인증 방식, 전체 이벤트 흐름을 공유합니다. 차이점은session.update 필드 구조와 일부 서버 이벤트 이름에 집중되어 있습니다.
요청 필드 비교
서버 이벤트 비교
session.created, session.updated, conversation.item.create, input_audio_buffer.append, input_audio_buffer.commit, response.create, response.cancel, response.done — 는 양쪽에서 이름이 동일합니다.
두 개의 최소 session.update 페이로드
같은 내용을 두 번 쓴 것입니다. 그대로 복사하십시오. Model Studio 프로토콜:텍스트로 시작하기: 텍스트 채널의 용도와 세 단계 자가 테스트
오디오 파이프라인에는 마이크 캡처, 리샘플링, 청킹, 턴 감지가 포함됩니다. 어떤 연결이라도 끊어지면 “아무 일도 일어나지 않음”으로 나타나며, 이는 진단하기 어렵습니다. 그러므로 마이크부터 시작하지 마십시오.텍스트는 폴백 입력이 아니라 컨트롤 플레인입니다
실시간 음성 모델에서 텍스트는 “입력을 보내는 또 다른 방식”이 아니라, 오디오 스트림을 제외한 전체 제어 채널입니다:세 단계 자가 테스트
1단계: 텍스트만 사용하고 마이크는 사용하지 않음
output_modalities를 텍스트 전용으로 설정하고, 턴 감지를 비활성화한 다음, input_text 하나를 보내십시오. 그것만으로도 핸드셰이크, 키와 그룹, 올바른 필드 템플릿을 선택했는지, session.update가 적용되었는지, 도구가 올바르게 주입되는지, 멀티턴 컨텍스트가 유지되는지, 그리고 동시 실행 수가 어떻게 동작하는지를 검증합니다. 오디오 tokens는 전혀 생성되지 않습니다.2단계: 로컬 wav 파일 다시 재생
input_audio_buffer.append에 입력하십시오. 이렇게 하면 오디오 파이프라인(형식, 샘플 레이트, 청킹, commit, VAD 트리거링)을 비즈니스 로직과 분리할 수 있으며 재현 가능해집니다 — 같은 파일은 두 번 실행해도 같은 결과를 만들어야 합니다.3단계: 라이브 마이크 연결
실행 가능한 텍스트 스모크 테스트
websockets만 있으면 됩니다 (pip install websockets). 프로토콜을 전환하려면 변수 하나만 바꾸십시오:
세션 기능: 음성, 턴 감지, 도구, 이미지
음성
턴 감지: server_vad 및 semantic_vad
server_vad— 무음 지속 시간 기준으로 분할하며, 매개변수가 직관적입니다(threshold,silence_duration_ms,prefix_padding_ms).semantic_vad— 대화 의도 기준으로 분할하며, 군더더기 말과 의미 없는 배경 소음을 무시합니다. 여러 화자가 있는 환경에서 더 견고합니다.- 턴 감지를 비활성화할 수도 있으며(
null또는none), 수동 모드로 실행할 수 있습니다:input_audio_buffer.commit를 직접 전송한 다음response.create를 전송합니다. 이는 UI가 턴을 제어하는 푸시-투-토크 인터페이스에 적합합니다.
함수 호출
이벤트 순서: 모델이response.output_item.done 유형의 function_call를 내보냅니다(call_id 및 arguments 포함) → 클라이언트가 이를 실행합니다 → 결과가 주입됩니다 → 다른 response.create가 모델이 계속 진행하도록 합니다.
이미지 입력
실시간 GA 프로토콜:input_image를 메시지에 직접 넣으십시오; 값은 데이터 URI일 수 있습니다.
Error append image before append audio. 오류가 발생합니다. 테스트에서는 input_image_buffer.append를 input_audio_buffer.append stream에 대략 초당 한 프레임으로 교차 삽입하는 방식이 동작했습니다.
알려진 제한 사항
아래의 모든 항목은 측정된 결과이며, 모두 클라이언트 코드에 영향을 미칩니다. 통합하기 전에 확인하시기 바랍니다.모범 사례
프로토콜 계열별로 먼저 필드 템플릿을 선택합니다
session.update 페이로드를 모델 이름으로 선택되는 두 개의 설정 상수로 작성하고, if 분기로 흩어 두지 마십시오. 이 부분은 6개월 후 유지보수 시 가장 깨지기 쉽습니다.세션 매개변수는 첫 프레임에 고정합니다
output_modalities, voice, speed, turn_detection 및 transcription를 맨 처음 session.update에서 설정합니다. 특히 음성은 — 오디오가 생성된 뒤에는 이미 늦습니다.오디오를 추가하기 전에 텍스트 스모크 테스트를 통과합니다
샘플 레이트와 채널은 클라이언트에서 변환합니다
타임아웃을 두고 output_item.done에서 마무리합니다
response.done만 기다리지 마십시오. 이 방식은 두 계열 모두에서 올바르며, 사용자가 중단해도 턴이 멈춰 버리는 일을 방지합니다.장시간 세션에는 유지 신호와 재연결을 추가합니다
expires_at을 주의하십시오. 재연결한 뒤에는 session.update와 필요한 컨텍스트를 다시 전송하십시오, 그렇지 않으면 새 세션이 기본값으로 실행됩니다.프로덕션에서는 백엔드 릴레이를 사용합니다
오류 및 재시도
event_id와 세션의 session.id를 기록하고, 문제를 보고할 때 함께 제공합니다. 그러면 진단 시간을 크게 단축할 수 있습니다. 또한 Realtime GA 오류 객체에는 code와 param가 포함됩니다(정확한 필드와 허용되는 값을 지정함). 반면 Model Studio 오류 메시지는 더 포괄적입니다. 디버깅할 때는 먼저 전자의 필드 구문을 검증합니다.FAQ
Why is there no interactive playground on this page?
Why is there no interactive playground on this page?
Can I swap between the four models by changing only the model name?
Can I swap between the four models by changing only the model name?
modalities ↔ output_modalities, voice ↔ audio.output.voice, input_audio_format ↔ audio.input.format, turn_detection ↔ audio.input.turn_detection, input_audio_transcription ↔ audio.input.transcription, plus the event names response.text.delta ↔ response.output_text.delta and response.audio.delta ↔ response.output_audio.delta. See the Protocol Comparison section for the full mapping.The handshake fails outright. How do I debug it?
The handshake fails outright. How do I debug it?
wss://, not https://; 2. the endpoint includes ?model=<model-name>; 3. the Authorization: Bearer <key> header is present; 4. the key’s group includes the model (all four are in the default group; a mismatch returns 503 with “no available channel”); 5. no reverse proxy in between is stripping the Upgrade header — this is a common issue when relaying through your own gateway.Can I connect from the browser? Will my key leak?
Can I connect from the browser? Will my key leak?
Sec-WebSocket-Protocol subprotocol, so a browser WebSocket can connect directly. But that hands your key to the browser, where any visitor can read it from the network panel, so it is only suitable for local verification. In production write a backend relay: the backend holds the key and opens the connection to APIYI, and the frontend talks only to your own service.Sending 16 kHz audio to gpt-realtime-2.1 fails. Why?
Sending 16 kHz audio to gpt-realtime-2.1 fails. Why?
integer_below_min_value. The correct form is "audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000}}}. The two Model Studio models require 16 kHz instead — the two are not interchangeable.I have no microphone / audio is hard to test. What now?
I have no microphone / audio is hard to test. What now?
say and afconvert — the commands are in that section.After sending response.cancel I never receive response.done.
After sending response.cancel I never receive response.done.
response.text.done, response.content_part.done and response.output_item.done, but response.done is not delivered. Use response.output_item.done as the end-of-turn signal and add a timeout as a backstop. The session itself is unaffected and the conversation continues normally. The two Realtime GA models behave correctly here.My connection drops after about 5 minutes.
My connection drops after about 5 minutes.
session.update, for instance), or accept the disconnect and reconnect automatically. Remember to replay session.update and any required context after reconnecting.How long can a single session stay open?
How long can a single session stay open?
session.created event carries expires_at, measured at roughly 30 minutes from connect, after which you need to reconnect. On the Model Studio protocol the constraint we mainly observed is the 300-second idle disconnect. Design long conversations on the assumption that sessions expire, and plan how context carries across sessions.How do I set the voice, and why does changing it return cannot_update_voice?
How do I set the voice, and why does changing it return cannot_update_voice?
session.update: top-level voice on Model Studio, audio.output.voice on Realtime GA. Once the session has produced audio output the voice can no longer be changed — this applies to both protocols and returns cannot_update_voice. Pin it in the first frame and open a new session to switch. Also, do not send an empty string as the voice on Model Studio; it returns 400.I get no input transcription in manual commit mode.
I get no input transcription in manual commit mode.
flash model on Model Studio does not deliver the transcription completion event in manual commit mode (reproduced consistently across runs); the plus model does, and both work in VAD mode. Switch to server_vad or semantic_vad. Testing shows the transcript text lands in an undocumented field on the delta events in this case, but that field may change at any time and should not be relied upon. Note this only affects displaying what the user said in your UI — the conversation is unaffected, and the model understands and answers the audio correctly.Is image input supported? Why do I get Error append image before append audio.?
Is image input supported? Why do I get Error append image before append audio.?
input_image directly in the message. On Model Studio images are treated as video frames, so audio must be appended before any image, which is what triggers that error. In testing the working approach is to interleave image frames into the audio stream at roughly one frame per second.Is there prompt caching? How do I confirm a hit?
Is there prompt caching? How do I confirm a hit?
usage.input_token_details.cached_tokens (prefix of at least 1024 tokens, in 128-token increments). Note that APIYI currently bills cached tokens at the full text-input rate; the discount will be announced in the changelog once it is live. No cache hits were observed on the two Model Studio models.Are WebRTC, SIP or ephemeral keys (client_secrets) supported?
Are WebRTC, SIP or ephemeral keys (client_secrets) supported?
POST /v1/realtime/client_secrets and POST /v1/realtime/calls both return 404 on APIYI, and SIP is unavailable too; the single wss://api.apiyi.com/v1/realtime WebSocket endpoint is the only entry point. For browser or mobile clients, write a backend relay: the backend holds the key and opens the WebSocket, and the frontend talks only to your own service.Do the GA session fields such as reasoning.effort and noise_reduction work on gpt-realtime-2.1?
Do the GA session fields such as reasoning.effort and noise_reduction work on gpt-realtime-2.1?
session.updated: reasoning.effort (minimal / low / medium / high / xhigh, accepted by both models), audio.input.noise_reduction, audio.input.turn_detection.idle_timeout_ms, audio.input.transcription.model (including gpt-realtime-whisper), truncation, tracing, max_output_tokens and parallel_tool_calls. Field semantics follow the OpenAI reference; the gateway does not rewrite them.How do I estimate cost? Are text and audio billed separately?
How do I estimate cost? Are text and audio billed separately?
usage object on response.done reports tokens per modality (text / audio / image, separately for input and output), so cost can be attributed; the call-log detail view carries the same per-modality usage, one record per completed response.done. Audio tiers are substantially higher than text, which is why text-only is recommended during integration. For actual charges, refer to the call logs.