Skip to main content
2026/10/11 17:59 (UTC+8) · Docs Update · Google / OpenAI / Anthropic / ByteDance / Alibaba / Moonshot 📊 The vision understanding docs have been fully rewritten from tests on 24 models We checked OCR, counting, multi-image comparison and public image URLs on every model with the same mixed Chinese/English order image. Key findings:
  • All 22 vision models work with the OpenAI-compatible image_url format; gemini-3.8-flash, claude-sonnet-5-5 and gpt-6-sol passed every test
  • The Responses and Anthropic formats only reach some models, and Grok silently drops images on the Anthropic endpoint
  • Downscale large images first: one 9000×6000 image was billed 63,658 tokens on gpt-6-sol, and Claude rejects images over 10 MB; after resizing the long side to 2048 px every model passed
The page also adds an AI agent integration prompt and removes the outdated note that GPT only supports temperature=1.
← Back to Live Updates · 📚 Monthly archive