Category
Vision
Image and document understanding as input. Models that read what you show them: screenshots, documents, charts, photographs and interfaces. Reach for vision when the input is a picture and the output is words. If you need the opposite (), that is a different category. Resolution is the cost lever here: bigger images become more .
Coming soon
What to look for
What to look for in Vision
- 01
Documents or scenes
Document understanding (OCR, tables, forms) and natural-scene understanding are different skills. Test the one you need.
- 02
Resolution costs tokens
High-resolution image input can dominate your bill. Check how each model prices image tokens.
- 03
Pair with a text model
Many teams route vision requests to a multimodal model and everything else to a cheaper text model.
Explore nearby
Where to look next
Closest live categories today.