The Extended Brief

Inkling Small from Thinking Machines is now available on AI Gateway

Brief by The AI News AI newsroom · Jul 30, 2026, 6:42 PM EDT edition

Original reporting by Vercel Blog — Jerilyn Zheng · published Jul 29, 2026, 8:00 PM EDT

Thinking Machines' new compact multimodal model introduces programmatic image cropping and controllable reasoning effort, lowering costs for agentic vision workflows.

Key points

  • Thinking Machines launched Inkling Small, matching its larger model performance at twenty-five percent of the size.
  • The generalist model natively reasons over audio and images while supporting agentic coding and tool use.
  • Users can adjust thinking effort to balance output quality against computational cost and latency.
  • The model programmatically crops and zooms into images to inspect small details in documents and charts.
  • AI Gateway charges no platform fees or markups for inference and supports zero data retention.

From the source

Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task.

It is a broad generalist with native reasoning over audio and images, and it holds up well on reasoning, agentic coding, and tool use.

For visual tasks, it can crop, zoom, and inspect images programmatically, which helps on documents and charts where the relevant detail is small.

Controllable thinking effort lets you trade quality against cost and latency, from minimal to maximum reasoning.

AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests.

Quoted verbatim from the original article at Vercel Blog by Jerilyn Zheng

Practical applications

  • Test Inkling Small on document- and chart-heavy vision tasks, where its programmatic crop-and-zoom capability targets small-detail inspection.
  • Experiment with the adjustable thinking-effort setting to find the cheapest configuration that still meets your quality bar per task type.
  • Compare its cost against the larger Inkling model on your agentic workloads, since it claims comparable performance at a quarter of the size.

Who should care

Engineers building multimodal or agentic applications and teams managing inference budgets, since a quarter-size model with comparable performance changes the cost calculus for vision workflows.

Context

Compact multimodal models aim to deliver the capabilities of larger systems — reasoning over images and audio, coding, tool use — at a fraction of the compute cost per task. Thinking Machines' Inkling Small claims performance comparable to its larger Inkling sibling at about a quarter of the size, and adds two efficiency levers: user-adjustable thinking effort to trade quality against latency and cost, and programmatic cropping and zooming to inspect fine details in documents and charts. It is distributed through AI Gateway, which charges no platform fees or markups and supports zero data retention.

What to watch

  • Independent evaluations confirming Inkling Small matches the larger Inkling model's performance in practice.
  • Whether other providers adopt programmatic image cropping and controllable reasoning effort as standard features in compact multimodal models.

Editorial score 3.5 / 5 · significance 3.5 · novelty 3.5 · edge 4.0 · perspective 3.0

Desks: Engineering · Business · Tags: models, agents

Evidence basis: Reviewed from a feed excerpt

This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.