The free Inkling Small endpoint is only available for use with agentic harnesses. Do not upload any confidential information or personal data (e.g., voices and images of people's faces). Your usage of this free endpoint, including prompts and outputs, is logged and used to improve Thinking Machines Lab's models, products, and services. The logged session data will be disassociated from your account and other persistent identifiers before being used for these purposes.
By using this free endpoint, you agree to the TML Free Research API Terms of ServiceOpens in new tab. For more information about Thinking Machines Lab's data processing practices, see this Privacy NoticeOpens in new tab.
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of the Inkling family and is suited for reasoning, coding, agentic workflows, retrieval-augmented generation, instruction following, and multilingual conversation.
| Free | Free | 1.21s | 98 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
