GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, coding, and multimodal analysis within the same workflow.
The model delivers improved performance in coding, document understanding, tool use, and instruction following. It is designed as a strong default for both general-purpose tasks and software engineering, capable of generating production-quality code, synthesizing information across multiple sources, and executing complex multi-step workflows with fewer iterations and greater token efficiency.
| $2.50 | $15.00 | $0.25 | 1.74s | 36 tps | ||
| $2.50 | $15.00 | $0.25 | 1.59s | 23 tps | ||
Not used in Standard routing:Why these endpoints are not used | ||||||
Flex | $1.25 | $7.50 | $0.125 | 12.14s | 15 tps | |
| $2.75 | $16.50 | $0.275 | -- | -- | ||
| $2.75 | $16.50 | $0.275 | 0.63s | 19 tps | ||
| $2.75 | $16.50 | $0.275 | 1.50s | 38 tps | ||
Fast | $5.00 | $30.00 | $0.50 | 1.65s | 94 tps | |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.