reasoning_content field, distinct from the final response.
Supported models
Models not listed here don’t support reasoning.
Enable thinking
For models marked opt-in in the table above, enable thinking by passingchat_template_args.
- Python
- JavaScript
- cURL
Pass
chat_template_args through extra_body since it extends the standard OpenAI API:enable_thinking.py
Control reasoning depth
Thereasoning_effort parameter controls how thoroughly the model reasons through a problem. DeepSeek V4 Pro, DeepSeek V4 Flash 0731, Inkling, Inkling Small, Kimi K3, OpenAI GPT 120B, GLM 5.2, and GLM 5.2 Fast support this parameter. Supported values vary by model:
Lower values return faster responses with less thorough reasoning; higher values reason longer and cost more output tokens.
reasoning_effort: "none" disables reasoning for every model in this table.
Inkling and Inkling Small ignore chat_template_args: {"enable_thinking": false}. Kimi K3 supports both reasoning_effort: "none" and chat_template_args: {"enable_thinking": false}.
GLM 5.2 and GLM 5.2 Fast return a 400 error for values outside their supported sets.
Some model templates also read reasoning_effort from inside chat_template_args (GLM 5.2 honors both placements). Use the top-level parameter: the API validates it and returns a 400 for invalid values, but doesn’t validate chat_template_args contents, so mistakes there fail silently.
- DeepSeek V4 Pro
- OpenAI GPT 120B
- GLM 5.2
- Python
- JavaScript
- cURL
Pass
reasoning_effort through extra_body since it extends the standard OpenAI API:reasoning_effort.py
reasoning_effort to low.
Parse the response
The model’s thinking process appears inreasoning_content, separate from the final answer in content. Both fields are returned on the message object.
- Python
- JavaScript
- cURL
Read
reasoning_content and content directly off the message object:parse_reasoning.py
Response
completion_tokens and count toward your total usage and billing.
Next steps
Model APIs overview
Supported models, pricing, and the feature support matrix
Structured outputs
Constrain reasoning models to a JSON schema