MODELS
GPT-5.6 Luna
by OpenAI
Overview
OpenAI’s fastest, lowest-cost GPT-5.6 model for cost-sensitive, high-volume text, coding, reasoning, and tool-use workloads via the OpenAI API.
Details
GPT-5.6 Luna is an OpenAI API model in the GPT-5.6 family. Supplied OpenAI sources describe it as the fastest, lowest-cost or most cost-efficient GPT-5.6 model and position it for cost-sensitive, high-volume workloads. The model-specific API page identifies GPT-5.6 Luna as a fast, high-reasoning model with text and image input, text output, a 1,050,000-token context window, 128,000 max output tokens, and Responses API tools. OpenAI’s GPT-5.6 guidance says to use gpt-5.6-luna for efficient high-volume workloads, and the preview system card covers safety evaluations, risk designations, and safeguards for Sol, Terra, and Luna.
When to Use
Use for cost-sensitive or high-volume API workloads where OpenAI’s GPT-5.6 family is appropriate. Use for text coding reasoning or tool-use tasks where speed and lower cost are priorities. Use when OpenAI’s API guidance recommends gpt-5.6-luna for efficient high-volume workloads.
Getting Started
- Review the GPT-5.6 Luna model documentation for supported inputs
- outputs
- context window
- max output tokens
- and Responses API tool support.
- Check OpenAI’s latest-model guidance for the current recommendation to use gpt-5.6-luna in efficient high-volume workloads.
- Review OpenAI API pricing for gpt-5.6-luna before production use.
- Review the GPT-5.6 preview system card and OpenAI safety-help guidance for safeguards that may affect biological or cybersecurity requests.
- Run a small evaluation on your own text
- coding
- reasoning
- or tool-use workload.
Key Features
- •Described by OpenAI as the fastest GPT-5.6 model.
- •Described by OpenAI as the lowest-cost or most cost-efficient GPT-5.6 model.
- •Positioned for cost-sensitive
- •high-volume API workloads.
- •Supports text and image input with text output
- •according to the model-specific API page.
- •Model-specific API page lists a 1
- •050
- •000-token context window and 128
- •000 max output tokens.
- •Supports Responses API tools
- •according to the model-specific API page.
Capabilities
- •text-input
- •image-input
- •text-output
- •coding
- •reasoning
- •tool-use
- •responses-api-tools
- •large-context
Last updated Jul 30, 2026