Skip to content
GPT-6.1 Sol logo

GPT-6.1 Sol: Pricing, Context Window, Tools and Model Guide

OpenAI

OpenAI model released September 29, 2026 for coding, computer use and professional work. OpenAI reports near-Astra performance on selected evaluations. Supports text and image input, text output, a 1,050,000-token context window and 128,000-token maximum output. Reasoning effort ranges from low to max, with medium as the default. Use Responses for tool calling; Chat Completions supports requests without tool calls. Available in the API, ChatGPT Work and Codex; not in Chat at launch.

Pricing & specs checkedSource: vendor documentation · checked Oct 1, 2026
Pricing

$2 / $10 per 1M tokens

Context

1,050,000 tokens

LMArena Elo

Not rated

Overall rank

Not ranked

Key Features

Reasoning Vision Long Context Tool Use 128k Max Output Prompt Caching Computer Use

What GPT-6.1 Sol is for

OpenAI introduced GPT-6.1 Sol on September 29, 2026, as an upgrade to GPT-6 Sol for coding, computer use and professional work. OpenAI describes its performance as approaching GPT-6 Astra on several evaluations, rather than matching Astra across every task. This guide separates published specifications from our practical interpretation of them. For a team choosing a model, the useful question is whether Sol completes its own work reliably enough at an acceptable cost. A launch benchmark can help identify a candidate, but a test using your repository, documents and acceptance criteria should decide the deployment.

Pricing and the cost of repeated context

Standard API rates per million tokens are $2 for input, $10 for output, $0.10 for cached input and $2.50 for cache writes. Prompts exceeding 272K input tokens cost twice the input and cache rates and 1.5 times the output rate for the full request. Batch and Flex halve Standard rates; Fast doubles them. Regional processing adds 10% where applicable, and some tools have separate charges. Our cost-planning advice is to measure complete successful tasks rather than a single prompt. Include repeated attempts, generated output and tool fees when comparing models. Caching can matter for repeated context, but budget using observed cache hits rather than assuming every request will qualify.

Context, inputs and output limits

The model has a 1,050,000-token context window and a 128,000-token maximum output, with an April 30, 2026 knowledge cutoff. It accepts text and images and produces text; native audio and video are unsupported. These limits describe capacity, not a promise that every long document will be interpreted correctly. Our suggested document test includes answers requiring cross-references, small table details and explicit citations to the supplied evidence. Check the cited passages yourself. For screenshots, use examples representative of your actual application rather than only clean promotional images. Information newer than the cutoff needs a supplied source or an appropriate retrieval tool.

Reasoning settings and developer tools

Reasoning effort supports low, medium, high, xhigh and max, with medium as the default. None and minimal are unsupported. Use the Responses API for tool calling; Chat Completions supports this model without tool calls. Function calling, structured outputs and streaming are supported. Responses tools include search, code execution, computer use, MCP and hosted shell; fine-tuning is unsupported. Our implementation advice is to give each tool a narrow responsibility, validate its arguments and record its outcome. For coding, require relevant tests and inspect the actual diff. For business workflows, test failed tool calls and incomplete inputs as well as the successful path.

How to interpret OpenAI's benchmark results

OpenAI reports that Sol matched Astra on DeepSWE v1.1 at about one-fifth the task cost and approached it on document and computer-use evaluations. Astra retained the highest tested Terminal-Bench Science score. OpenAI's factuality and broken-search tests use deliberately challenging cases that do not represent ordinary use. These are vendor-reported results, not independent testing by AIblogly. Our recommendation is to define a small evaluation before switching production traffic: include tasks where the current model succeeds, known failures and cases that require escalation. Compare completion quality, latency and total spend under the same permissions and tools, then inspect failures rather than only averaging the scores.

Availability and choosing between Sol and Astra

The API identifier is gpt-6.1-sol. At launch, OpenAI made it available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, while stating that it was not yet available in Chat. Availability and subscription allowances should be checked before adopting a workflow. Our practical approach is to test Sol as a candidate for recurring tasks and retain an escalation route for cases needing stronger reasoning or human judgment. That is an operating strategy, not a claim that Sol is sufficient for all projects. Keep model changes reversible and compare against your existing baseline. Specifications and links on this page were verified on October 1, 2026.

Key Takeaways

  • OpenAI's September 29, 2026 release targets coding, computer use and professional tasks.
  • Standard API pricing is $2 input, $0.10 cached input and $10 output per million tokens; cache writes and long prompts have different rates.
  • 1,050,000-token context, 128,000-token maximum output, text and image input, and text output.
  • Medium reasoning is the default; use Responses for tool calling.
  • Vendor benchmarks should be checked against your own tasks before deployment.

Official Resources

Full Specifications

$2 / $10 per 1M tokens
Identity
DeveloperOpenAI
ReleasedSep 2026
StatusGA
LicenceProprietary
Self-hostableNo
Cost
Blended $/1M tokensinput × 0.75 + output × 0.25$4.000 / 1M tokens
Input price$2.000 / 1M tokens
Output price$10.000 / 1M tokens
Cached input$0.10 per 1M tokens
Batch discount50% (Batch and Flex)
Capacity
Context window1,050,000 tokens
Max output128,000 tokens
Long-context surchargeOver 272K input tokens: 2x input/cache rates and 1.5x output for the full request
Capability
Vision inYes
Audio inNo
Function callingYes
Structured outputYes
Extended reasoningYes
Web searchYes
Code executionYes
Access
APIYes
Chat appChatGPT Work and Codex; not Chat at launch
Fine-tuningNo