Focus · 焦点

GPT-5.6 Luna降价80%
OpenAI的清场时刻,Token走向免费

GPT-5.6 Luna 80% Price Cut
OpenAI's Clearance Sale, Tokens Go Free

GPT-5.6 Luna输入降至$0.20/MTok,降价80%,大模型价格战进入清场阶段,中间层公司生存空间被碾平。

GPT-5.6 Luna drops to $0.20/MTok input, an 80% cut, as the LLM price war enters market-clearance phase, crushing middle-tier players.

No.026 2026.07.31 约 10 分钟阅读 ~10 min read

7月30日,OpenAI在官方博客扔出了一颗定价炸弹:GPT-5.6 Luna降价80%,输入价格从每百万token 1美元直降至0.20美元,输出价格从6美元砍到1.20美元。中端模型Terra同步降价20%,旗舰Sol维持原价但新增2.5倍速的Fast模式。

没有发布会,没有预热,甚至没有上首页头条——一篇3000字的技术博客,配一张价格对比图,就把大模型行业的定价体系掀翻了。

Luna现在的价格是什么水平?$0.20/MTok输入、$1.20/MTok输出。作为对比,2025年夏天GPT-4o的价格是$2.50/$10,2026年初Claude Opus 4是$15/$75,一年前的前沿模型价格是Luna的75倍。而Luna在Agents' Last Exam基准上的表现超越了年初发布的Fable 5,单次任务成本低约99%

这不是正常的产品降价。这是清场。

三层定价体系:OpenAI的"智商税"分层

仔细看这次调价,OpenAI实际上搭建了一个精密的三层定价金字塔。

第一层:Luna——"不要钱"级别的工作马。$0.20/$1.20的价格意味着什么?一个每天处理1000万token的自动化客服系统,月成本从$1800降到$360;一个跑CI/CD的代码审查Agent,月成本从几千美元降到几百美元。Replit的总裁Michele Catasta说Luna是"intelligence too cheap to meter"(便宜到无法计量的智能),Blitzy的CTO说从GPT-5.4 mini切换到Luna后成本降低了87%,同时处理的上下文量是2.2倍。这个价格带,直接把所有"廉价模型"公司的生存空间碾平了。

第二层:Terra——"刚刚好"的生产力工具。$2/$12的价格,比Luna贵10倍但能力更强,定位在日常知识工作:文档分析、代码编写、数据处理。Notion的AI产品负责人说Terra在内部评测中达到了GPT-5.5的质量,但成本只有一半、速度快60%。Terra是企业客户的"默认选择"——不太贵,也不太弱。

第三层:Sol——旗舰"智商税"。价格不变,但加了Fast模式:2.5倍速度,2倍价格。Sol是给真正需要前沿能力的场景准备的:复杂推理、多步Agent任务、科研辅助。Fast模式本质上是"优先排队权"——你付双倍价格,可以插到队伍前面。这和AWS的EC2 Spot Instance/On-Demand逻辑如出一辙:基础资源极便宜,但你要快、要稳、要优先,就得多付钱。

"$0.20/MTok是什么概念?一年前沿模型的价格是$15/MTok,现在Luna用1/75的价格提供去年同期的前沿能力。这不是降价,是大模型的'摩尔定律时刻'——只不过这次的周期不是18个月,是3周。"—— Dawn Vision编辑部

自我吞噬的飞轮:GPT-5.6自己优化自己

这次降价最恐怖的细节,藏在OpenAI博客的第三段。

OpenAI明确写道:GPT-5.6 Sol在人类监督下自主重写了生产环境的推理内核,设计并运行了数百个实验来优化token生成效率,还监控训练过程在出现问题时主动干预。这些工作让模型端到端服务成本降低了20%,token生成效率提升了15%以上

翻译成人话:GPT-5.6在帮OpenAI赚更多钱的同时,还在帮OpenAI把自己变得更便宜。

这形成了一个可怕的正反馈循环:更好的模型→帮助优化推理效率→成本下降→可以降价→获得更多客户和数据→训练更好的模型→继续优化效率→继续降价。量子位的报道标题一针见血:"GPT-5.6自己优化自己实锤了,新的左脚踩右脚已经出现。"

这个飞轮一旦转起来,竞争对手会非常痛苦。因为你面对的不是一个固定的价格点,而是一个持续自我压缩成本的移动靶。今天Luna是$0.20,三个月后可能是$0.10,半年后可能是$0.05。每次降价都不是底线,而是下一轮降价的起点。

Cognition(Devin的母公司)的联合创始人Walden Yan说他们已经把Luna放进Devin Fusion里做日常编程任务的pair programmer;Dust的联合创始人Stanislas Polu说Luna比他们之前的默认模型快40%、便宜40%。这些客户证言不是白给的——它们在告诉市场:迁移成本几乎为零,降本效果立竿见影。

谁会被这轮降价碾死

每次大模型降价,都有人问同一个问题:谁会受伤?这次答案很清晰:中间层模型公司。

大模型行业的格局正在快速分化为三层:

顶层:前沿模型公司。OpenAI、Anthropic、Google DeepMind——这几家有资本烧钱、有人才、有算力、有数据,能持续打价格战。他们的竞争逻辑是"赢者通吃":用低价占领市场,用规模摊薄成本,用飞轮持续领先。

底层:开源/免费模型。DeepSeek、Llama系列、Mistral的开源模型、各种小参数模型——这些模型走的是另一条路:免费、可私有化部署、社区驱动。它们不和闭源模型拼API价格,拼的是自由度和定制化。

中间层:这轮降价的最大输家。那些"价格比OpenAI便宜一点、能力差一截"的闭源模型公司,会被Luna的$0.20直接碾过去。以前你的卖点是"GPT-4质量的70%但价格是1/5",现在OpenAI直接把价格砍到你的1/3、质量比你还好,你怎么打?

更危险的是,降价的冲击波会沿着产业链向上传导。AI应用层公司短期是受益者(成本大降),但中长期会面临更残酷的竞争——因为AI能力的门槛被拉平了,你能用GPT-5.6做的事,竞争对手也能用,而且一样便宜。护城河不再是"我接入了GPT",而是数据、工作流、行业know-how。

云厂商的AI inference业务也会承压。Azure OpenAI Service、AWS Bedrock、Google Vertex AI本质上是在转售模型能力并加价。当模型本身越来越便宜,云厂商的加价空间被压缩,他们必须靠增值服务(安全、合规、企业集成)赚钱,而不是靠倒卖token。

终局判断:Token走向免费,价值向上游迁移

把时间拉长看,GPT-5.6这次降价指向一个清晰的终局:基础模型的token价格将趋近于零,价值将从模型本身向上游(数据)和下游(工作流/Agent/行业解决方案)迁移。

这不是预测,是正在发生的事实。2023年GPT-4刚发布时,$0.03/1K input、$0.06/1K output的价格被认为"便宜";2024年底GPT-5的价格是$1.25/$5;2026年7月Luna是$0.20/$1.20。三年时间,等效能力的价格下降了超过98%。按这个速度,2027年高质量模型API可能接近免费——就像2010年代云计算从昂贵变成水电费一样。

当token不再值钱,什么值钱?第一,专有数据。你拥有别人没有的数据(医疗记录、金融交易、企业内部知识),你就能在开源模型上fine-tune出别人无法复制的能力。第二,Agent编排能力。把模型调用、工具使用、记忆管理、错误恢复串成可靠的工作流,这比选哪个模型重要100倍。第三,行业深度。理解医疗、法律、金融等行业的真实工作流程和合规要求,把AI嵌入到业务中去,这不是API调用能解决的。

OpenAI自己也明白这个逻辑。他们在推Codex、ChatGPT Work、Custom Agents、GPTs——这些都不是卖token,是卖工作流和场景。这次降价本质上是"刮骨疗毒":把token这个商品的利润让出去,换取在Agent和应用层的更大生态位。

对创业者和开发者来说,一个残酷但重要的真相是:别再做"更好的大模型"了,那是巨头的游戏。用大模型去解决真实问题,去改造真实行业,去构建真实的工作流——那些才是token免费之后仍然有价值的东西。

$0.20/MTok不是终点。当你看到这个价格的时候,下一轮降价已经在路上了。

明天见。

Sources · 参考来源

声明:本文为 Dawn Vision 基于公开信息的二次创作与独立分析,标题、观点、行文均为原创,仅供参考,不构成任何投资建议或决策依据。如有侵权请联系删除。

本文基于 Dawn Vision 认知引擎处理的 18 个源信号生成,经编辑部人工审核。素材来源:OpenAI官方博客、量子位、今日头条、搜狐科技。

相关入库笔记:GPT-5.6 · OpenAI · Luna降价80% · Token价格战 · 大模型定价 · 自我优化 · 中间层危机

On July 30, OpenAI dropped a pricing bomb on its official blog: GPT-5.6 Luna is getting an 80% price cut, with input prices plummeting from $1 to $0.20 per million tokens and output from $6 to $1.20. The mid-tier Terra is dropping 20%, while the flagship Sol holds its price but adds a 2.5x Fast mode.

No press conference, no teaser campaign, not even a homepage feature — a 3,000-word technical blog post with one price comparison chart upended the entire LLM pricing structure.

Where does Luna's new price land? $0.20/MTok input, $1.20/MTok output. For context, GPT-4o cost $2.50/$10 in summer 2025, Claude Opus 4 was $15/$75 in early 2026 — Luna is 75x cheaper than frontier models from a year ago. Yet on Agents' Last Exam, Luna outperforms Fable 5 released earlier this year at roughly 99% lower cost per task.

This isn't a normal product discount. This is market clearance.

A Three-Tier Pricing Pyramid: OpenAI's 'IQ Tax' Ladder

Look closely at this pricing adjustment, and you'll see OpenAI has erected a carefully engineered three-tier pyramid.

Tier 1: Luna — the "basically free" workhorse. What does $0.20/$1.20 mean in practice? An automated customer service system processing 10M tokens daily drops from $1,800/month to $360; a CI/CD code review Agent drops from thousands to hundreds of dollars. Replit President Michele Catasta called Luna "intelligence too cheap to meter." Blitzy's CTO reported an 87% cost reduction switching from GPT-5.4 mini to Luna while handling 2.2x more context. At this price band, every "cheap model" company's oxygen gets sucked out of the room.

Tier 2: Terra — the "good enough" productivity tool. At $2/$12 — 10x Luna's price but more capable — Terra targets daily knowledge work: document analysis, coding, data processing. Notion's AI Product lead said Terra matched GPT-5.5 quality in internal evals at half the cost and 60% less latency. Terra is the enterprise "default choice" — not too cheap to be scary, not too expensive to resist.

Tier 3: Sol — the flagship "IQ tax." Price unchanged, but add Fast mode: 2.5x speed at 2x price. Sol is for scenarios that genuinely need frontier capability: complex reasoning, multi-step agentic tasks, research assistance. Fast mode is essentially "priority boarding" — pay double, skip the line. This mirrors AWS EC2's Spot vs On-Demand logic: base resources are dirt cheap, but if you want speed, stability, and priority, you pay the premium.

"$0.20/MTok — let that sink in. A year ago, frontier models cost $15/MTok. Now Luna delivers last year's frontier capability at 1/75th the price. This isn't a price cut. It's LLMs' 'Moore's Law moment' — except the cycle isn't 18 months, it's three weeks."—— The Dawn Vision Editorial Desk

The Self-Devouring Flywheel: GPT-5.6 Optimizes Itself

The most terrifying detail of this price drop hides in the third paragraph of OpenAI's blog.

OpenAI explicitly states: under human supervision, GPT-5.6 Sol autonomously rewrote production inference kernels, designed and ran hundreds of experiments to optimize token generation efficiency, and even monitored training to intervene when problems arose. This work reduced end-to-end serving costs by 20% and boosted token-generation efficiency by 15%+.

In plain English: GPT-5.6 is helping OpenAI make more money while simultaneously making itself cheaper to run.

This creates a terrifying positive feedback loop: better model → helps optimize inference efficiency → costs drop → prices can be cut → more customers and data → train better models → continue optimizing → continue cutting prices.

Once this flywheel spins, competitors face agony. Because you're not competing against a fixed price point — you're competing against a moving target that continuously compresses its own costs. Today Luna is $0.20; in three months it could be $0.10; in six months, $0.05. Every price cut isn't the floor — it's the starting line for the next cut.

Cognition (Devin's parent company) Co-Founder Walden Yan said they've integrated Luna into Devin Fusion as a pair programmer for routine tasks; Dust Co-Founder Stanislas Polu reported Luna is 40% faster and 40% cheaper than their previous default. These customer testimonials aren't charity — they're telling the market: migration cost is near zero, cost savings are immediate.

Who Gets Crushed by This Round of Cuts

Every LLM price cut prompts the same question: who gets hurt? This time the answer is clear: middle-tier model companies.

The LLM industry is rapidly stratifying into three layers:

Top: Frontier model companies. OpenAI, Anthropic, Google DeepMind — these have the capital, talent, compute, and data to sustain price wars. Their logic is "winner takes all": use low prices to capture the market, use scale to amortize costs, use the flywheel to stay ahead.

Bottom: Open-source/free models. DeepSeek, Llama family, Mistral open models, small parameter models — these play a different game: free, privately deployable, community-driven. They don't compete with closed-source models on API pricing; they compete on freedom and customization.

Middle: the biggest losers of this round. Closed-source model companies selling "70% of GPT quality at 1/5 the price" will be steamrolled by Luna at $0.20. When your selling point was "cheaper than OpenAI" and OpenAI cuts prices below yours while offering better quality — how do you compete?

More dangerously, the shockwave travels up the stack. AI application companies benefit short-term (massive cost reduction) but face fiercer competition long-term — because AI capability is being commoditized. Whatever you can build with GPT-5.6, your competitor can too, just as cheaply. Moats will no longer be "I integrated GPT" — they'll be data, workflows, and industry know-how.

Cloud AI inference businesses also face pressure. Azure OpenAI Service, AWS Bedrock, Google Vertex AI essentially resell model capability at a markup. As models themselves become cheaper, cloud markups get compressed; they'll need to earn from value-added services (security, compliance, enterprise integration), not token arbitrage.

Verdict: Tokens Approach Free, Value Moves Upstream

Pull back the lens, and GPT-5.6's price cut points to a clear endgame: base model token prices will approach zero, and value will migrate upstream (to data) and downstream (to workflows/agents/vertical solutions).

This isn't a prediction — it's happening now. When GPT-4 launched in 2023, $0.03/1K input, $0.06/1K output was considered "cheap." By late 2024 GPT-5 was $1.25/$5. In July 2026 Luna is $0.20/$1.20. In three years, equivalent capability has dropped over 98%. At this rate, high-quality model APIs could be near-free by 2027 — just as cloud computing went from expensive to utility-billing in the 2010s.

When tokens aren't valuable anymore, what is? First, proprietary data. If you have data nobody else has (medical records, financial transactions, internal enterprise knowledge), you can fine-tune open models into capabilities no one can replicate. Second, agent orchestration. Chaining model calls, tool use, memory management, and error recovery into reliable workflows matters 100x more than which model you pick. Third, vertical depth. Understanding real workflows and compliance in healthcare, law, finance, and embedding AI into business — that's not solvable with API calls.

OpenAI understands this. They're pushing Codex, ChatGPT Work, Custom Agents, GPTs — none of these are selling tokens; they're selling workflows and scenarios. This price cut is fundamentally "jettisoning the cargo" — ceding margins on the token commodity to capture a larger position in the agent and application layer.

For founders and developers, a brutal but essential truth: stop building "better LLMs." That's the giants' game. Use LLMs to solve real problems, transform real industries, build real workflows — those are the things that still have value when tokens are free.

$0.20/MTok isn't the finish line. By the time you read this, the next price cut is already on its way.

See you tomorrow.

Sources · 参考来源

声明:本文为 Dawn Vision 基于公开信息的二次创作与独立分析,标题、观点、行文均为原创,仅供参考,不构成任何投资建议或决策依据。如有侵权请联系删除。

This article was generated by the Dawn Vision cognitive engine processing 18 source signals, with human editorial review. Sources: OpenAI Blog, QbitAI, Toutiao, Sohu Tech.

相关入库笔记:GPT-5.6 · OpenAI · Luna 80% cut · token price war · LLM pricing · self-improvement · middle-tier crisis

$0.20/MTok是什么概念?一年前沿模型的价格是$15/MTok,现在Luna用1/75的价格提供去年同期的前沿能力。这不是降价,是大模型的'摩尔定律时刻'——只不过这次的周期不是18个月,是3周。

—— Dawn Vision编辑部

$0.20/MTok — let that sink in. A year ago, frontier models cost $15/MTok. Now Luna delivers last year's frontier capability at 1/75th the price. This isn't a price cut. It's LLMs' 'Moore's Law moment' — except the cycle isn't 18 months, it's three weeks.

—— The Dawn Vision Editorial Desk
GPT-5.6 · OpenAI · Luna降价80% · Token价格战 · 大模型定价 · 自我优化飞轮 · 三层定价体系 · 中间层危机 · 价值向上游迁移
GPT-5.6 · OpenAI · Luna 80% cut · token price war · LLM pricing · self-improvement flywheel · three-tier pricing · middle-tier crisis · value migration upstream
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 18 个源信号生成,经编辑部人工审核。素材来源:OpenAI官方博客、量子位、今日头条、搜狐科技。

This article was generated by the Dawn Vision cognitive engine processing 18 source signals, with human editorial review. Sources: OpenAI Blog, QbitAI, Toutiao, Sohu Tech.