Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
> We've gone from 80% in some places to 80% in some more places.
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before.
I don't have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you'll get nowhere. And i've never seen the term 'brute-force' more abused than these LLM discussions. Basically none of the results have been brute force.
"This machine-intelligence stuff is overrated, they are just using <insert particular machine-intelligence technique here>" isn't the resounding verdict it may have sounded like when you typed it.
The plateau doesn't have to be perfectly flat, but it's not a straight line upward anymore either (kind of like our work on transistors, where we've kind of hit the bounds of speed in clock cycles but are improving on miniaturization and power efficiency)
Very true, the sharp increase in difficulty (as measured by human passrate plummeting from 1->2 and again from 2->3) gives an even more stark view of AI capabilities over time.
I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.
I agree that they're not hitting a plateau and I see it in my reserach. I had a math/code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
> sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
Those points were true at the time and most are still true now. But they aren’t predictions.
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
I was thinking the same thing in terms of running out of data a few months ago. But aren't most gains in the past year+ due to reinforcement learning in some form? Which doesn't need "fresh data" per se, as the model effectively creates the data as it goes. As long as engineers can come up with proper environments, tasks/goals, rewards, and actions, I don't really see data being a limit to model improvement in an agentic sense. Maybe as a knowledge base
The new hardware (TPU v8 and VR) are more expensive but they are significantly cheaper per flop. e.g. many multiples more performance for only 2x the price.
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
>Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
I think we'll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical/physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I'm young enough that I'm expecting with the rate that we are advancing, I will see AI / LLM analogues developed and run on a quantum computer in my lifetime.
Yep. I've made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark's Jarvis at our fingertips.
Incredible how many times I read similar comments over the years, containing 'on the cusp' and 'what a time to be alive'. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users
Not a single thing? In my house, we are using LLMs to:
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
> I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
The difference now is that they've hit the "good enough" point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
This is how I feel about it. I've stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I'm not the only one either, from comments above like
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
>The difference now is that they've hit the "good enough" point.
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
> ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
I didn't use the word LLM. I'm talking AI capability, you're focused on this or that current approach to AI. I think it's fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years.
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off
This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.
it's also why there have been so many calls for regulation and slowdowns.
Exactly what I've been doing. I don't need the all-powerful GPT-6 Math Scoopa, or Opus T-1000, just to write react, svelte and C# for me; my local Qwen3.8 is more than capable, and I can switch to Deepseek and GLM on OpenRouter when I need speed. I just pop in to read the comments on HN for the latest drama and navel gazing, then I click the Hide button and move on. Couldn't give a wooden nickel what their latest and greatest models are capable of anymore, it's just PR buzz.
There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
Probably one of the factors.
Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.
Luckily it's not a mistake as now we have access to
.
.
.
dots.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
They have also cut allowances for subscriptions in half. So even in the best case scenario it's about 2.5 times cheaper for Codex users. They just seem to have matched Claude Sonnet 5.5 *API pricing*, but from what I see online, it seems Claude Code now has a much more generous subscription allowance.
I love free market competition. We're getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence
This is great. But maybe part of the motivation is that 6-Sol wasn't as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.
Yes, obviously. They're both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.
I wish they'd list the environmental cost. My employer has an unlimited AI budget so I don't care about using Astra if it's just more profit for OpenAI. I care more if it actually uses 5x more energy.
I don't understand the point of this, why just now when it comes to llms. Why wasn't anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.
Edit: for context, just Steam alone has ~200million monthly active users.
Considering many games make use of cloud computing for online play and similar functions, they probably make up a pretty goot bit of global cloud compute capacity. Likely quite a lot less than the big AI players, but not an insignificant amount.
The energy costs of the cloud computing required for gaming are substantially less in power - not to mention overall demand - than LLMs. Come on, we're not in the same energy ballpark here.
You're not including the physical supply chain energy consumption of distributing video game equipment in this analysis. Nobody ships LLMs to big box stores and tries to sell them to consumers.
If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.
There are movements against consumerism and the environmental impacts of industry in general. Greenpeace is over half a century old.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
Because people find video games fun, though I suppose there's some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.
Yes I do. I've got a spare desktop that isn't too efficient (probably ~100W idle but annoyingly I've lost my power meter) so I don't leave it on even though I would like to use it as a server.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar/battery/nuclear.
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
You need to use the Blender MCP. There is an official plugin for this now, so the third party one can be avoided.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
I've only tried animating models in Astra-6, and I was quite impressed! It's rarely able to one-shot things perfectly, but it usually gets pretty close.
From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It's slightly worse at spatial detail but better at aesthetics and little touches.
After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.
And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
I'm not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me.
Can you give an example? For me I find that one shot prompts are pretty good it’s only when working with large codebases and complex, multi prompt workflows, that I find the real limitations of models
While I have no experience comparing this brand-new model, OpenAI themselves call it "near-Astra" intelligence. I set Astra and Opus 5.5 independently working on the same large research/coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both.
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it.
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
For my personal experience, antropic model have better user experience except for 4.7 and 4.8 though. 4.7 and 4.8 feels like expensive downgrade of 4.6 to me (I didn't know why these two should even exist)
However it's less willing to obey your instruction so it's less usable for general runtine flows.
For me Anthropic models from 4.7 to 5 including where bad and ate tokens like crazy. Task delivery was worse than GPT 5.6 and token usage was 2-3x higher.
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?
You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
yeah I use sol constantly and have done maybe $15 of spend in the past week. it's solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.
It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.
You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.
Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.
Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.
The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).
I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don't think this release in particular contradicts it too much.
Let's all boycott and move to Claude until they release 6.1 Astra. I don't like to be teased.
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
Is this due to a similar safety concern or just because it's not ready yet for one (or more) of a myriad of possible reasons?
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
They need to fix Astra first. My main issue is with GPT in general is that unless steered it goes into building AI “sloppiness”/machinery that is not “needed”.
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
> It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
I work on a very large code base (millions of LOC) and I've had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I've used can definitely get that refactor done, but not autonomously. It needs to be small slices. I've yet to see it 1 shot anything really complicated.
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.
I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
When GPT 6 Sol & Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can't trust it to do anything big alone anymore without babysitting.
Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?
In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
That's not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit.
Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: https://help.openai.com/en/articles/10306912-sharing-feedbac...
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
I'm in the same boat, I'll give 6.1 a shot but I'll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
I have had the same exact experience. I feel like I'm working with 5.3 again. It is alarming how degraded the experience has become over the last month.
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
GPT 6 Sol is obsolete after only one week! I am glad that they are not afraid to update the models more frequently. The Navier-Stokes thing revealed that it took them only a week or two to train a model more capable than Astra, and I want the pace of public releases to keep up with that.
I got a popup in my Codex just now saying "Try out 6.1 Sol!" and so I clicked the button to try it, and intriguingly, it set my model selector to "GPT-6 Astra Light" which makes me think 6.1 Sol may be in some way just a lighter/distilled version of Astra? defo interesting, not sure if I should read too much into it though. I see no option for directly selecting 6.1 Sol in my Codex Desktop UI.
So their original plan was to axe Terra, but then introduce an "Astra Light" model a week later? They had a nice lineup named for a whole 3 months, and they're already messing with it.
“OpenAI's new Pro 500 plan offers OpenAI's highest usage allowance and comes with access to its new "Ultrafast" feature — it also costs $500 per month.
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
It really does look like OpenAI is trying to gradually get rid of their subscription plans. Every week there is noticeably less usage available to them while each new model release boasts substantially cheaper API token pricing. If this continues then the two pricing models will eventually be at parity.
This is not true at all, at the most fundamental level. There is a reason why all the businesses (IT, gyms, cars, restaurants, streaming services, music, games, stores, apps, food delivery, magazines, newspapers, shaving blades, parfume, etc) are doing everything in their to get subscribers and are willing to decrease prices in order to get customers who are paying the monthly (or even better, a yearly) fee.
This is pretty typical product positioning. You want to sell to both high-end and low-end users, so you offer products at a few price points. Then it turns out that that middle is a much better fit for most users. So you start making the middle a worse fit to push most of those users into the higher tiers.
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
The 200$ plan was appealing because you got 4x usage for 2x the price.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
This is what I did. Hope it works out. The other benefit is you have a more natural method to avoid lock in. A lot of "improvements" to the agent harness I believe are attempts to build customer lock in.
I'm ok with whatever price they give out given they are not a monopoly and have competition, the lock in is minimum for me. This means they have legit reasons to send us this price plan. I don't believe they would shoot themselves in the foot when there is cut throat competition (Claude/opensource) out there.
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
Well, here is a breaking-news for you: the 20x from Claude is not a 20x on the weekly usage, it's a 20x on the 5h usage, while the weekly usage is simply double the $100 plan...
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
I'm describing what I got from 20x Codex vs Claude 5x.
Codex is just not worth the money, at least for me.
What's flipped is the value you get for each of those
"For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. " - Dario a couple weeks ago.
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
You are painting half of the picture, perhaps on purpose? The other half is this: Opus 5.5 is significantly better than both Sol 6.1 and Astra, and with the newly increased limits across the board, it is quite difficult to run out (unless you're spamming agents at Max effort). So it is a much, much better deal than OpenAI's Pro 100.
> Opus 5.5 is significantly better than (..) Sol 6.1
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
If you've used both you know the OpenAI plans don't compare to Anthropic plans _at all_. Claude code subscriptions are probably worth 4x as much in API spend compared to the same OpenAI subscription tier.
I think you have probably started using OpenAI recently -- one draw used to be that it was really, really hard to ever hit limits. If you did, you probably had usage resets available.
I think this is still true provided you're not using Astra.
The shitty thing about OpenAI's resets is that, unlike Anthropic, they also reset the limit (on the next natural weekly reset). It means that of you pushed the reset button 5 days into the week, you only get 2 days' (2/7 of weekly) worth of extra tokens.
> According to the company, existing subscribers will keep their current limits for a time, and will later receive a one-time credit to help them make the most of their new reduced allowances
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
I think that's by design - they're going to IPO soon so if they can get a significant percentage of users to switch from the $200 to the $500, they can 2.5x projected revenue.
For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can't justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it's only getting better from here.
Yes, either US AI corps reduce the cost of their top tier personal subscriptions down to what people are already paying for other expensive personal apps (e.g. Adobe), so ~$50-100, or open weights are going to eat their lunch very quickly. We're not there yet, as current hardware doesn't allow you to do things like multiple parallel agents, but we'll get there soon enough.
People said the same about $200 a month. I think the ceiling is probably much higher. Companies regularly spend 10% or more of employee cost on offices, SaaS, equipment. I could see these costs going to 10% of white collar income.
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
Yup. I got an R9700 recently for exactly this reason. Figured if I'm going to spend $2400 a year I may as well have something to show for it at the end of it.
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
What models are you running locally? Are you banking on them improving or do you think they're good enough today? 32GB of VRAM there wouldn't be close to enough to run the best local models.
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
Ultrafast uses 6x the usage. They probably realize that people will complain if the plan limits are too low. In any case, the TCO of the newer chips is supposedly lower. Hopefully everyone is on ultrafast eventually.
It's more like they released GPT 6 Sol too early because they were under pressure and now they are releasing the real version. You cannot do anything more than minor post-training in a week.
Implying they don't have like 3 or 4 "models" (different quants, post training, plain renaming) on the back burner at any point in time to do exactly that
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
Typically how long does codex take to update with the right model metadata for the release of a new model?
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
Impressive improvements, but GPT 6 Sol came out 7 days ago, and this one will behave differently. The panicked pace is becoming a liability, maybe they should have waited and released this as the 6.0 release
I don't know why anyone was saying that when Anthropic clearly knew Opus 5.5 significantly outperformed Astra at the time of Astra's launch. I think it might be a good exercise to go back and find out who called Astra a "gut punch" and lower your credence in their future claims.
I actually heard the opposite, the folks at Anthropic felt pretty confident they were ahead after the Astra release because it wasn't as good as they were expecting.
I guess they released this because GPT-6 Sol was underwhelming, they didn't even release it to ChatGPT. It was basically GPT-5.6 Terra for the price of Sol. However, who doesn't like price cuts? Astra for the fifth of the price? Wow, OpenAI have been quite generous recently, I still have not forgotten their 90% price cut with GPT-5.6 Luna, and now this? Astra was truly a milestone, and now they are offering similar "intelligence" for cheaper price. Incredible.
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
It’ll be interesting to see what happens to the economics of this business if we hit a wall on peak intelligence but keep finding cool ways to lower prices.
Do these benchmarks have any meaning anymore? And do the announcements seem less exciting now? (Not taking anything away from the advances we are making but it seems more incremental now?) The reliable way to tell if you'll like a model is reliable collage/X reviews to gauge a model's capability and then trying it out to see if you like the style.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
I fully believe that these models perform better in benchmarks versus their predecessors, but in real world usage inside of real, production codebases? They feel just as flawed as ever. I honestly have not seen any significant improvement in a few months. The last thing where I felt "wow" was `/fast` mode and Deepseek.
Wholeheartedly agree. Astra was some improvement over 5.6-sol in the sense that I'd "argue less" with it, but still frustrating and still sloppy. I'm starting to feel people are not honest about their experiences, they do very simple things or have very low standards. The biggest improvement i've seen from Astra so far is speed.
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
> I'm starting to feel people are not honest about their experiences...
I think it falls under:
1. They don't actually look at/care what the agent is producing as long as it works (not planning on maintaining/ops yet).
2. They are using 3rd party benchmarks (which is fair given how widely real-world workloads change from day-to-day, feature-to-feature, making it difficult to really know how well the models would perform).
3. They are doing greenfield work where there is no scaffolding, no existing code, no legacy code, nothing to guide the agents along. I believe in these cases, new models can possible do better from a blank slate. But in existing codebases, I feel like the agents are more likely to simply follow existing patterns and existing guidance to begin with so things are a wash and more reliant on harness and existing code hygiene.
Surprisingly (or maybe not) it matches the performance of Astra on my benchmark[1], but is much cheaper. It is also head to head with Opus 5.5 on both the price and pass rate, but edges it out slightly.
Sol is so good, honestly - the sweet spot for me. I've only ever found it stumbles when you don't give enough direction. But for idea execution - Sol is the GOAT.
I’m a bit disappointed with Sol 6.1. I suspect they didn't show the benchmarks and test results because it would have been embarrassing to reveal that their flagship model can't compete with the capabilities of Sonnet 5.5. That said, I still think this model is useful for a great many things, but it looks like Anthropic has the upper hand this time.
> At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
very surprised by the sentiment against GPT 6.0 Sol, I've been using it exclusively since release and it feels like a cheaper astra to me. admittedly I haven't tried any anthropic models in a while other than small tests since i can't use my anthropic subscription in other harnesses (like OpenAI has supported natively for a long time).
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
This is the first time I've seen praise for GPT 6.0 Sol: it's widely disparaged on Reddit and here in the HN comments too. My own experience likewise shows 6.0 making loads of silly mistakes, both for things 5.6 Sol is good at and things 5.6 Luna Xhigh is good at.
well i could certainly be in the wrong; i'm just speaking from my personal and likely flawed experience but i feel like i've noticed silly mistakes in every (llm) model that has been released (and that i've sufficiently used) and it hasn't felt like 6 Sol was much of a regression from 6 Astra (more than reported in both model cards), both of which ive very extensively.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
There was a model called Astra-Minor, found in the files a few days ago. I assume Sol 6.1 is this, as a last minute panic rename due to Sol 6 being underwhelming while Opus 5.5 turned out really strong. I can't really explain releasing Sol 6 in any other way, especially mere days ago.
I mean, when I upgraded my pipeline from terra to sol saying "it's the same price basically!" I was excited. Probably would not have felt as excited if it was just a version bump.
Not sure that's why they did it. But that was my experience.
I like it. It's cool and better than Opus, Fable, Sonnet etc.
Like who can figure out the ordering? With Luna < Terra < Sol < Astra it's obvious at a glance.
I propose for some third company to name after monsters: Cyclops < Minotaur < Ettin < Cerberus < Hydra < Kraken < Nyarlathotep (the AGI singularity stage)
Many people would get confused on the order of Opus, Fable and Mythos. In my mind, they are even from different groups; fable is a synonym forba fairytale (perhaps with songs), mythos is communicating importance and status instead of a size, while opus is the only one which in my mind communicates a big size. "What did you think about the Tolstoy's book? It was not a book, but a real opus, a tedious, incomprehensible, sluggish monumental opus."
It’s not… why do I need to translate from Latin (or wherever these come from) to understand them? Opus, Sonnet, and Haiku do the same thing, and are widely known words. Not to mention all LLMs do is generate tokens; I prefer the homage to writing over a space reference.
Terra is the "missing middle" model and had no positive brand recognition. Sol was for intelligence, Luna was for efficiency. Luna max was cheaper and smarter than terra light and sol light was better than terra max.
Was it a last-minute panic, or just OpenAI releasing an update when the had a bit more training under their belt to make 6.1-sol a whole lot better? Either way, I'm extremely pleased and will be giving this model a shot.
For most sane people, OpenAI is the way to go... A lot of usage with very good models, but you know that Anthropic is laughing all the way to the bank with Opus 5.5 being "the best" model right now... There are a ton of people (and companies) that will just refuse to use anything else than the highest benchmarking model in existence.
Did Medium not get a response, or is this a display issue?
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
I always notice this too. Getting it right seems (psychologically for me anyway) to be a big part of "a good pelican" whenever I look at these. But doesn't always seem to correlate with increasing intelligence of models (measured via benchmarks, experience with the model etc.).
It's not frontier pelican without the back leg behind the bike frame IMO.
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
noting that input:output:cached is 10:50:1 for astra and luna but 10:50:0.5 for sol. this doesn't mean a lot without "tokens per task" information but it's still interesting for there to be a "dip" like this instead of a monotonic change in one direction or the other
I was wondering why GPT-6 astra has been performing so incredibly bad on codex for the last week. This seems to be a repeating pattern, to dial the settings on the current models to the idiot setting, and then release a new model about a week later.
sol-6 is terra-6. They figure that no one was using terra and they could bring the speed and cost saving of terra distilled on astra, but rebranded as the more popular sol.
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
It's not that terrible then if that's the case. 5.6 Sol was superb in my eyes, and while I've tried a million things, there hasn't been anything that 6 Astra improved or did better than 5.6 Sol, while costing a ton of time and resources. Including research topics, where it should have excelled. So even if 6.1 Sol is no worse than 5.6 but cheaper and faster, the gutted Pro200 might still make sense.
I am yet to spend $200 on deepseek this year. Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500. Deepseek is faster, IMO intelligence difference is negligible and it so much cheaper that i no longer care about how much i use it. I never hit any daily/weekly quota or anything like that while working or tinkering. At this point i am OK with being 6 months behind the "frontier", purely on bang-for-buck basis and who cares which shadowy government gets my data.
A bit tired of spending $200 out-of-pocket for openai. What do you use as harness? (for me the harness if half of the benefit... controlling my PC, working from phone, etc.)
I have a big beefy desktop at home which I ssh (using EternalTerminal instead of raw port 22) into. I built it last summer right before the prices got very expensive. It's headless so I use a cheap Macbook Air to connect to it at home and I use Termux on my phone to continue working from everywhere.
Dude prompt your agent to set up always on remote connections via systemd… then you can drive from Claude or codex mobile apps natively over their native hookup. Works great.
I don't want to rely on Claude of Codex mobile features at all. Also none of this works when you use third party LLMs which is what the current comment tree is about.
curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git am
I have a server living in my home office, always on. I have a tmux session on it with vanilla Codex and Claude Code CLI, I can via my Ubiquiti network stack wiregaurd in to this box anywhere on the globe with just my laptop. Works super well for me. I also have some cheap shelley power plugs that I can use to cycle my PC’s power state if needed.
This is basically my setup but I'm using tailscale and zellij. I don't have any contingency plan in place for my power or home internet going down though..
It's like choosing between vim and helix. I started my career with vim in the early 2000's, customized the whole thing and had my config in a version control.
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
I've been trying to use DeepSeek V4.1 Flash more and been very impressed. My current (very rough) rule of thumb is that an Artificial Analysis score of ~40 is the crossover point for "good enough" for most of the things I need to do with coding agents.
47 is my crossover for serious things (e.g. Grok 4.7 is below the line and GPT 6 Sol is above the line). I mean, Opus 5.5 is way better, but GPT 6 Sol still gets the job done for anything that doesn't require design thinking.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
It’s easy to hit your quota. “Speed up the compilation time of this C++ codebase. Feel free to use several subagents to search through the files in parallel.” That’ll cost you about $200 for a codebase of ~1,000 files.
Subagents are like trading derivatives. You can lose as much as you want.
What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
And it's all measured in "intelligence", a completely meaningless term. For coding i'd be much more interested in how much context actually works, what the complexity of algorithms it can understand and create is, for what languages. How much it manages to follow existing structures or that is just adds ad-hoc machinery to pass the test, etc etc.
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
It's still insane that they stopped showing you all the tokens you pay for. They could inflate the billed reasoning token amount by a lot before it would raise any eyebrows.
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
This truism is intuitive to everyone but always funny to me how everyone never has any time, needs to save time, needs to hire staff workers for every mundane job and robots can't come soon enough… all so we can binge watch Game of Thrones and 90 day Fiancé.
And watch 10 hours of football on Sunday for our DraftKings bets.
> Time is money. Parallelism is very helpful optimising one to get the other.
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
You can do a code review on a "less capable" model that costs less, and the key model gets its output / summary, then you can have that model build a plan, and feed it to cheaper models. It's a more efficient approach than just running everything through Opus, and now that Sonnet is a lot better I'll probably use them more frequently, one thing to note is don't ask it to spin up endless subagents, I'd cap it to 2 or 3 at a time, otherwise, yeah you'll hit your limit extremely quickly.
The models get dumb as context fills. Subagents allow them to accomplish a task with minimal context rot. You can also use cheaper models for subagent tasks
Sadly this is true - for individual folks on the lower end of the spend spectrum.
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
Recently, drivers for a bunch of obscure hardware. Lots of c\c++, that i am ok with but not enough to make hardware drivers (i am just impatient). Just using pi agent with a few plugins.
How complex are they? Drivers vs an entire application would be a big difference in token usage, and it could also depend on the type of work being done.
A car that feels safe to be driving at 200 mph is going to feel more comfortable at 60 mph, compared to one for which 60 mph is at the very limits of its abilities. Analogies only go so far so I'm not sure there's anything to be learned from that though.
The issue is once you solve the hard problems, the lower models start messing things up that were working and reverting all fixes for the hard problems. They'll just go off and do dumb stuff.
In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.
There's a difference between "write this function for me" coding agents and "build this prototype from end-to-end". If you're doing the former, deepseek is fine. If you're doing the latter, it's not gonna work, and that's where the extra intelligence is most valuable.
This is how ChatGPT, Cursor apps are built. They spent so much money on PR stunts, but "thousands of agents" can't make an app that doesn't freeze on each keystroke. Not even talking about user-friendly ui
Early on their software was like this, the original ChatGPT desktop app was basically unusable and full of memory leaks that would tank the software. They’ve long since fixed that though
You're not gonna get something that's ready to ship, but as a first pass to get something running yes. Let's you explore far more ideas with only a few hours of agent time.
I started KeenLore (an emotive audiobook creator) that way. I gave it software specifications, languages, JSON schema definitions, container requirements, hardware configuration (8GB NVIDIA T1000 GPU, 96GB RAM), and zero user interface mockups. For the second round, I asked it to build a completely independent, re-entrant, and data isolated demo system on top of the web application. The demo application included voice generation using one of its voice designs. Here's the output:
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
This was my experience 3 months ago. I had an Android app that interacted with a Bluetooth device that I wanted to reverse engineer and build my own Linux app for it. DeepSeek was struggling really hard. Claude did it end to end after 3 or 4 prompts. To be fair, I was using a web interface for DeepSeek and the CLI for Claude; maybe that makes a large difference.
"build this prototype from end-to-end" works fine with DeekSeek V4.1 Flash, the problem occurs if you're not only building a prototype but want a finished product.
It's trivial to hit that kind of quota if you're trying to execute on major projects. Especially as you start having dozens or hundreds of subagents investigating, prototyping, and working on different things.
So I took Deepseek V4.1 Flash for a spin maybe 2 weeks ago now (before Luna 6 and Sol 6 were announced), and I racked up $100+ in about 2-3 days. It was pretty great, but it uses way more tokens (TPS is fast, but it's way more tokens per turn) than Sol 5.6 which I found to be about it's equivalent at the time (on medium or high, with DS on max). My cache rate was around 98-99%.
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
Maybe CC does something that breaks the cache? I cannot recommend Oh My Pi enough. Every default is galaxy brained, and it plays incredibly well with deepseek flash 4.1. My favorite coding harness rn for sure.
I found omp used quite a bit more tokens than my fairly basic pi setup... but most of those tokens would be cached with DS V4.1, so maybe worth if there are gains elsewhere.
For readers wondering, OpenRouter isn’t capable of caching as effectively as DeepSeek is because they will, for instance, switch inference providers in the middle of a session.
Fireworks directly. At the time they were the best value of cost, speed, ZDR. They got slower on me though, but I think they are retooling, so maybe things have or will get better again. I think fireworks is primarily for when you want to do your own training on top, which I wasn't doing.
How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.
I’m lost when I read these sort of comment chains. Free Gemini works just fine for me. Maybe it’s because I don’t use it for programming? How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race.
First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. Humanities will stake it out a little bit longer because AI isn’t human, but AI companies would be dammed if they don’t try. Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible. Not to mention all the roles like tech support and customer service. Once they have gotten as far as they think they can go, they will try to turn up the prices. However, they might struggle to do so as models are becoming a commodity. This is why they are arguing for regulation and stating that only they can tame these beasts.
I have deepseek agents doing email responses, with real tools (think running quotes, gathering info, scheduling things) and running business processes that used to be done by $35/hr administrative type people. And the capabilities are expanding every day as I learn how to build scaffolding around the model.
pi. I wasn't even going that hard. I checked the logs for Sep 18 and I did just shy of 3b input with approx 98.5% cache and 5.8m output, which cost around $35. Most of the was a Rust code review exercise with 1 driving agent and a varying number of subagents (up to 6 some times). I do the same with with Sol med/high driving and Luna x-high reviewing and get at least as much done if not more in a day, but I'd use up two x20 weekly allowances for the week. Worth noting that token cost isn't super meaningfull on it's own, because DS is super token heavy (but also great at caching) compared to Sol. (my stats show DS uses 3x the tokens as Sol)
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.
There's a bunch of skepticism in the replies but I ran over 100 tasks against DeepSeek 4.1 Flash and Sol (among others) and I can confirm, it is in fact a little smarter than Sol and a little more expensive than Luna. https://slopcop.com/power-ranking?pricing=api
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
Same experience. I often see people say how little they spend on DeepSeek v4.1 flash, but when I put 60 bucks into my account, it was gone in a few days of non-exclusive use. I'm actually curious what the difference is. I used it through pi and opencode, but the harness seemed to have no obvious impact on usage.
I spent the weekend trying Deepseek 4 Pro on a Linux porting project and it led me down a complete rabbit hole where Linux wouldn't even boot by the end of the weekend. Waste of $120. Switched back to GPT 6 on Monday and Linux is booting again and I'm making progress.
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate.
Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance.
Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff.
Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established.
Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation.
Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration.
Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers.
Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
That's an unfortunate experience. Think of v4.1 flash as actually v5.0 flash. It's night and day compared to the 4.0 flash (and 4.0 flash was unintuitively better than 4.0 pro). I would re-evaluate with v4.1 flash. I'm not saying it better than Sol or anything, but it's in the ballpark.
I just ran a huge text/image extraction grudgematch against all the current inexpensive models except gpt-5.5/5.6/6 (due to some issues with openrouter and bugs in my code) and DS4 ranked very poorly. Accuracy winner was Gemini 3.8 flash with minimax M3 and qwen 3.8 placing, and the chinese models beat the incumbent (Gemini 2.5 Flash) on cost whilst keeping like 95% of the accuracy.
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
The speed of deepseek is insane to experience after using claude code with opus for so long. Not only is the tps roughly 3x faster, but the round trip times are magnitudes faster.
I've been using Claude Code at work and OpenCode for side projects for a few months. Every OpenCode model I've tried always felt subpar compared to Claude, but good enough. But it changed with DeepSeek 4.1 Flash, I've been using it for the past few days and I've come to forget I was not using Claude, it's a really good model and it's basically free for my usage (I used it almost all the weekend and spent ~$5)
A lot of people having different pricing experience. I think it’s important to understand that caching can differ, than if the agents spend waiting on code, or consume a lot of content. It depends on how you structure you codebase and how explorable it is, how much effort you set and probably some other issues.
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
I think the only answer to this is you're just not using agents enough, because even with the very cheap pricing, it's still easy to rack up a large bill.
I feel a little salty about the plan changes. I wanted to upgrade to the $200 plan a day after it was blocked. Now it only includes half the usage unless for those that got grandfathered into the x20 usage.
Opus 5.5 is on another level, especially when it comes to mathematics implementations. You can drop it a PhD-level physical simulation (for example, a contrast-injection simulation for angiography in my case), and it just...implements it. With full-on WebGL rendering in the browser, from scratch (or using an existing library, if you prefer).
Just got access in Codex, looking forward to trying it out. Opus 5.5 has blown me away with what it's capable of doing, hopefully 6.1 will actually be a worthwhile contender.
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
6.0 Sol was literally a week ago... Basically continuous integration for model releases at this point.
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
Interesting idea, but at the same time it is just so cheap that you can just run it with API pricing. I sometimes do even if I have available usage that I'm going to cap so I'll save it for bigger models
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
The only real question that matters at this point is what do the economics look like for OpenAI. If they make good money on this with sane accounting principles then great. If this is just throwing more gasoline on the pile of burning cash to avoid losing more inference business then this bubble can’t pop soon enough.
More thought can cause the important info to leave the context or hallucinated info to be enshrined in the context and later acted upon, especially in long horizon benchmarks like DeepSWE.
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
> Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
All of a sudden getting competitive on token pricing over the past couple of releases tells me they’re about to kill subscription pricing big time. The subsidised tokens aren’t going to survive the IPOs but if they can capture baseline dev tasks at a cost competitive with open weight models through Luna then capture the frontier token spend as well they could be pretty well placed. The Jarvis bros aren’t going to be able to afford their dashboards though.
It seems pretty clear that this is a much larger model than Sol 6, and you can see this in the much lower generation times. I think this is also the main explanation for the $200 plan being cut in terms of API usage.
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
Frontier models are being used to obtain training data from users. We burn tokens teaching OpenAI how to make a cheaper model that is almost as good. I think the new $500/month pricing strategy is a significant misstep by someone who has clearly not tried Gemini 3.8 Flash or Deepseek 4.1 Flash.
hlynurd | 4 hours ago
tedsanders | 4 hours ago
thejazzman | 4 hours ago
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
pkulak | 4 hours ago
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
cmrdporcupine | 4 hours ago
Will see if this remedies things.
pkulak | 4 hours ago
prodigycorp | 4 hours ago
These moves all make sense when you take into account the enterprise market.
https://news.ycombinator.com/item?id=49889873
aaronbrethorst | 4 hours ago
t-sauer | 4 hours ago
nsingh2 | 4 hours ago
SirMaster | 4 hours ago
algoth1 | 4 hours ago
oh_no | 4 hours ago
gradus_ad | 4 hours ago
nojito | 4 hours ago
I remember when bandwidth was super expensive and now it’s dirt cheap.
vanviegen | 4 hours ago
iAMkenough | 4 hours ago
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
djfjkfkffkkf | 4 hours ago
Razengan | 4 hours ago
It's not even anything controversial..
simlevesque | 4 hours ago
necovek | 3 hours ago
Razengan | 3 hours ago
neta1337 | an hour ago
alch- | an hour ago
bogrollben | 4 hours ago
Razengan | 4 hours ago
CamperBob2 | 4 hours ago
ok123456 | 4 hours ago
mixdup | 4 hours ago
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
semiquaver | 4 hours ago
Edit: removed a comment that was uncharitable and rude, for which I apologize.
ActionHank | 4 hours ago
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
CuriouslyC | 4 hours ago
ActionHank | 3 hours ago
We've gone from 80% in some places to 80% in some more places.
CuriouslyC | 3 hours ago
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
ActionHank | 2 hours ago
mixdup | 4 hours ago
phoghed | 4 hours ago
arctic-true | 4 hours ago
famouswaffles | 4 hours ago
semiquaver | 4 hours ago
CamperBob2 | 4 hours ago
serf | 4 hours ago
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup | 4 hours ago
password54321 | 4 hours ago
delillos | 4 hours ago
password54321 | 3 hours ago
JacobAsmuth | 3 hours ago
LPisGood | 4 hours ago
theturtletalks | 4 hours ago
colechristensen | 4 hours ago
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
CuriouslyC | 4 hours ago
omalled | an hour ago
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530
xienze | 4 hours ago
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead | 4 hours ago
azan_ | 4 hours ago
People were talking about plateau for years already.
sebzim4500 | 4 hours ago
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
luma | 3 hours ago
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
dgellow | 3 hours ago
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
moosehater | 3 hours ago
JacobAsmuth | 3 hours ago
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
john_strinlai | 3 hours ago
do you think it will be exponential forever?
RobCat27 | 3 hours ago
spathi_fwiffo | an hour ago
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
trentnix | 3 hours ago
What a time to be alive.
neta1337 | an hour ago
trentnix | an hour ago
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
FiberBundle | an hour ago
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
digdugdirk | 3 hours ago
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
willchis | 3 hours ago
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
famouswaffles | 2 hours ago
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
OliveronData | 3 hours ago
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
luma | 2 hours ago
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
neta1337 | 2 hours ago
interestpiqued | 3 hours ago
chamomeal | 45 minutes ago
jorblumesea | 4 hours ago
it's also why there have been so many calls for regulation and slowdowns.
LeBit | 3 hours ago
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
nozzlegear | 3 hours ago
0cf8612b2e1e | 3 hours ago
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
jimbob45 | 4 hours ago
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
simianwords | 4 hours ago
minimaxir | 3 hours ago
eli | 11 minutes ago
amelius | 4 hours ago
skulk | 4 hours ago
amelius | 4 hours ago
aleph_minus_one | 4 hours ago
amelius | 4 hours ago
mholm | 4 hours ago
condour75 | 4 hours ago
SkyBelow | 4 hours ago
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
Mkengin | 2 hours ago
phpnode | 4 hours ago
jesse_dot_id | 4 hours ago
lxgr | 4 hours ago
dandellion | 4 hours ago
lxgr | 3 hours ago
sharpshadow | 4 hours ago
wg0 | 4 hours ago
Wheen | 4 hours ago
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
LPisGood | 4 hours ago
mckirk | 4 hours ago
system2 | 4 hours ago
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 | 4 hours ago
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
copperx | 4 hours ago
andybak | 3 hours ago
thraway3837 | an hour ago
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
jonatron | 4 hours ago
SwabbyNat74 | 4 hours ago
tjwebbnorfolk | 4 hours ago
colpabar | 4 hours ago
infamouscow | 4 hours ago
Aboutplants | 4 hours ago
scrollop | 4 hours ago
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
sockaddr | 4 hours ago
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
geeky4qwerty | 4 hours ago
pythonaut_16 | 3 hours ago
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
mattnewton | 4 hours ago
motoboi | 4 hours ago
mynameisjonny_ | 4 hours ago
agluszak | 4 hours ago
orbital-decay | 4 hours ago
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
jchw | 4 hours ago
esafak | 4 hours ago
toasty228 | 4 hours ago
copperx | 4 hours ago
copperx | 4 hours ago
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
denysvitali | 4 hours ago
blmarket | 4 hours ago
az226 | 4 hours ago
MisterMunchkin | an hour ago
cmrdporcupine | 4 hours ago
Lapalux | 4 hours ago
fraywing | 4 hours ago
Astra is a pretty impressive model. Excited to try this.
gobdovan | 3 hours ago
Tadpole9181 | 2 hours ago
gobdovan | 2 hours ago
Nevin1901 | 4 hours ago
jeffybefffy519 | 57 minutes ago
glimshe | 4 hours ago
SirMaster | 4 hours ago
scottyah | 4 hours ago
IshKebab | 4 hours ago
rs_rs_rs_rs_rs | 4 hours ago
Edit: for context, just Steam alone has ~200million monthly active users.
paulryanrogers | 4 hours ago
How many DCs are devoted solely to gaming?
lp92 | 4 hours ago
HelloMcFly | 4 hours ago
rs_rs_rs_rs_rs | 4 hours ago
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
empthought | 4 hours ago
rs_rs_rs_rs_rs | 4 hours ago
An entire planet. Just Steam alone has one or two hundres million monthly active users.
lbrito | 4 hours ago
rs_rs_rs_rs_rs | 4 hours ago
Yeah? Show me the big movements against computer gaming.
lbrito | 3 hours ago
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
JDups | 4 hours ago
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
IshKebab | 3 hours ago
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
rs_rs_rs_rs_rs | 2 hours ago
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
IshKebab | an hour ago
I dunno what you're not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.
Also gamers tend not to spin up dozens of other machines to also game for them.
otterley | 4 hours ago
lp92 | 4 hours ago
codehorses | 4 hours ago
IshKebab | 3 hours ago
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Bolwin | 4 hours ago
sergiotapia | 4 hours ago
barrenko | 4 hours ago
Starlevel004 | 4 hours ago
dcchambers | 4 hours ago
cmrdporcupine | 4 hours ago
Huge misstep releasing it.
algoth1 | 4 hours ago
cmrdporcupine | 4 hours ago
tandr | 4 hours ago
iamdelirium | 4 hours ago
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
hyperpape | 4 hours ago
squidbeak | 4 hours ago
A_D_E_P_T | 4 hours ago
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
ekun | 4 hours ago
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 | 4 hours ago
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
A_D_E_P_T | 3 hours ago
lukan | 4 hours ago
A_D_E_P_T | 3 hours ago
CuriouslyC | 4 hours ago
A_D_E_P_T | 3 hours ago
CuriouslyC | 2 hours ago
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
kroaton | an hour ago
therealdrag0 | 4 hours ago
minimaxir | 4 hours ago
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
TuxSH | 4 hours ago
bigwheels | 4 hours ago
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Infinity315 | 4 hours ago
toasty228 | 4 hours ago
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
copperx | 4 hours ago
AndrewKemendo | 4 hours ago
squidbeak | 4 hours ago
edgyquant | 4 hours ago
AndrewKemendo | 3 hours ago
colinhb | 4 hours ago
rspeele | 4 hours ago
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
agar | 3 hours ago
phoghed | 3 hours ago
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
beering | 3 hours ago
ex1fm3ta | 3 hours ago
mmis1000 | 4 hours ago
However it's less willing to obey your instruction so it's less usable for general runtine flows.
krzyk | 3 hours ago
Looks like 5.5 is the new 4.6
jauntywundrkind | 4 hours ago
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
dotancohen | 4 hours ago
peterbell_nyc | 3 hours ago
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
Starlevel004 | 3 hours ago
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
TuxSH | 4 hours ago
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
sobiolite | 3 hours ago
beering | 3 hours ago
dom96 | 4 hours ago
1 - https://bench.killswitch-lang.org
zeroonetwothree | 2 hours ago
dom96 | 2 hours ago
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
joshstrange | 4 hours ago
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
codewithcheese | 3 hours ago
redox99 | 3 hours ago
shimman | 2 hours ago
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
jorblumesea | an hour ago
Aeolun | 20 minutes ago
apitman | 3 hours ago
antonvs | 3 hours ago
ChickeNES | 2 hours ago
Foobar8568 | 2 hours ago
Marha01 | an hour ago
_davide_ | 2 hours ago
onlyrealcuzzo | 2 hours ago
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
ngruhn | 2 hours ago
sally_glance | 38 minutes ago
SyneRyder | 33 minutes ago
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
https://community.openai.com/t/why-does-codex-report-a-258-4...
AmazingTurtle | 2 hours ago
manmal | an hour ago
verdverm | 4 hours ago
crazylogger | 4 hours ago
throwitaway222 | 4 hours ago
Guess not?
murbard2 | 4 hours ago
minimaxir | 4 hours ago
nimonian | 4 hours ago
zf00002 | 4 hours ago
skybrian | 4 hours ago
alvis | 4 hours ago
slopinthebag | 4 hours ago
sehw | 4 hours ago
Aboutplants | 4 hours ago
Alifatisk | 4 hours ago
iosjunkie | 3 hours ago
mekpro | 4 hours ago
jdw64 | 4 hours ago
godwinson__4-8 | 4 hours ago
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
ColonelPhantom | 4 hours ago
godwinson__4-8 | 4 hours ago
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
wren6991 | 4 hours ago
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
vb-8448 | 4 hours ago
tandr | 4 hours ago
az226 | 4 hours ago
TuxSH | 3 hours ago
objektif | 4 hours ago
vb-8448 | 4 hours ago
tultra | 4 hours ago
formvoltron | 4 hours ago
Alifatisk | 4 hours ago
gavin_gee | 4 hours ago
nicce | 4 hours ago
thefounder | 4 hours ago
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
the_duke | 4 hours ago
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
nxc18 | 4 hours ago
user43928 | 3 hours ago
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
nxc18 | 3 hours ago
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
sigbottle | 3 hours ago
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
nxc18 | 3 hours ago
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
sigbottle | 2 hours ago
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
moshegramovsky | an hour ago
moshegramovsky | an hour ago
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.
I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.
jstummbillig | 4 hours ago
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
the_duke | 4 hours ago
nicce | 4 hours ago
copperx | 4 hours ago
Marha01 | 3 hours ago
Eridrus | 3 hours ago
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
phoghed | 3 hours ago
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
btbuildem | 3 hours ago
wkcheng | 3 hours ago
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
trentnix | 3 hours ago
chronogram | an hour ago
r0l1 | an hour ago
setnone | 3 hours ago
ozgung | 3 hours ago
bitexploder | 2 hours ago
NorthSouthNorth | 2 hours ago
sunaookami | 2 hours ago
stldev | 2 hours ago
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
moshegramovsky | 2 hours ago
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
jrflo | an hour ago
soulofmischief | an hour ago
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
beebmam | an hour ago
jsw97 | an hour ago
pampas | an hour ago
jeffybefffy519 | 58 minutes ago
epolanski | 4 hours ago
modeless | 4 hours ago
pmdr | 3 hours ago
samuelknight | 3 hours ago
mkaic | 4 hours ago
slekker | 4 hours ago
recursive | 3 hours ago
skerit | 2 hours ago
solarkraft | an hour ago
lynx97 | 4 hours ago
Aboutplants | 4 hours ago
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
surgical_fire | 4 hours ago
The only way is for prices to go up. Way up.
mrtesthah | 3 hours ago
surgical_fire | 2 hours ago
machomaster | 23 minutes ago
moregrist | 4 hours ago
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch | 4 hours ago
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
jpadkins | 2 hours ago
TomGarden | 4 hours ago
Our VC-backed subscription days are numbered
m3kw9 | 4 hours ago
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
glaslong | 4 hours ago
onlyrealcuzzo | 2 hours ago
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
honkycat | 4 hours ago
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 | 4 hours ago
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
diffuse_l | 3 hours ago
spiderice | 3 hours ago
diffuse_l | 3 hours ago
cmrdporcupine | 3 hours ago
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
enraged_camel | 3 hours ago
WinstonSmith84 | 3 hours ago
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
MCArth | 3 hours ago
nostrebored | 2 hours ago
I think this is still true provided you're not using Astra.
machomaster | 31 minutes ago
andriy_koval | an hour ago
the_duke | 2 hours ago
You have to do a lot of things in parallel.
InsideOutSanta | 2 hours ago
spiderice | 3 hours ago
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
honkycat | 56 minutes ago
torginus | 4 hours ago
adonese | 4 hours ago
scottLobster | 3 hours ago
glub | 3 hours ago
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
latentsea | 3 hours ago
glub | 3 hours ago
$500 for the old $200 is definitely a fumble.
latentsea | 3 hours ago
seizethecheese | 2 hours ago
glub | an hour ago
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
seizethecheese | an hour ago
glub | an hour ago
Now you can use it in coding harnesses that call the API.
Computer0 | 35 minutes ago
LeBit | 4 hours ago
Madmallard | 2 hours ago
girvo | an hour ago
killingtime74 | 34 minutes ago
latentsea | 3 hours ago
tripleee | 3 hours ago
latentsea | 3 hours ago
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
tripleee | 2 hours ago
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
user43928 | 3 hours ago
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
cactusplant7374 | 3 hours ago
intenex | 4 hours ago
zarzavat | 4 hours ago
toasty228 | 4 hours ago
nater5000 | 4 hours ago
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
tripleee | 3 hours ago
ychnd | 2 hours ago
sergiotapia | 4 hours ago
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
TomGarden | 4 hours ago
enraged_camel | 3 hours ago
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
atonse | 3 hours ago
The best thing is that we benefit from these constant back and forth gut punches :)
JacobAsmuth | 3 hours ago
jhonof | 2 hours ago
Alifatisk | 4 hours ago
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
arctic-true | 4 hours ago
MetaverseClub | 4 hours ago
neosat | 4 hours ago
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
CharlieDigital | 3 hours ago
paskejl | 2 hours ago
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
CharlieDigital | 2 hours ago
1. They don't actually look at/care what the agent is producing as long as it works (not planning on maintaining/ops yet).
2. They are using 3rd party benchmarks (which is fair given how widely real-world workloads change from day-to-day, feature-to-feature, making it difficult to really know how well the models would perform).
3. They are doing greenfield work where there is no scaffolding, no existing code, no legacy code, nothing to guide the agents along. I believe in these cases, new models can possible do better from a blank slate. But in existing codebases, I feel like the agents are more likely to simply follow existing patterns and existing guidance to begin with so things are a wash and more reliant on harness and existing code hygiene.
dom96 | 4 hours ago
1 - https://bench.killswitch-lang.org/
MisterMunchkin | 4 hours ago
joduplessis | 4 hours ago
itzikkatz | 4 hours ago
moinism | 4 hours ago
prometheus1992 | 4 hours ago
dangoodmanUT | 3 hours ago
ghm2180 | 3 hours ago
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
ChaseRensberger | 3 hours ago
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
unsupp0rted | 3 hours ago
ChaseRensberger | 3 hours ago
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
BrokenCogs | 3 hours ago
unsupp0rted | 3 hours ago
oh_no | 3 hours ago
varispeed | 3 hours ago
revolvingthrow | 3 hours ago
spwa4 | 3 hours ago
nsingh2 | 3 hours ago
swalsh | 3 hours ago
Not sure that's why they did it. But that was my experience.
r0b05 | 2 hours ago
outside1234 | 2 hours ago
Razengan | 2 hours ago
Like who can figure out the ordering? With Luna < Terra < Sol < Astra it's obvious at a glance.
I propose for some third company to name after monsters: Cyclops < Minotaur < Ettin < Cerberus < Hydra < Kraken < Nyarlathotep (the AGI singularity stage)
ngruhn | 2 hours ago
iamdelirium | an hour ago
machomaster | an hour ago
Quarrel | 39 minutes ago
recursive | an hour ago
It.. is?
wmichelin | an hour ago
Razengan | an hour ago
sydd | an hour ago
Razengan | an hour ago
Many of which are thiccer than our beta ass sun
Pretty badass name tbh
voiceeh | an hour ago
n8m8 | an hour ago
Razengan | 30 minutes ago
When was the last time you heard anybody say "opus" or "sonnet" before this?
Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
asa123 | 21 minutes ago
clear would be something like
piss-cheap - it’s-alright-i guess - okay-relax - ouch-my-wallet
HDThoreaun | an hour ago
jrflo | an hour ago
hawk_ | an hour ago
ttul | 3 hours ago
jauntywundrkind | 3 hours ago
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
jumploops | 2 hours ago
ylsilva | 3 hours ago
xyzzy123 | 3 hours ago
user43928 | an hour ago
So yes, it is clearly cheap in comparison.
jfrbfbreudh | 3 hours ago
Stevvo | 3 hours ago
par | 3 hours ago
sanex | 3 hours ago
simonw | 3 hours ago
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
agar | 3 hours ago
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
dankben | an hour ago
UnboundedContex | 59 minutes ago
It's not frontier pelican without the back leg behind the bike frame IMO.
simonw | 53 minutes ago
thefourthchime | 24 minutes ago
GPT 6.1 Sol — 91, ~9 min, $0.51 https://jonclegg.github.io/pacman-bakeoff/#gpt-6.1-sol
Opus 5.5 — 99, ~9 min, $2.00 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
GPT 6 Astra — 87, ~10 min, $2.42 https://jonclegg.github.io/pacman-bakeoff/#gpt-6-astra
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
Full gallery: https://jonclegg.github.io/pacman-bakeoff/
tensegrist | 3 hours ago
spwa4 | 3 hours ago
codewithcheese | 3 hours ago
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
JacobAsmuth | 3 hours ago
However, that doesn't say much. You can just run a smaller model at a larger batch size to get higher throughput but lower interactivity.
5555watch | 2 hours ago
gxs | 3 hours ago
I can handle issues much better if they are predictable even if the model makes mistakes — much more frustrating when the model is erratic
I find codex wanders off road more often and fails to see the “bigger picture” (as much as LLMs can see the bigger picture at least)
And tbh when it was first released Astral felt even worse
I’m being forced to use it right now and at the end of the day I’m making do so it’s fine, but Claude makes for a smoother experience
proxysna | 3 hours ago
zzleeper | 3 hours ago
dolebirchwood | 3 hours ago
simlevesque | 3 hours ago
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
IOT_Apprentice | 3 hours ago
simlevesque | 2 hours ago
cruffle_duffle | 2 hours ago
simlevesque | 2 hours ago
windexh8er | an hour ago
ThomasGlanzmann | 3 hours ago
proxysna | 3 hours ago
phyalow | 2 hours ago
marknutter | 2 hours ago
pimeys | an hour ago
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
kleinishere | 57 minutes ago
pimeys | 31 minutes ago
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
apitman | 3 hours ago
mchusma | 2 hours ago
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
apitman | 2 hours ago
hmontazeri | 2 hours ago
dzink | an hour ago
apitman | an hour ago
sillysaurusx | 3 hours ago
Subagents are like trading derivatives. You can lose as much as you want.
abixb | 3 hours ago
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
catigula | 2 hours ago
zer00eyz | 2 hours ago
MintsJohn | 2 hours ago
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
KeplerBoy | 2 hours ago
ldng | an hour ago
mlmonkey | an hour ago
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
kruipen | an hour ago
deadbabe | 3 hours ago
phyalow | 2 hours ago
apitman | 2 hours ago
apsurd | 2 hours ago
And watch 10 hours of football on Sunday for our DraftKings bets.
knollimar | 2 hours ago
georgemcbay | 53 minutes ago
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
FearNotDaniel | 2 hours ago
qarl | 2 hours ago
nater5000 | 2 hours ago
giancarlostoro | 2 hours ago
HDThoreaun | an hour ago
proxysna | 2 hours ago
gbacon | 2 hours ago
Excellent pithy warning.
sheepscreek | an hour ago
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
rmaxdev | 3 hours ago
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
iammrpayments | 2 hours ago
proxysna | 2 hours ago
marknutter | 2 hours ago
FailMore | 2 hours ago
nullbyte | 2 hours ago
tripleee | 2 hours ago
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
nkjoep | 2 hours ago
_benj | 2 hours ago
fragmede | an hour ago
bel8 | 2 hours ago
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
linuxftw | 2 hours ago
zozbot234 | an hour ago
jrflo | 2 hours ago
sreekanth850 | 2 hours ago
outside1234 | 2 hours ago
poilcn | 2 hours ago
marknutter | 2 hours ago
SOLAR_FIELDS | 2 hours ago
jurgenburgen | an hour ago
jrflo | 2 hours ago
thangalin | an hour ago
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
mitthrowaway2 | an hour ago
piterrro | 2 hours ago
proxysna | 2 hours ago
throooooo | 2 hours ago
logicchains | 2 hours ago
alfalfasprout | 2 hours ago
giancarlostoro | 2 hours ago
rapind | 2 hours ago
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
christophilus | 2 hours ago
Edit: others have noted the provider and harness matters. My experience is with opencode.
guluarte | 2 hours ago
taylorfinley | 2 hours ago
rapind | an hour ago
taylorfinley | 2 hours ago
faitswulff | an hour ago
Computer0 | an hour ago
roarkeful | an hour ago
pests | an hour ago
https://openrouter.ai/docs/guides/routing/provider-selection
pwython | an hour ago
swingboy | an hour ago
rapind | 33 minutes ago
pimeys | 2 hours ago
What harness you are using?
the__alchemist | an hour ago
Mistletoe | an hour ago
internetter | 41 minutes ago
drewnick | 40 minutes ago
pimeys | 30 minutes ago
rapind | an hour ago
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
ricardobeat | 28 minutes ago
jbellis | an hour ago
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
cpursley | an hour ago
jbellis | 32 minutes ago
MisterMunchkin | an hour ago
InsideOutSanta | 46 minutes ago
m3kw9 | 2 hours ago
pnw | 2 hours ago
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
pimeys | 2 hours ago
forsalebypwner | 44 minutes ago
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
rapind | 29 minutes ago
pdntspa | 2 hours ago
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
LarsDu88 | an hour ago
And no it did not deliver. A lot of it was re-done by Astra
dzink | an hour ago
UltraSane | an hour ago
ApolloFortyNine | an hour ago
thiht | an hour ago
jwpapi | an hour ago
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
holbrad | 46 minutes ago
mrbonner | 40 minutes ago
lukehandcool | 3 hours ago
AnodicElegy | 3 hours ago
https://artificialanalysis.ai/?models=gpt-5-6-luna-low%2Ccla...
According to this, at Max it's better and cheaper than 5.5 Medium, but worse than 5.5 High. At Medium, it's better and cheaper than 5.5 Low.
dzogchen | 3 hours ago
InsideOutSanta | 2 hours ago
m4rtink | 3 hours ago
aabajian | 3 hours ago
demibabs | 3 hours ago
Either way, a little ironic…
seaal | 2 hours ago
Excited to tryout Decisions API as well.
nzoschke | 2 hours ago
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
A few more thoughts here https://housecat.com/blog/introducing-housecat-agent
jdprgm | 2 hours ago
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
bayesianbot | an hour ago
poisonborz | an hour ago
machomaster | 22 minutes ago
alright2565 | 45 minutes ago
KingOfMyRoom | 2 hours ago
medler | 2 hours ago
whatifitoldyou | 2 hours ago
AmazingTurtle | 2 hours ago
it's 500$ for 25x the plus usage, thats pro (max).
this implies that the old 200$ 20x pro (more) is now more like 10x the usage of plus.
they are slashing our subscriptions in half and make it "but we're more efficient!"
user43928 | an hour ago
However, I don't know how future larger models such as the cancelled 6.1 Astra will be priced.
If the price stays high, this would indeed be quite bad for the $200 subscription..
jjcm | 2 hours ago
GPT 6.1 Sol: https://html.non.io/lcars-gpt-6.1-sol
Opus 5.5: https://html.non.io/lcars-opus-5.5
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
jjcm | 2 hours ago
cmiles8 | an hour ago
hacker_88 | an hour ago
andsoitis | an hour ago
Eventually: black hole.
jumploops | an hour ago
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
[0]https://x.com/thsottiaux/status/2105007628460109953
EugeneOZ | an hour ago
nilslindemann | an hour ago
cannonpalms | an hour ago
jrflo | an hour ago
kenzic | an hour ago
pazimzadeh | an hour ago
For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive
https://openai.com/index/introducing-gpt-6-1-sol/#deepswe
Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..
zamadatix | an hour ago
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
xkcd-sucks | 59 minutes ago
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
zerof1l | 57 minutes ago
dools | 36 minutes ago
holbrad | 35 minutes ago
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
resters | 33 minutes ago