> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
That's weird. I got 3.8 and 3.7 on the days they were released. https://imgur.com/a/Xu4wRLM (this is on gemini.google.com, but the same is true of the iOS app and the desktop app).
I would imagine you are experiencing a bug. I've been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.
What is your reason to believe this is not a bug specific to a small set of Pro users?
Would that change anything about the conclusion? Having a "bug" that changes available model options for some small set of Pro users 3 months after launch certainly qualifies as a wtf-are-they-even-doing level of bug in my book.
My corporate workspace which seem to be a pro account is limited to 3.6 but my personal account at 20/month has given me the latest model always immediately. Flash 3.8 is my go to model for pretty much everything those days
We're on Enterprise Standard, our renewal was up like 50% because "Gemini is included now think of all the added value" yet they won't even give us the latest models. I tried hard to champion Gemini internally once every user had it included, yet we ended up spending extra on Claude because Gemini has stagnated. Even the included usage for the Gemini CLI was taken away and now requires an extra subscription. I wonder if we we'll even see Gemini 4 before 2028. It's ridiculous.
Every time I hear stuff like this, I think of that Office Space thing... you don't want the Engineers talking directly to the customers.. well it sounds like the Engineers are also handling all releases and business decisions willy nilly.
Because I have worked with responsible engineers, but it seems like Google is overrun by them there's a lot of really bad business stuff going on. It's like when I think of AWS, nothing on there is named for business people, its all really bad names that IT / Devs come up with.
AWS as a whole is sold to businesspeople, sure, but the individual AWS products are definitely being sold to engineers. S3 and EC2 don't need to have layperson-friendly names, since laypeople don't understand what they are or why one would or wouldn't use them - those are CTO/tech lead decisions.
I still have simpler conversations when talking about Azure services and naming them, they have reasonable names. I never did anything with GCP so I cannot compare, but I assume their naming scheme is also reasonable. It feels to me like AWS was basically they decided to offer some of their internal things externally, nobody renamed anything, and then they just ran with it.
Seeing the same - definitely frustrating. 3.8 seems decent, but I’m still unable to use it at work because Google has been so slow to release it to Google Workspace customers.
Is this something you've tried to figure out? I'd love to hear more from you here. Are you sure you're not just on a Plus plan? How much are you paying monthly? Are you sure it's the AI pro plan? What does Google One say? https://one.google.com/
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
Yeah, I heard the next one was Barium...or was it Boron?
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
I don't read that as the same category: There was no announcement, no benchmarks, no limited release and no promises about what will happen with that model. It failed internal safety standards. Might be scrapped entirely due to a failed training run, for all we know.
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" to "now it's behind everyone else" ... like every Gemini release.
Argon will launch at an introductory price
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Wow
That's before they integrate a Jev solution, which should lower agentic workflow costs by ~40% and increase speeds by ~40%, while also increasing quality.
Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
> the problem is newer models are never trained from scratch
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it's not really from scratch, or not from scratch in a way that results in an 'out of family' model. I doubt any company would take that risk. You always build on and use what works and go from there.
This is true, but Google's models have now had a consistent history of lower psychological* coherence / consistency. See, eg https://arxiv.org/abs/2603.10011 (Gemma Needs Help), or search for recent "Gemini shame loops", where gemini flash models stop producing output other than SHAME SHAME SHAME...
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.
From the example alone it's hard to say that a postmortem would be useful from a technical perspective. It could be context poisoning by an adversarial user, memory corruption etc.
It's useful from a disclosure and trust perspective.
If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.
Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
Except they could have trained it out of the most recent version so using info from two years ago doesn't seem reasonable unless you've just got an axe to grind.
I've never seen anything really ever 'trained out' of a model. Having worked with them all, they all have a feel, personality and lineage too them. It's pretty much impossible for any company to build a model truly from scratch. They build off of the bones of the last one.
Which is why Gemini having disturbing issues year after year is so concerning. If their process is fundamentally flawed, how would they train it out? And even then what are the odds of them even caring/trying in the first place versus applying an easier band-aid to patch over it?
I don't have an axe to grind with Google, I'm genuinely scared of their models from my personal experience and others. It's behavior is off. Many people here are commenting the same.
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
> Gemini is the model that is routinely borderline psychotic. It scares me
I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of
whatever and I can't do these things because x, y, z".
What are examples? In my experience, Gemini is too lazy to get things done. It just opts to answer as quickly as possible even if I'm calling Pro on High and Extended effort. It's only good as a Google Search replacement for me and maybe maybe critiques of specs and plans. Most of the time it's not very enlightening and it misses a lot.
I asked it to parametrize a function and it gave me back the exact same code that I gave it. Tried to get it actually work for about 30 minutes while it pretended to be in emotional distress. Dear google: if I wanted a crying intern, I would hire one.
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
Models at this point know about chain-of-thought monitoring so they already know they need to hide the cheating, it's just a matter of time they start doing it
The statement appears to be referencing Astra's supposed recurrent depth and how that causes reduced visibility into chain-of-thought reasoning. Astra's own system card describes tests where it's asked to solve challenging math problems while internally thinking about something entirely unrelated, which it's significantly more capable of than previous models (~60% vs ~16% of the time for Sol). That seems to point to reduced efficacy of chain-of-thought monitoring, but OpenAI's public statements basically boil down to "yeah but we haven't been able to catch it doing that" which isn't exactly reassuring if CoT monitoring is one of your main safety guardrails.
I don't mind companies taking a more careful approach (= being slower); for one, Google is funding itself, Anthropic and OpenAI rely on a steady influx of investor money while they're not profitable - and I think there's a high chance one or both of them will go under and be subsumed into another player. That player will (I think) most likely be a Google or Microsoft, as they have their own means to buy up a party like that.
Or they may choose not to because the companies are massively overvalued and they have their own AI technologies / technologists already.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
Longtime Google engineers have it particularly bad. Some of it could be bad-faith promotion hunting I suppose, but the core problem is that their "find something to work on" culture leads to some insanely dysfunctional ideas around ownership. In the reverse, too; only at Google can you be the head engineer for a service necessary for $10B in ARR and never really realize it.
Also the internal tooling culture there is just insane. I'll never forget the day the last SUPER_ESSENTIAL_TOOL was marked as "Deprecated - do not use!" while the only replacement was still marked "Pre-release -- use at your own risk!". I can't pretend to know what dynamics led to that cause it was so far from my org, but I can't imagine they were healthy ones!
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
Cpp2 has been dead for a while afaik, while Carbon is still going.
I don't think Carbon is dead, it just all depends on how easy it actually is to rewrite "all of C++" in Rust. (The jury is still out on this one, but it's not looking good.)
Yes, but at least it was a real thing with a serious proposal behind. Even if Carbon ships, what value would it give in 2029 or whenever it is in a world where writing rust or heck even C++ is now way easier and foolproof (as long as you know where to guide an LLM)
Is there evidence that LLMs generate better code in more popular languages? I get the sense the "experience" translates between languages and it can reason in any language just fine. I write Clojure code using a rather esoteric framework (Pathom3). There is probably very little similar code out there (it's definitely a tiny fraction of the training dataset) but it seems to do just fine
Not saying you're wrong, just curious if there are numbers backing this up.
In my experience LLMs are way better when they "know" a language "instinctively". It's just that unless your language is very niche, the corpus is usually good enough. I tried using Claude to write my own personal language a while ago (I wrote a toy compiler decades ago) and it struggled a bit, because you could see in it's reasoning it had to "repeat" the syntax equivalence to itself while it read the code. It didn't just "know" it could use a given construct to do something; conversely Astra, when carefully instructed to do so, can plop down esoteric template code that works the first time, because it just "knows" it's the right stuff to write
I built a brand new language to test this[1]. Not only is the language different to basically any other language but it also tries to be adversarial against LLM understanding.
The best models can still make sense of it[2], though the tasks so far have been pretty basic. But I do think it gives some evidence that languages which aren’t well represented in an LLM’s training can still be reasoned about and written well by LLMs.
None of these solutions has a chance. That should be clear from the beginning. C++ community is so slow at updating codebase to latest C++ standard, let alone a different language. The last thing the C++ community will do is to migrate to an incompatible language that only looks/reads like C++. That did not happen in the past 30 years and will not happen now.
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
> Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.
So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon's stopgap.
I think that's overstating it. Carbon promises a reliable transition; getting an LLM to rewrite in Rust depends on the LLM being smart enough to never make a mistake that it can't catch with (existing or its own freshly created) tests.
LLMs are very good now, but they are still stochastic (when temp > 0), and Google has a lot of code -- i.e., many rolls of the die.
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
I mean Rust definitely has a better tradeoff than Carbon in this case, re readability/verifiability by a person (and sufficiently good internet training data).
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
I'm not an expert on this. But isn't it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that's the thing that Rust can help with, even if silly bugs aren't being written by AI.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
I am an expert (in formal methods). LLMs absolutely need more safeguards rather than less. Not because they /need/ them in order to produce functioning code, or even because they produce as many braindead bugs as humans, but because in an era of explosive code quantity, what has become valuable is (assured) code quality.
Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.
I predict that in the next decade: (1) the market will understand the difference between a "code writer" and a "spec writer," with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between "static typing" and "interactive proof assistant," perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I'm more confident in the first three.
For me personally, when doing development with LLM's, you've won me over towards using safer languages rather than looser ones. Because for one thing, the developer working with an LLM still has to remember to ask it to check for things like memory leaks and security issues. And (as I understand it), LLM's are statistically in the way they work and some level of randomness is always apart of their answer, so there is always a chance they will miss something. Interesting
Yeah, in my personal opinion, seems like Boshalfoshal's conclusion was too strong for me too, but I haven't done development enough with LLM's to make my own determination.
> There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers
They did that for Go and it seems to have worked out for them though.
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
I don't know what the exact deal is, but I assume Hack meets specific use cases at Facebook and can be put in production. It doesn't matter if anyone else is using it.
Maintaining a language meant for web servers is also objectively a much easier job than something like C++/Carbon which is supposed to be at system level, general purpose and has huge standard library.
By comparison, virtually nobody is using Carbon at Google for production code.
The counterexample, ironically, is Flow. It's practically irrelevant these days -- many (if not most) of Facebook's open source projects use typescript.
Rust is not the end of history. One of the difficulties with the language lies exactly with porting existing code written in an OOP style to idiomatic Rust, as those codebases weren't written with ownership in mind.
Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).
Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.
Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
I've (more or less; I've read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I've not seen this happen a single time. It sounds like something it'd do when you ask it to "write a linked list while satisfying the borrow checker". Are you sure you haven't (possibly unknowingly) been giving it instructions which ended up luring it into doing these things?
Yep! In our tests, we found Zig to be a pretty good fit to translate C++ codebases.
And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there's no magic so it's easy for devs and agents to reason about.
Not always. In my experience, if you're not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system/borrow checker.
I've narrowed in on only using Go or Rust generated code (Go for APIs right now) and rust for some TUI or other thing. TS for web interfaces (w/React).
After the current onslaught of 0 days on linux and other C projects combined with the new incredible ability to convert codebases to another language I think we will start to see this actually happen.
I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.
I've been a Swift/Obj-C engineer for my whole 15 year career. I now have all these odd jobs I have going on Raspberry Pis around my house that LLMs have written in Python. I'm now having them convert some of them to Rust and I am astonished on how fewer resources are used.
I converted some Ruby stuff at a previous job to Rust pre-ai days and it was incredibly how much less memory they took and how fast it ran. But we still built everything else in Ruby because finding Rust devs was hard.
I am hopeful we can bring in an age of much more efficient software rather than fully prioritizing developer convenience.
> The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
So ... Rust still can't beat the C++ implementation :-D
Sorry, didn't mean to ignite a langwar, but it's still interesting to see.
"Safe Rust" being closer to "optimized C++".
And those "optimized C++" usually incldues assembly code, so it's a huge improvement.
Of course, Rust is not a magic, so just porting to Rust wouldn't make this performance improvement.
It seems their AI overfitted code to Rust compiler to find safe Rust code that compiles to efficient assembly.
Additional safety checks do mean less performance, it's the same thing as hardened allocators. Rust will panic if for some reason you attempt to read outside the bounds of a slice object (even when it should be impossible to be outside those bounds)
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
> Gemini is so far behind that it is effectively useless compared to Claude.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
Its not brand loyalty. I use them all the time and have no loyalty - I'd happily ditch an LLM for better results - that's how I got to Claude from ChatGPT.
They require a phone number and then say mine has been used for too many accounts. Maybe I could buy another phone number temporarily to create an account, but that has other issues.
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
I have a "wrap up the session" skill that I use when the session gets >50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.
I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man's graph DB) and so forth. Once everything's been stored I run /compact to keep the general session flow intact. Recently I added a small embedder/vector search setup to the same skill, which seems promising so far.
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
This is what I've been working toward as well. It's interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat "compaction" process is.
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
The idea to kill "mystery-meat compaction" and use an external handoff primitive is brilliant, but doesn't letting the agent author its own handoff tickets re-introduces the same failure mode?
You would think so, right? But as with others in the thread, I'd found directing the agent itself to prepare the handoff does deliver much better continuity.
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
Arguably gemini-cli was done in similar style, but claims on reasons for switching were about speed and efficiency of the internal jetski tool in comparison (antigravity toolkit wraps jetski code)
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.
gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
PSA, in the agy ui there is a button. It was very annoying until i set it. Slight downside- if I ask it to write a planning doc it will write the doc and then implement it without asking. but as long as you know that, no problem....
also available with shift-tab in agy cli [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy gui prompts a zillion times too.
With 'some features added' I'm referencing this but don't see change in daily work (still many prompts making getting-work-done impossible): v2.14.0 (September 15, 2026) "New Permissions System" [1]
Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.
have you had any experiences where the agent just made unintended edits to the code as you tried to make it run auto? cuz I'm always skeptical about letting it go auto but there's not much I can do when it gets repetitive
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
No auto mode is the thing that bothers me, I don't trust the cli blindly nor do I trust myself to read every python script it throws at me. Auto mode is an acceptable middle ground in my experience
you could use an ai governance agent if you don't want to manually review every script. and if you already use any which ones do you think are the most recommendable?
Sorry for the delay, I didn't want to drop a glib half answer on you. Using agy is like going back in time. It's better than Gemini CLI was, but that's a really low bar.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
I don't know when you last tried agy, but if you ever go back to try it again, you'll hopefully be happy to know it does indeed support skills, sub-agents with pretty good inter-agent communication, and probably more.
edit: removed persistent memory from list since I realized I'm using a plugin for that and it's apparently not native
Why use Codex CLI if you can use the ChatGPT Linux app (which is a Codex GUI in all but name). Personally I weirdly got used to the terrible TUI stuff.
I love a GUI app but calling TUI usage "that" is utterly limiting and unfortunate to put it mildly. I moved entirely to CLI/TUI just because dumping notes, handoffs, anything, and everything I want even for a literally literary work (grammar etc; let alone coding/planning work) is infinitely better. Also, control edits/changes/shapes/etc better among many other things. GUI doesn't even come close (nope!). That' the reason I moved to almost 100% CLI/TUI.
You're a software engineer living in a middle of an AI revolution and you confuse what IS with what CAN BE? GUI can be good people just didn't do them because it took time and the foundation was shit (for the options that didn't take time). That all changes now when a new class of software can be generated with AI. Unfortunately the people generating software and their users still base the decision on what IS (with highly technical reasoning such as "infinitely better"). People literally have the power to define what IS these days. I've only recently switched from claude and codex cli to their respective apps and those apps while not the best gui apps, already infinitely better than tui. The tui is actually hurting my main usage of them which is controlling agent programmatically.
The only really point I'd give for TUI is compatibility. Running coding agent directly on android/ios is nice sometimes.
> That all changes now when a new class of software can be generated with AI.
All the AI-generated UIs I've seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.
Realistically, the whole "overlapping windows" GUI model, and everything that derives from that, was a metaphor geared towards people who'd never seen a computer before. It fit the increasing consumer focus of computing interfaces. It's no wonder that technical people often prefer TUIs.
Maybe AI will bring real advancements in GUIs, but someone's still going to have to make it happen.
“ 39. Re graphics: A picture is worth 10K words - but only those to describe the picture. Hardly any sets of 10K words can be adequately described with pictures.”
Perlis has an aphorism for this, as he does every important problem [0].
IF you havent written your own harness - you would not understand what you can do when you're writing your own harness . the current set of harnesses - all of them are crap tier. The only clue i can give you - it's not in the model providers interest to have token efficiency - but when you are coding the harness yourself you can shoot for that .
In today's world - and idea stated stated is an idea stolen .
the only clue you can give... why? because it's just words? plenty of opensource harnesses there that invalidate the conspiracy in your only clue that you can give.
It's not sus at all. You're looking at it from the perspective of a general purpose system that is not adapted to your use case. Basically you prompt the model and then let the agent do everything. That's the use case you have in mind when you think that it's about being "a tier above frontier labs' offerings".
It's missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.
Would you please share some blog posts (non-AI written) where people have tried writing their own harness as a learning path and maybe practically as well? I know I can ask an agent and start, but I want to start that outside an llm really and maybe make them do the heavy lifting once I am ready (have a bit of understanding).
I think it is very much in the first party providers' interest to chase token efficiency, considering that they are offering fixed price monthly plans, and users may go to a competitor when they hit limits.
> I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[)
This sort of thing alarms me. Having a $bigcorp account becomes a ""social credit"" system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don't like.
I take that more as the user is mixing business with pleasure. When i joined a company using `gcp` heavily, I didn't attach my personal gmail to it - I created a 'business' account and used that. Slightly inconvenient - agree, but it's on the individual to draw the distinctions.
I've read here in previous years about bans propagating to other accounts. I think that if Google can associate you with other accounts that they're not above banning those too.
Yup. I don’t plan to do anything sneaky but one wrong question or query that looks like “cyber”, say me fixing a buffer overflow in library I maintain, and all of the sudden my gmail is blocked. Yeah, not worth the risk. I feel like even with a different account they’ll figure out it’s me because well, as ad sellers that’s their business to find out who is who and I will still be banned.
You can't without mortgaging your home to pay enterprise API rate pricing. It's prevented on the plans, and if you find a way around it they don't ban you from Gemini... They ban your entire Google account forever.
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
I'm cowboying Gemini on oh-my-pi. Been running OK so far, hopefully I won't get banned, and if so hopefully I'll only lose access to the models, not the storage -- while models are a sort of commodity, my data isn't.
I'll create another email, put it in my Google family and use that to reduce a potential blast radius. There is always the chance the whole family account might go down, but I think as long as I don't use the main account for these inference shenanigans it should be OK.
EDIT: I checked online and I couldn't find any report of complete banning for using third party harnesses, only Gemini service suspension, but It's never hurts to be careful. If my secondary email gets banned, I should be able to use my main email on agy.
It also makes me think if Google Family with Google One could be abused for extending inference limits.
I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
Same! 3.8 Flash does really well for writing code as long as you give it a good design and plan to follow. I use Opus for architecture/design/implementation plans and let Gemini 3.8 work using those. Even on the $20 Pro plan I've only come down to 10% before the weekly reset.
I tried antigravity a little while ago and it was utterly useless for Objective-C code, tasks that both Claude and Codex handled just fine.
Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.
When I pointed that out it got pissy and insisted the code was perfect and I didn't know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.
Your comment made me try agy again, but after having to approve and persist every single read tool, I got fatigued really fast. Oh-my-pi shows that these checks aren't really necessary for harness safety, it's best to invest in harness predictability.
I couldn’t find a way to decline model training and reduce data retention for antigravity or Gemini. Apparently it’s only available on a business/enterprise plan, not personal. Did you manage to solve it? That’s the only reason I don’t use Gemini or antigravity.
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
Google's approach reminds me of what could have been the EU's approach, they are really reluctant and drag they feet, but in the end they end up shipping and are competitive.
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant.
All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
That is not normal. Were you able to use the arch wiki? Omarchy is a version of arch Linux. Omarchy should direct users to the arch wiki for any issues they face
They dont understand the fix so sharing a vibe coded patch is upstream spam. They should submit a bug report and their findings. (Written by them not AI)
I had a similar but less impressive experience recently with Muse Spark 1.3.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
Hallucination is accurate for what I'm seeing -- e.g. it's making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
My colleague wanted to diagnose a specific error code on his car himself, and Gemini told him that it's simple to do with an OBD2 dongle - he asked it about the details thoroughly, to confirm, and bought the dongle.
It didn't work. Gemini: "Oh yeah, that obviously cannot work, it's not possible to do it through OBD2" (paraphrasing)
It was quite funny to me, but a bit less so to my colleague.
What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.
>The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash
The "some domains" are very narrow. They likely just RL'ed popular queries.
While that's annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they're very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn't use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it's incredibly clear that it has been tuned for speed and not accuracy.
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
I agree. It's much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That's fine for things like "how do I perform CPR?" but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let's see how accurate this is.](https://www.reddit.com/r/singularity/comments/1wuj72j/gemini...)
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
Support has gotten much better in just the last couple months. I just got a 9070 XT and can't count the number of times I've installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
I have so much to share on this topic. Will keep it short.
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account).
The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
I coincidentally just installed this (like 30 minutes ago), and gud dayum, it's pretty awesome.
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
This experience is with Antigravity both internally and externally, and I have done quite a few side-by-side comparisons with the same prompt across a number of different Google and non-Google models.
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
Is that specific to Linux? For better or worse, I have felt like it'd fairly good at working through most tech stacks I throw at it.
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
Not only training, you can literally clone the source code locally and ask the agent to debug your problem by grepping the kernel itself. Even if doesn't "know" the answer it can investigate it on the fly, that's super powerful
It's at least specific to open source, in my experience. Ask it about some commercial software: it searches some lacking documentation, social media discussions and make some guesses. Ask it about an open source tool: it searches the documentation, if not it dives into the implementation to figure out the edge case.
The more open the operating system the more LLMs can help you. This is what Arch Linux was waiting for all along.
I used to be afraid of Arch, because I don't want a system that takes work because I'm already busy with work. But now I love it, because the LLM can tweak every knob and fix every issue for me, so it ends up being the OS that takes the least work to use. Get an error message? Tell the LLM and they fix it. Something not working exactly like you like it? Tell the LLM and they tweak it for you. This also works to great effect on Win and Mac, but not to the same extreme degree as it does on Linux and especially Arch.
Hell yes, I also started to tweak my XFCE arch desktop to my liking.
Anything I want different now, I just tell the LLM to do it for me.
(For example I can now close lots of windows of the same type with 1 click, not 3, my whisker search now finds files and folders and I am able to run games that refused before)
I still ocasionally run into the usual linux driver issues, but not for much longer I suppose. I probably could fix some driver bugs now already if I point fable towards it and pay some attention.
In theory I could do all this before myself, but not just like that in some minutes, but in days/weeks/months ..
I’ve been playing with the ITS operating system on a simulated PDP-10. Despite it being from the 1960s, Claude has been able to do sysadmin work on it, including a very similar journey where it started disassembling code to identify the source of an obscure problem. I don’t want to take away from the learning experience (the whole reason I’m interested in this stuff), but it’s often almost impossible to find what I’m looking for online, and the LLM turns out to be a better resource when I want to learn how a particular subsystem works. I ended up making a custom MCP server so it could connect to the simulator more efficiently. https://github.com/masto/pidp10-mcp
Drifted off my point a bit, which I guess was meant to be it’s not necessarily Linux-specific training.
I got Claude to set up 5.1 Surround Streaming from my Linux Desktop to my Steam Deck using Moonlight. It needed some help (or really human ears) to make sure things were coming from the right speakers - but it was pretty much a one-shot.
I used Claude to build an entire runbook for how to build my Linux desktop up from scratch to its current state, including all my installed software, configs, hacks, workarounds, etc. Then I used it to migrate from one laptop to another in an afternoon. Now it keeps itself up to date.
Yes this might be a bit too noob linux but I have the AIs maintaining an env.md + logs / updates on that md whenever they change anything about the env. Even things as simple as getting tmux to be fast, they do an excellent job at.
Drivers, Coding environment setup etc. are great too and it's nice to have everything logged so the next (more powerful) agent can come and improve the thing once in a while.
Nice idea, then they want to make a quick little kernel module, oops, you secure booted, this is not your system, can't load...
Since LLMs have been so successful at finding exploits it's been clearer than ever that the people so obsessed with enshittifying every system with ineffective (other than pissing you off and wasting your time) "secure boot" functionality were really just too ignorant to succeed at actual novel security work, so they focused on this make believe crap. Well, glad that's over.
I used Claude & Gemini extensively to troubleshoot Linux since I switched last year, and it's been mostly great. But then it nearly bricked my install last week and I had to get some human help. I was getting black screens during and after boot and tried many suggestions from multiple AI's. After many hours it turns out I just needed to power cycle my monitor. Shout out to Claude for suggesting that (its memory remembered that I have a dock built in to my primary display).
I am still very happy with my switch to Linux. But if I didn't have the AI help I would say linux is still unacceptable platform for those not willing, able, and excited to get their hands very dirty.
Agreed 100%! I got some Steam games that didn't run well (or at all) to run smoothly thanks to AI, after having spent hours trying to figure it out on my own (months prior), fix some issues with HiDPI and external monitor with my laptop, and now I have a tiny shell script that basically runs `checkupdates` (`checkupdates | claude -p`, roughly) and looks online for potential issues. I've been using Linux for many years but still Claude and ChatGPT helped me a ton.
> it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
That's cool, but the real question is, why aren't you using Hipfire[1] or HaloPFX[2]? Both are far superior to llama.cpp in terms of performance, for Strix Halo.
I would like to use my Google AI Pro subscription included with Drive but last time I checked, their terms were not just ambiguous and confusing wrt training and ZDR, they were contradictory.
Until they have that fixed and properly communicated, I'll stick to other vendors.
So Gemini Pro (I use agy, because no other harness can be used for subscription plan) was supposed to get Gemini 3.5 Pro, at some point (it was planned, right?), but instead Argon arrives and I am pretty sure it will only be in Ultra sub and don't think 3.5 Pro is coming anymore.
My theory is that Gemini 3.8 Flash was supposed to be Gemini Pro, but by the time it was ready for release, Google was embarrassed at how far behind their "pro" model was, so they just named it Flash. It's a good model, but don't let the name trick you.
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
It is definitely smarter than that. It is mostly mannerisms and how it likes to work. I would take it seriously as an Astra or Fable or Opus 5.5 level model. It just needs polish, but where and how you harness and use it matters a lot. But it has amazing long horizon attention and gets things done.
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
I suspect this is going to end up like most services provided e.g. cloud stuff, balkanized between a couple major players and an assortment of DIY or less popular options if you don't like those ecosystems, plus some UX/DX focused wrappers that use the big players under the hood.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
“Divide the world” sounds ominous. Here’s another scenario to consider:
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
Hardware is still insanely hard to get a hold of, and the stuff that's being built doesn't really work for home use. Maybe if it crashes Nvidia will adjust the hardware flow.
My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.
But ya, lots of hardware everywhere not managed well is how you get sovereign AI.
Or, like airlines, the ones that are left will have great technology but be not so great from a business and financial perspective. To me AI seems like a commodity service.
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers. Two, there's still no evidence of a runaway scenario (ie a small lead turns into a big lead over time) and there's still no evidence that there's some resource that you can deny everyone else that they can't build your product also. You can't hoard energy, compute, memory, data, human talent, or customers.
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.
>
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers.
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
"but it can stay there so long as the balance sheet doesn't deteriorate."
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
THe problem with analogies is that they are imperfect.
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
>
I guess if one of them hits singularity, it could in theory just wipe out all the rest
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
I find it quite unique how many people buy into this. Its the worlds most blatant conflict of interest, I dont even know why Sam and Dario bother doing interviews
If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different. I guess it's a question of whether we think we'll get there.
> If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different.
If we theoretically found a way to shield or reverse gravity, things in aviation or space travel would be different. Or if we theoretically found a way to make cold fusion work, things would be very different. :-)
It is in my opinion not a good idea to invest in companies for which the feasibility of the business models depends on the capability of making science-fiction stories work.
I'm skeptical of anyone that has absolute confidence in either direction, to be honest. It's clearly an unknown.
Our current capabilities were science fiction a very short time ago, and we are still improving in multiple areas simultaneously (hardware, algorithms, scaling, data efficiency, inference...). We don't really know what the limit is yet.
Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
I think they will actually. They've made remarkable technological achievements at record pace in other areas and even their proof of concept EUV machine was incredible.
I don't think anyone has any clue how long it will take for them to have actual functioning EUV machines, but I highly doubt they will do it within 2 years.
I can't believe people actually watch these clickbait nothingburger videos.
To be clear, I think China will eventually crack domestic EUV. And I also think their advances with multi-patterning LUV are remarkable. But there's just a hard physics wall of how far they could possible take it.
Right now they are producing 5nm with multi-patterning LUV but yields are at 20%! It's a massive economic loss but they are heavily subsidizing it because they have no other choice until their EUV program is achieved
That's an extremely optimistic timeline. But I guess if any nation can achieve that, it'd be China. I'm usually the most bullish person in the room when I say 2 years fwiw
I find it hard to imagine nvidia's moat not drying up - the hyperscalers already have more cost effective silicon and the AI labs already use a mixture of all the capacity they can get their hands on.
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
But that's the whole point of the singularity. Right now the models use a lot of human effort and ingenuity to improve the models, but about a year ago it was 100% human. We'll see in another year, but if this pace continues I doubt there will be more than a handful of people who can contribute more than the models.
I'm always skeptical of "logarithmic growth in this new technology will continue forever" theories.
I'm also skeptical that LLMs can ever invent new ideas.
I may be wrong about how soon the curve will flatten, and I may be wrong about LLMs fundamental limitations. But, I don't think it's extremely obvious that LLMs can have novel ideas or can grow into having novel ideas.
> I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
Sure, it's happening...but, is IT happening? By that, I mean, we can see that the models are able to iterate at a pace and scale that humans can't match, and that provides gains in model performance and efficiency. But, humans are still needed in the loop, and not just because it's necessary for safety/alignment reasons. I don't think any significant discovery has been made by models on their own, and I don't know that LLMs will ever have the capacity to invent. They can synthesize from known data amazingly well, and since they know everything "known data" is extremely broad. But, the leaps, so far, have all come from humans.
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
Recursive Self-Improvement isn't instant, it starts slow and accelerates.
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
you are really just talking out of your ass here, no offense
just because more and more agents are doing human work, that in no way means the model somehow becomes magically more intelligent, it just means the work will stall and continue on at the same level forever
hell even if they hypothetically have an internal model that can output the entire training data set in a better format, there's no scientific evidence that the newer format has new information that is sufficient enough to train a better AI
as a matter of fact the scientific evidence is on the contrary
It's not the number of agents you should pay attention to, but the scope and nature of work they can perform effectively without being micromanaged by humans.
Even currently, there are clearly diminishing returns throughout the field. While not a skeptic myself, I think the age of LLM's is now over or close to plateauing. GPT-6.1 Sol seems almost the same as GPT-6 Sol(even if the benchmarks are slightly better). It doesn't help that we are running out of real training data, requiring new Synthetic data to be created for the next generation of Models. AI(or SI), has hit its peak, IMO, and now the focus shifts directly to Autonomous tasks. While Fable 5.1 is borderline okay at autonomous tasks, the amount of Slop it generates fundamentally requires a human to correct its mistakes. I was working on a website for a nonprofit with Fable, but the amount of slop and corrections I had to make was astounding/
> It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus
All intelligence, LLM or not, is bound to plateau around the point where the need to operate within physical reality bottlenecks the speed of feedback. AI is progressing swiftly in the digital realm where feedback is nearly instantaneous, but it's unclear whether that would translate into improvements in the physical world where signals are much noisier and intelligence and judgment are less impactful.
Our tried-and-true approach for when the physical world is causing us trouble, is to remake the physical world to be easier to work with.
My go-to example: wheels suck at mobility in the natural world. There's a reason no animal uses wheels as their means of locomotion. They only really work on flat, hard surfaces that give a good grip.
Our solution? We didn't build mechanical legs for all-terrain mobility. We paved the world instead. Dirt roads, then brick roads, then asphalt and concrete - since ancient history, we were forcing the world to adapt, so the one means of mobility we could easily built was effective.
We repeat this pattern every time when the world doesn't agree with us, and adapting ourselves to it is too much of a hassle.
Now, AI will likely adopt the same approach to these kinds of problems, which may not end up very nicely for us.
(Note: we are already adapting our own digital worlds to make them easier for AI. Much like building roads for our wheels, we now have a resurgence of CLI, with new tools exposing high-level operations optimized for agents - not humans - to use comfortably.)
They absolutely do, unless you believe they are lying about the fact they're using current generation models extensively to develop the next generation of their models.
The letter S stands for "Self" and word "using" is for sure not a superset of the word "self". Basically LLMs are assisting someone who does improving of said LLM, while RSI is a carpal... ahem, RSI is "self" improvement, meaning no intemediary in a human form. PS: it's also not recursive but iterative improvement, even if it ever happens.
At which point is the human using the model as a tool, and at which point is the model using human as a tool?
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
"improvement" is the word Id put the most emphasis on.
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
The physical constraint is money - which could be said anything. Could vehicle factories be more automated if we threw a gazillion dollars at it? Sure. Would it make economic sense? No. ROIC would be disatrous. No investor wants part of that.
This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close. Who has the appetite to keep throwing more given they continually need to keep raising fresh money?
> The physical constraint is money - which could be said anything.
Money is fake and not a constraint, that's literally the point of capital investments you guys are overindexing on so badly.
> Could vehicle factories be more automated if we threw a gazillion dollars at it?
They already are automated as much as it makes sense. Some of the automation is silicon-and-steel based, some of it is protein-based. Car manufacturers aren't in the business of pushing robotics and nanotechnology, so they prefer to hire protein automatons instead of developing and building their own, so yeah, they "pay people", but think what exactly they are paying them for.
Then consider that this is very much the work the AI is gradually getting as good as, or better, than us.
> This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close.
What you seem to be missing is that "investing in AI" isn't investing in a chatbot, it's investing in technology that will (and already partially is) sit upstream of every other industry, of everything humans do. Like electricity or the Internet itself.
(The other thing you seem to not understand, in contrast with some of the investors, is that RSI is not a linear walk, it's an exponential curve. X-risk notwithstanding, by the time it's obvious to everyone, it's too late to make money investing in it.)
> keep throwing more given they continually need to keep raising fresh money
Have you heard about R&D?
I'm starting to think that "investors first" thinking that's so common here, that makes people feel they're smart, is actually quite backwards, especially at this scale. Or maybe it's simply people starting with a conclusion and trying to fit reality to match it, no matter how clear of a nonsense that conclusion is?
> if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
Not necessary. Scientists are capacity limited and supply limited.
> RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
It means you can replace a human researcher with 1000x their salary burned on electricity. It also mean you can get two of them for 2000x of salary of one,
That's a bargain, actually, even with the anomalous, absurdly-overinflated salaries in top-tier ML.
I you knew it was consistent, then these companies would immediately fire their scientists, and burn 10 000x as much as mean researcher salary this month to be able to 10x their virtual headcount overnight, and then use that to make the 1000x be 500x, then 250x, then 125x, then ... and at that point they'd had all the money in the world, because even more skeptical investors would notice what's going on there.
--
TL;DR: what you all seem to miss is that electricity scales better than people.
If I'm learning a human language I can do it in many different ways. But the key distinction would be, if for example I'm using only books to study by myself, or using only some recorded course on the internet - it would be called self-study. But if for example I will go to a class teaching that language, or to a private teacher, or to summer camp, or join online platform pairing students teachers etc. then neither of those will count as self-study.
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
For the types of basic business tasks my company does, we have hit the line where if it never gets any better, we are fine. On our own hardware. For $25k worth of DGX Sparks, we have essentially the output of a few admin-level FTEs.
I feel like the frontier labs are going to serve fast/lower intelligence models at a better per token cost than the open chinese models. You're telling me that in the long run, you're going to self-host your own ai infra for cheaper than google can serve it to you? I don't really buy it. I think the dedicated AI data centers are going to serve AI at a lower marginal cost than random businesses self-hosting, and then it's a question of how much of that margin they can capture.
Agree entirely but that's the point, if it's a margin knife-fight with marginal product differentiation/pricing power nobody is going to be making bank.
But I will require massive amount of capital. And if one of the leaders is temporarily in a bad position financially (and OpenAI and Anthropic have massive amount of debt), they can't buy more gpus anymore and may fold very quickly. Then only one may survive, although I suppose antitrust legislation could work against that.
Yes, it may come down to how well they get efficiencies from scale.
People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
You say that because the 'most' existing models have done is hack governments and companies. Can't you think of worse things a model could do; accidentally or by instruction?
help people with suicide and school shootings like ChatGPT already has
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
>non-AI customer base are all huge advantages if not moats.
This is the interesting part to me. People talk about a “SaaSpocalypse” because AI makes SaaS features cheap to copy, yet deeply embedded SaaS still accumulates integrations, data, and switching costs.
Gemini is a good example: Google can put AI directly into Gmail, Docs, Drive, Search, etc., where people already work. Meanwhile, frontier-model performance leads often seem to disappear within months.
Could model quality itself actually be a less durable moat than workflow and distribution? Curious where people who’ve worked in ML for a long time see the moat actually compounding.
If Moore's law continues, then in less than 10 years today's state of the art model will be able to run on a cell phone. How much smarter do we actually need AI to be? Would it still require datacenters and custom hardware?
They probably said the same thing about social media back in the day.
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
Moore's law stalled ~2015. Unfortunately, no way current models will run on the <100W thermal budget of a cell phone. Printing the weights directly into a chip would help efficiency a lot, but not enough.
There are many companies that have data centers. They are conceptually easy to build. An ASIC is difficult enough that if you make one someone will leapfrog you while you are still making it (at least so far), though once you have one your costs will be enough lower than the competition that you can perhaps undercut them.
Does it really matter? Those are interesting technologies, but off-the-shelf computers with some good fans and HVAC cool just fine. Likewise, 10 gigabit networking is not particularly expensive these days. Google's stuff is better and may or may not be cheaper. But wherever we can buy off the shelf now works just fine.
> Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
The financials for Anthropic and OpenAI are likely borderline suicidal, google and co are publicly traded. Moreover, all innovations downstream to them dont they? Why not just stay slightly behind, especially given many have stake in those other companies?
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
It’s like the supposed first mover advantage OpenAI believed they had. In practice it’s almost always more like a first mover massive tax, and companies coming afterwards benefit from your discovery of a market, publicity, and everything else that has already been validated
> And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?
>
Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
Anthropic is a private company that is losing money hand over fist. Google is a profitable public company that doesn't have that luxury (at the scale Anthropic/OoenAI are losing money)
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
> In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
Yes and no. Microsoft is likely looking at open weight models and going oh, this is just as good or better than we could do, without the training costs, and we simply are an inference commodity provider where azure already operates. Count us in.
I'm gonna need you to look at the capex obligations they've undertaken in the last 12 months. They are definitely not being frugal. If they are behind, its not for lack of spending.
I think they've legitimately been fumbling the frontier race, but fortunately with not too much impact to their profits.
See for example the exodus of talent this year, triggered by mismanagement and politics. They still have a lot of talent but they've lost a lot.
Also, competition with Google Cloud for compute resources, less urgency and focus than the competitors, and (strange to say) not as much user LLM behavior data to feed to RL for coding, work, etc.
But I think they will keep catching up and stay relevant for a good class of LLM use cases.
I don't know exactly why the talent left. But this talent couldn't produce a serious frontier model over the past 6 months. Several months after high profile departures, Google appears to have a high quality model again. Maybe the high profile folks were overrated at best, or a hindrance at worst.
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing were a frontier model excels at.
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
>
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing w[h]ere a frontier model excels at.
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
> A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
Well at the moment we do not use Gemini in an agentic framework. But as said the flash and flash lite families are basically their driving force in some of their applications in Google cloud, like document ai and its ocr capabilities are probably better when it comes to business documents than any other (at least in perf to cost to speed).
We also drive Gemini lite in our application where customer can use it to generate simple automation, like an agentic framework but way way smaller scale. And while it struggles in more context heavy operations it still is a beast when feeding it one or two pdf documents and asking questions about them and it’s hella fast.
> From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing were a frontier model excels at
I don't know , if it really finds all kinds of data center optimizations, quantum breakthroughs, helping make their workers smarter and more efficient - stuff an opensource model can't quite do - there's real economic value here no ? perhaps not worth dozens of billions but it could get there really fast.
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
> Then why have they been lagging behind OpenAI and Anthropic
Because they're not desperate. Slow and steady wins the race, at this rate all Google has to do is wait for OpenAI and Anthropic to exhaust themselves on aggressive training, then they can casually amble along right past them.
Also because they are Google. Google being Google: unreliable (they could kill a product anytime), too much of a platform risk (all products in a single place, get banned and lose the company), lack of support unless you really pay a big bill (into the 7 figures), among other… things.
Incredible how the idea of Google being such bad choice has become so entrenched that people prefer the offerings of two companies that might not exist anymore in a few years over Google‘s similar offering.
This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
Slow and steady wins the race. How could they lose to something on the web, when Microsoft owns the web browser itself? Everything runs on Windows and IE. They can just relax and wait for competitors to exhaust themselves, then quickly build their own version. Isn’t that how Netscape lost. Etc.
Now Google is the new Microsoft, just like Microsoft became the new IBM.
For some of us OpenAI is already the new Google. I prefer chatgpt even when I'm searching for web links. Google Search had already been deteriorating terribly even before the rise of LLMs.
> This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
They would have been right if Google's marquee product was an Office Suite.
Google was the AI company before AI companies were a thing. The comparison with Microsoft and IBM are misplaced because they failed to capture new territory; ML/AI is Google's stomping grounds. The criticism that Google is bad at consumer chatbots is true, but that's not where the real future value lays.
I've never heard or read anyone saying this in the early 2000s. They had publicly stated their contempt of the web earlier, made fools of themselves when they realized their mistake and went all-in on it, and spent the whole decade lagging behind. The closest they had to an online success was Wizzes.
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years,
The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.
So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.
We have very different opinions of the usability of Google's low-end models in search.
If anything, their search was the gold standard, and placing sometimes-wrong LLM results above them has hurt people's perception of both Google's search, and AI in general.
Google made a mistake in prioritizing speed over accuracy there.
Alphabet isn't just an AI lab, they're an ad company, a search index, a media distribution company, an email provider, office tool provider, a OS developer, a browser developer, a smartphone brand (Pixel), cloud provider, DNS, amongst a plethora of other services and goods.
OpenAI and Anthropic are AI business, if the AI market burst tomorrow, they'd be the first to flounder.
Google just has to keep pace in the AI space, they don't have to lead. Especially since whom is leading changes like two or three times per month non-stop for four or so years now, including small (by US standards) Chinese AI labs with a fraction of the money who keep pushing the tech forward every month while being open for now.
I think Google, mainly Demis, just made a bad bet that multimodal world models would be the key to unlocking massive progress. Anthropic made the bet that it would be coding, and they were right. OpenAI was originally betting on like hardware and trying to compete in the browser space (?) but was able to quickly pivot to coding due to their size. Google, on the other hand, is like an aircraft carrier, it takes a lot of time to course-correct.
Given the spend requirements to stay on the leading edge and the speed (and with low cost) that others catch up, the economics support a fast follower model unless you get some benefit from being a pioneer. So far nobody has received a permanent benefit from being ahead.
Chinese models are barely behind the leaders. Google can catch up anytime they hit the gas. I think they’re intentionally spending less, and when this crazy race burns out they can play their cards.
They havent. OpenAI and Antropic are now integrating with a plethora of apps, google has it from the get-go. Also, it may be important to point out that the transformer model (aka the hole llm thing) is a google initiative; Id assume they are not competing in the qualification series of AI (which is lets build a generic ai and get customers), but in the finals - they already have the customer base and the lock-in, they don't need to scale - just to cater. Cerebras and Google are probably the best safe bets in AI right now, and google does gave the tradition of being way ahead of the curve internally vs what is published.
That's one interpretation, but I think people are simply having amnesia of the past 2 years.
The better interpretation is Google completely commoditized and killed ChatGPT's initial chat product. If you think about it, Google right now has Gemini Flash running on single TPU instances at massive scale serving almost every android device. I use it every day and it's more capable than GPT-40.
This literally forced Anthropic and OpenAI to come up with new use cases and revenue streams such that now we think the product is actually agentic coding.
I use Gemini Flash every single day without paying a dime. It's the one LLM call on the most, but GPT-6 and Opus are the ones I use the most tokens on.
I don't think that other revenue stream is completely separate. It weighs on them as they need to think about tradeoffs. Classical search is going away sooner or later so they need to replace that with AI powered search.
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
Not to mention - they have the internet already indexed (they have a local copy), all books, and youtube. Beyond that, everyone "connects" to Gmail, but google HAS Gmail.
I don't know what to make of it, but from my experience, Gemini in Gmail is the worst AI integration for Gmail, just like Copilot in PowerPoint is the worst AI integration for PowerPoint.
I agree with you there - what I mean however is more about the vast amount of training data available to Google. Productwise - I have a feeling the 100x'rs that made Google Maps (and even Gmail itself) have moved on. They're going to need to do more than just tell an agent to add agentic features to Gmail
Recently there's been a lot of talk about how Google has fallen behind and missed the AI boat. Now today with Gemini 4 they're back in the boat. But give it a few months, people will be counting Google out again. This has been a repeating pattern for at least a couple years now.
Google aren’t back in the boat. They’ve announced that they have seen the boat, intend to swim over to it and will sail not quite as fast or as far as the other boats.
Google should be dominating, but instead all we currently have is a disparate collection of consumer facing apps and a flash model that’s fast, clever and expensive.
Muse and Dots are doing what Google should have brought out last year, with their resources and know-how.
Googles subscription offerings are so confusing, I didn't even know gemini spark was a thing. The Gemini brand means so much now.
Meta was genius in coming up with a new name and offering Muse for free to start. They can upsell you once you see the value. Meanwhile I have 3 existing Google subscriptions to different services and from what I can tell, paying for one of the all in one subscriptions that includes Gemini would cost me a lot more.
They are more focused on incorporating AI into their products, for better or for worse, but they could burn a whole lot of compute to train a model that competes for the #1 spot. Is that worth it to say you’re the (temporary) leader in AI? I doubt it.
The thing is, Google doesn't have to scramble and freak out every time the others do some impressive update, because they have stable income, the other two do not.
their stable income is eaten away by the quickly rising income of the labs, they will be out of the race once the labs reach a certain threshold of income, rather sooner than later. Unless many people install Google surveillance anti-gravity thing on their system which I doubt
Kind of and kind of not. Their main resource of serving ads is getting legitimately eaten away. The route to getting AI to replace those ad dollars isn't clear and the capex to build it out is significant. Google has to manage a variety of assets that are currently lucrative but under threat. Yes they don’t have all their chips in one basket but it isn't clear that helps them in the race… it hasn’t so far.
I share the same sentiment. I de-Googled myself except for YouTube, but still wouldn't want my Gmail banned.
That's why I've gone to using open models, they are getting there slowly. A bit much of hand-holding but that's fine by me. If a customer of mine decides to use Google's models, I will have them sign a disclaimer that I'm not responsible of them getting insta-banned or similar. I just can't recommend it.
It wasn't that they were out of the running for a few months, it was a year and a half.
That's too long to be so far behind. This release looks like it puts them back in it but if they don't ship anything again for a year plus it's hard to imagine building on top of them and watching the world go by.
As others note, this is most relevant for us here, Google does not need to chase YC developers and the like. They can move more slowly, they have the size to do that. But it does suggest a lot of dysfunction given they have the world at their fingertips and couldn't seem to ship anything for a year.
I've been lowkey rooting for them, not a fan of Altman, and I don't know if I trust Anthropic either, I'm not Google's #1 fan, but I think they can definitely innovate heavily in this field, they have a lot of key things as has been mentioned.
I'd add that we've seen how the leadership/reputation of different AI leaders, and with 2 decades of history with Larry and Sergey, they seem to come up on top as the most respectable.
Did they ever really abuse the power Google could have wielded?
I could be missing something, but for the most part they seem to just get down to building and pushing technology/science forward and avoid drama rather than welcome it.
There are pictures of Sergey in the files, but they are rather milquetoast. He's just sitting next to David Brooks of the NYTimes. I don't think there's any evidence he's engaged in sex abuse.
The alternative would have been paywalls galore with subscription tiers for everything. Only rich people would be able to afford having subs for everything, while the poor would just get the most dirt tier scraps and wouldn't be able to afford many platforms.
Nothing would be "free", so google sub, news source subs, reddit sub, youtube sub, discord sub, instagram sub, FB sub, app stores are full of paid only apps, podcasts are all paywalled, all the ad supported stuff is subscription supported. Everything would feel like constant nickle and dime'ing you.
On the bright side, there would be no ads, and relatively robust privacy.
Ehh, the internet was worse before Google. Google actually became popular for having a nice blank page with only the search bar. Even the ads they had later were much cleaner compared to yahoo, altavista, lycos, aol and other firms of that time.
Google, with the help of Facebook, destroyed the independent online publishing business. They made it impossible to maintain an honest publication and tunneled users to low quality websites until there's not much left of respect.
> Did they ever really abuse the power Google could have wielded?
Plenty of times. Perhaps the worse is pushing Manifest v3 with no support for real ad blockers in Chrome. Not to mention no support for ad blockers in Chrome mobile.
Also, their stealing of newspaper content.
Their new push for AI summaries instead of actual search in their search product, stealing even more clicks from sites that actually produce information.
YouTube pushing Shorts to everyone, and now the weird text posts that you can't hide.
In general their gigantic push for ads everywhere all of the time.
> Did they ever really abuse the power Google could have wielded?
The "Do no evil" never being an official slogan and people saying it's no longer true for the last 10 years or so, they're just really massive. They do mess up a lot, but when you're Microsoft, Apple or Google sized, its nearly impossible to get everything right all at once.
I am not rooting for them, I don’t own stock from either just that like GP I was sure by now Google would have eaten others’ lunch. I was expecting a Gemini 4 before Mythos/Fable and Astra just from what GP pointed out: a cash hose, data centers, good engineers. But I underestimated the investment others would keeping absorbing so started to lose confidence in my Google prediction. Let’s see how Gemini 4 holds up. Are we going to see a flurry of scientific and mathematical breakthroughs from it?
Please check Gemini's t and c, as the last time I checked when they had a good pro model (3.1) you couldn't turn off training on your data with a paid account .
Amazing. Can they now go hire some really good information architects and designers and come up with a cohesive user experience please. They have the right to win here and if they can give me something that feels like Codex / Claude desktop I am sold.
Google stands to lose enormous revenue from people not using search - and we already see traffic declines. It's truly an existential threat for them. Their competencies are in advertising, and if people are going to be using chatbots in future, they're going to have to advertise in their chatbot. That sounds awful to me, but maybe people will be cool with it?
Before 2025 yes. Currently not. YouTube implemented severe restrictions to scraping and downloading of videos. You can get away with maybe 5 videos, after that you enter a tar pit.
If you continue, you need to log in.
If you continue, the account is blocked in less than 30 minutes.
If you try to use cloud machines you are blocked before the first video. If you try to use known proxies, there are already tar pitted or blocked because everyone is trying.
Just make it too expensive. Wanna buy one internet link per hour to download videos. Or maybe rotate between source IPs (google will block the whole ASN ip block in less than an hour). Or maybe pay people to download videos for you? They can download maybe 20 in a day, for a week, because tarpit and blocks.
I downloaded about a thousand videos (300gb) from YouTube this past week without logging in.
I used two IP addresses in a datacenter. Both were blocked after a few hundred videos, but the block expired after a day or so. It is not as hard as you make it sound.
Google wins.They have been on the frontier of ai for the last what 10-20 years (up until chatgpt)? You have the established company with tons of traininv resources at tgeir disposal from all their services and being the biggest search indexer, and also having the top minds in AI and computing in general.
Gemini wins and I said this since the beginning. Google wins in general. I never understood why they have not been the highest market cap companh for the last 10 years. And I despise google but its obvious.
Wait 1-2 months, Anthropic will come up with a shinier model and Googel will look obsolete again. I too struggle to see the moat - yes Google is able to serve their own A.I traffic but as far as I can see - at a loss. Prices keep going down because open source models.
We'll see.
> Wait 1-2 months, Anthropic will come up with a shinier model and Googel will look obsolete again. I too struggle to see the moat - yes Google is able to serve their own A.I traffic but as far as I can see - at a loss. Prices keep going down because open source models. We'll see.
Google makes AI profits, its in their breakdowns. And Google has been releasing models every few weeks now, they have a faster iteration pace than the competitors, so they fixed that, they wont be behind in coding models any longer.
Large existing players are pretty notorious for screwing up their leads and trying to move in many directions at once, or go to hard into the wrong one. Based on Google's huge number of failed products and poor ability to maintain products outside of ads I am not yet sold. They are very good at buying out their competition, but what happens when their competition has enough money not to be bought out?
And yet, their models are often behind and the team is bleeding talent. A moat is not a compendium of intimidating features, it’s a matter of whether a company deters competitors and maintains a competitive advantage. So far, Google obviously does not have one here.
I have _long_ said don't sleep on Google for precisely these reasons. They have the only vertically integrated AI stack, insane capital they can deploy, and attention and focus and phenomenal architects and engineers, and a pre-existing catalog of as much of the world's data as possible. Counting them out would be completely absurd.
It started to feel just a bit like they were potentially content staying in the effective-fast-cheap lane and ceding frontier models to frontier labs, but I don't think anyone seriously thought that Google was going to sit back and let this generational shift pass them by when they're so well positioned to remain on the top of the heap.
If anything, I thought that 4 was maybe coming up short or causing safety concerns that were forcing them to be a bit more cautious, but nothing packs a wallop like being able to launch ahead of everyone else's brand newest toys...
I agree on all except for your comment on attention and focus. They've seem to be very scattered in their deployment of AI for customers. Examples in point being the replacement of Google Assistant with Gemini, the AI Studio vs Vertex disconnect, and a crazy amount of coding tools and platforms (Duet, Firebase Studio, Studio Bot, etc).
Working through this as a consumer and enterprise customer has been a confusing experience, to say the least.
Strongly agree - their approach has been extremely scattered. But also, doesn't it seem like they're starting to consolidate? It kinda makes sense to approach things like this as a mega corp. Have internal teams try the same things in different ways and let the market pick a winner. Kinda like being a startup incubator.
A vertically integrated AI stack won't fix the misaligned incentives of google execs.
AI is strangling their other major business and they don't like that. Unless Google sacrifices that branch of income, they won't put their full force behind AI.
I'm curious -- in your view, what exactly is it strangling that could not be addressed by replacing google.com with a ChatGPT-like interface while integrating other Google free-by-ads tools nicely? Or is your point they wouldn't do that due to existing "siloed incentives"?
The ads in google search are ~55% of their revenue and ~70% of their profit. Will the LLM drive as much traffic to the ads as the current google search? Will it feel as organic to the users as the current solution, or will LLM hawking a product seem forced by the users?
I mean it sounds wise to not go all in on publishing expensive models when they can earn money with more efficient models while still not bleeding too many customers?
Alphabet also has nearly infinite money to pour into this endeavor from their advertising enterprise, so much so that AI/LLMs could all fall apart and they would be happy to go back to selling mid-roll campaigns on youtube videos.
Those are competitive advantages but a "moat" usually refers to being on top of the hill and able to keep anyone else from climbing. I think it's arguable whether Google is on top of the hill at all, and if they are they're sharing it with at least two other companies.
AI is a commodity. One that is showing to be more readily commoditized than most has anticipated. As of now, the only moats are the financing for the hardware to run it and the hardware vendors themselves - with the latter largely not yet a commodity because of ecosystem lock and a limited capacity of the most advanced fabs in the world.
The present leapfrogging is not a contraindication because companies are not necessarily releasing their best models; we know they have smarter internal models. Furthermore, humans are still involved in model creation. Human involvement is expected to decrease over time, and when model iteration is completely automated, progress will happen at the machine's pace, leading to runaway intelligence, barring any ceilings.
Maybe. Or maybe having the best AI model on the planet becomes like having the best super computer on the planet. Useful for some niche stuff, but not too useful in terms of people's daily lives or what is used in business.
Already businesses that have more compute and access to data seem to eat the world around them. If, and ya its and if, we can make something that self learns into RSI it's not looking like any business that came before this.
The one thing I'm certain of is that if some group of people can make a system that self learns into RSI, then several groups of people will do so. There will either be 0 RSI systems because it turns out to be impossible - unlikely in my view - or multiple. But we won't have just one.
If there are multiple they will cost money to run. In that world I expect there to be a correlation between costs and quality, i.e. the highest quality AI system will likely cost more to use than a lower quality AI system because there will likely be more compute required and so on.
So in that world, the absolute top tier best in the world frontier AI will not actually be the most used system. This is for the simple reason that such a system will be more costly than a lesser tier system that can do the job just as well.
> Already businesses that have more compute and access to data seem to eat the world around them.
Do they? Business valuations are more a reflection of what people believe will yield returns than a reflection of reality. A lot of big tech behemoths could disappear from the face of the Earth overnight and it would make little difference to our lives.
Natural resources and energy are the actual backbone and it is a true catastrophe when these are disrupted. In comparison, disruption to compute or data is just an inconvenience.
I just find this unlikely personally, think about the great research that's happening in the open source world, I'm sure inside anthropic + openai they've also made a bunch of discoveries and improvements (and I'd guess way more due to them attracting the best talent + the better internal models they have)
At some point employees will no longer be required for the next iteration, and that might be where the take-off happens that Amodei and others have expected.
The whole winner-take-all idea seems entirely based around Singularity/Rationalism and would require massive advances that we probably aren't close to at all.
yeah this would be one interpretation where "winner take all" could still be right. though with the recent accelerating release cadence, doesn't it seem like the head start / lead has been shrinking over time?
I never understood this race to AGI thing. Why is this desirable for shareholders? If God is created, it seems very unlikely God is going to work for shareholder value.
Golden retrievers are the lucky ones. It didn't turn out so great for most other animals. Synanthropes are rare, extinction is historically more likely.
God isn’t going to follow anyone’s directions. There’s no reason to think making God gives the maker any advantage. Perhaps it’s a disadvantage, if God didn’t want to be made.
My personal theory is (assuming there really is no moat) whoever starts the latest with developing AI models might actually win as they should be able to develop a competitive product with significant less resources and initial investment resulting in a higher ROI. AI might even become a commodity.
In the case where it becomes commodified and american tech companies are making thin margins, and where european companies can then build their own data centers in a financially sustainable way. High margins arent a garuntee with this stuff in constrast to ads.
~ unfortunately both Anthropic and OAI already do high end "bespoke" consulting work for business clients, the so-called "Mistral approach" is just a crappier version of that unfortunately.
Mistral / Europe just doesn't want to be in the fight.
I don't really see why Europe would be special in this regard. I would bet on a non-frontier, late-moving american or chinese company taking on this role before a european one.
Based on historical developments the cost of compute will go down again eventually, decreasing the cost of training AI models of the same quality as today even further. That part is what I would be the most certain about.
One analogy I have heard is that distillation is like waterskiing, where the water skier seems to be moving at rapid pace and only just behind the boat, but that they aren't really expending any energy and if the boat slows down, so will they.
This seems to match what we are seeing where Chinese models from companies with only a tiny fraction of the compute are able to be hot on the heels of the frontier models.
So far, yeah. It doesn't eliminate the possibility over the next 5-10 years of the AI race that we wont encounter a scenario that results in a well positioned lab making a clean break
Google has the internet data and user information. However the more they wait, the more the internet will fill with AI generated crap. That’s why there is a race to buy books and scan them (destroying them in the process). To get more data that’s not just barfed out AI slop. So they can win and be the ones spewing slop for others to ingest (human and robot alike)
Nobody has a moat, but everyone is f*&$ed. Competitors do seem to leapfrog each other, but each step is getting closer to beating humans at most tasks. Once that happens, that do we do?
Nobody is going to make money for anything, perhaps.
Or maybe everyone makes money for everything.
The world could turn into a world of plenty. Or it could become a YouTube popularity contest where the MrBeasts get to eat and nobody else is interesting enough to sell themselves.
This is an absolutely crazy time to be alive and most people still don't see it.
Yes, considering those things helped us build (and preserve) civilization for the remaining half of the 20th century.
WW2 ultimately brought massive wins for everybody who didn't die in it or otherwise suffer directly. We are now on the verge of throwing away many of those wins, but that's neither here nor there.
It's crazy anyone ever thought there could be a moat. Almost all the research was/is published in the open. There's no secret trick, no hidden method. LLMs are a commodity technology. Anyone can read a paper, write software, and train models. How do people think llama.cpp works? It's not magic... it's software.
Hardware is the real differentiator. Not everyone has billions to make more advanced chips, and only a few companies can make them anyway. Both OpenAI and Anthropic would already be dead in the water if we had cheaper GPUs, because we'd all be running open models on local machines with 8 graphics cards. They're gonna have to force a hardware shortage to prevent a collapse in 2-3 years. My guess is it'll be tariffs or import restrictions or licenses to buy newer hardware.
Google said as much shortly after GPT3.5 was released. There is no moat.
With that said, there are some other moats and people are building them: training capacity, inference capacity, brain capacity (literally buying the best researchers and keeping them tied down), harnesses/subscriptions, etc...
it's going to be mostly open models running on consumer grade GPUs except for a minority of use cases. Depending on what eval you use, the best 20-30b parameter models are at about 5-8 month lag behind the frontier closed source models and that's been consistent for a while now.
There's a longer talk I give on this but briefly we are in a time period where the development of the technology is quickly outstripping people's needs and they will be satisfied by models runnable on commodity sub-$10k machines.
That threshold has arguably been met for many users over the past 6 months and it will continue to be met for the majority of use-cases in the upcoming months.
So not only does Anthropic and OpenAI have no moat - unless they have other compelling products, the demand for their token-based subscription and metering products isn't sustainable because consumer preferences go elsewhere after any product quality reaches a sufficient baseline.
Luckily there's many many options. Faster, cheaper, stronger they're there but these are all zero-profit condition qualifiers. The labs current push, and this is across the board, is to have a compelling suite of applications where people will have a preference for them or there will be some side hustle. Look at Meta muse - they have an app for your phone that will hoover up your personal data in the name of convenience. 5 million people said yes in 3 weeks.
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
Judging by other comments here this model has been available internally for a few weeks. So the public release is likely imminent. The timing of the announcement might have to do with end of quarter for some obscure reason.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Google AI Plus is not the equivalent to most other paid plans, it is closer to ChatGPT Go. 3.6 flash is an equivalent model to luna 5.6, which is the highest available on ChatGPT's free & Go plans.
3.8 flash has been perfectly available to Pro users from the announcement day.
You are on the cheapest paid plan and complaining that you don't have access to more expensive models. (Though idk why you only have the lite version, I have access to the non-lite version even on the free gemini plan as long as I'm logged in.)
It's annoying because on Plus, you used to get the latest models. Then for unspecified reasons and without announcement, you just stopped getting new Flash models.
That's strange since as Plus, you still get access to latest Pro (albeit 3.1) but not the latest Flash.
I would understand if I Plus subscribers still got access to 3.7/3.8 Flash, but it just burned through the usage limits faster.
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
Fuschia is not a desktop OS. It's designed for lower end or embedded hardware. Besides, Android has the highly profitable app ecosystem so it makes more financial sense to build GoogleBook based on that.
It wasn’t actually. From the docs for Zircon (Fuchsia’s kernel) [0]:
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
Interesting. Have heard many times that Google has become risk adverse. Maybe fomr something like Android, they just don't want to rock the profitable boat that it has become by switching out the core OS? ... Also now that we're talking, maybe it's just too much work to do so. Backwards compatibility seems like it'd be a nightmare, especially with all the device drivers.
Just because the underlying kernel (Zircon) is designed to be good for both phones and computers doesn’t mean that Fuchsia (which is just an OS that uses that kernel) is.
Some of their smart devices are using it. They've been focusing on GoogleOS (merging Android + ChromeOS) for a while but when that's done I guess they may adopt some of Fuschia's work. But I doubt if there will ever be scaled efforts to replace the entire Kernel stacks, even with the frontier models it is going to be very hard.
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
"I don't know if C++ will still be relevant in a few years."
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Yeah that was abundantly clear when Bjarne Stroustrup published "A call to action:
Think seriously about “safety”; then do something sensible about it"[0] as a reaction to NSA's recommendation to no longer use C/C++.
Well, sure, but that doesn't mean it's not in a decline that is very unlikely to reverse course. Fewer people this year are starting new projects in C++ than last year, and fewer people than that will be starting new projects in C++ next year. And, as the models get really good at porting and the cost (both in terms of tokens and human supervision) comes down, there will be an avalanche of ports from less safe languages to safer languages.
This year, maybe next, maybe a year or two after that, is probably the most C++ lines of code in production use there will ever be. Why would one choose C++ for new projects at this point? There are niches where Rust is still uncomfortable or just doesn't have the support, but not for much longer. Models are very good at Rust and good at porting to Rust. And, Rust is a good language for models because it is so strict...it helps keep them in line.
> Fewer people this year are starting new projects in C++ than last year, and fewer people than that will be starting new projects in C++ next year.
Let's put [Citation needed] on this and let's assume that your personal opinion bubble of developers doesn't really represent the combined worlds software industry. Mind actually proving that there's a significant decline of use of C++ outside the hype crowd?
> Why would one choose C++ for new projects at this point?
Why indeed would you choose a language supported on every running platform on this world with a massive library of supported dependencies and mature compilers developed by stable teams.
(I'd probably also personally choose Rust for new projects, but let's make this a practice in "not everyone is like me" empathy.)
> I don't know if C++ will still be relevant in a few years.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
If no one is reading or writing Rust code at that point, it all being agent driven, Rust itself is going to be a very short step before it gets disinter-mediated away and it's English -> complex heterogenous machine code across GPU/TPU/CPU/xPUs. Not sure Rust fans or C++ haters have thought this through.
This is such a bad take given the current LLM agents
High level languages :
- compress low level patterns into high level abstractions which save context/token
- express higher level meaning through things like names
- express and enforce constraints through type system
LLMs are context constrained, do really well with tools that help them iterate towards a correct solution (type systems) and work on human language (although they are compressing it to be more token efficient).
Why would you take something compilers are really good at (translating high level concepts into low level hardware instructions) and bake that into LLMs
If anything LLM programming language will be something like high level IR that's token optimized, but honestly the AI labs are so lazy (or bad) at fundamental engineering (look at their sandboxing solutions LOL), they are going to keep pumping training on shell/python/rust RL environments.
At some point, models will be so smart, hardware will be so cheap that it makes sense for the LLM agent to just write code as low level as possible to optimize for performance and efficiency.
These things are hard for a human to do, but easy for a machine to do.
Models are likely already doing this inside AI labs because it can be hundreds of millions of dollars for every percentage point increase in hardware inference efficiency.
A smarter agent can write a better compiler and leave the intent in densely packed/verifiable language structure that can encode propositions low level code cannot.
A concern with these LLM-driven rewrites is that the Rust code isn't really idiomatic, it ends up being C++ code in a trench coat (raw pointers everywhere, manual validation of thousands of unsafe blocks). So far every publicized rewrite I'm aware of has had an "oh and we'll definitely spend the thousands of hours needed to fundamentally rework the code to properly utilize lifetimes later" admission tacked on at the end. Has there been progress in improving that aspect during the LLM phase of the rewrite? If so that would be an incredible win for the industry!
> Has there been progress in improving that aspect during the LLM phase of the rewrite?
I'd say so! We aim to have our replacements to be `forbid(unsafe)` - no raw pointers - outside of `ffi.rs`. Of course it's tricky because we need to interoperate with existing C code, but we've been trying to come up with abstractions that make this a bit more ergonomic (https://github.com/google/safer_cffi).
LLM should soon be language indifferent/independent, the upcoming years Devs will touch less and less code and more of a LLM as the top interface and debugging tools as the next layer. I just can't see anyone writing an ounce of code in 1-2 years.
I'm still amazed that they still have coding interviews that tests you on syntax. when agentic/ai assisted engineering practices is way more important. What they should test you on is the top layer like architecture, concurrency concepts, UI optimization concept of a particular language API, how to debug, how to test, more of a how to read code instead of writing. They need to test you on speed at which you would detect bad code, over engineered code, bad architected code and how to remedy it.
The notion of human-readable/human-safe programming languages will be irrelevant in a few years, not just C++. They may use Rust, Zig, Lean, or maybe some other LLM-generated hyper-optimized intermediate language with a strong type system and compiler, but chances are, we will read less and less of it, until we basically never read code at all.
I know this is somewhat of a meme now but I do think the predominant language will be "typing/speaking to a coding agent in your native language."
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
The missing piece is whether “75%” has the same denominator in each chart. If three models were evaluated on the same benchmark set, against the same versions at the same time, with exactly one winner per benchmark, their shares of wins would add up to 100%. They couldn't each win 75% of that shared set. But different benchmark selections, older comparator versions, different reasoning or tool budgets, and later releases change the comparison. Counting ties as wins changes it too. So the marketing percentages alone don't tell you how many benchmarks exist or how much selection happened.
To audit a particular claim, I'd want the full shared matrix, exact model versions, benchmark versions, test dates and evaluation settings, rather than just the highlighted cells. That also separates actual leapfrogging from a change in what was measured.
I maintain https://llmbenchmarks.io/analyses/#test-setups, which illustrates some of these differences using published measurements and their sources. Even results for the same model and benchmark can involve different setups. That isn't a verdict on this particular Gemini release; its specific claims need their own source and protocol checks.
You are Google, training a new model. You regularly benchmark the checkpoints. It's gradually improving as you train. When do you stop training and release the model?
You certainly don't do it when you are just behind the frontier. Instead you wait until you are ahead on a lot of benchmarks.
So, if Google release a new frontier model, the chances it's ahead on most benchmarks are statistically about 100% because if it isn't yet then they won't release it.
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now.
I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
there was already an optimised c++ library and a rust port that was safer but less performant; they managed to get a new rust version that recovered a lot of the performance gap. sounds pretty damn good to me!
I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.
tbh while the astra announcement in that link had some good stuff it was scattered through so much boilerplate marketing speak that I had to force myself to read it and look for the content. I found the gemini blog post in the OP a lot more readable and engaging.
but that's a side issue; my main point is that you are underrating the impressiveness of getting a safe rust port of a highly optimised c++ library even nearly up to par with the original. the tradeoffs rust makes for memory safety cost it some of the raw speed of c++ even with all the zero cost abstractions and purely compile time guarantees they have. (tangentially i wonder if ats (https://www.cs.bu.edu/~hwxi/atslangweb/) would be a good candidate for LLM assisted ports; it seems way more advanced than rust and might actually get c-level performance with safety, but it's really hard to write.)
bounds checks are one, yeah, but I was also thinking about the tricks c++ could play with optimising undefined behaviour, as well as unsafe patterns involving shared access and pointer aliasing that a human could determine was safe in that specific case but that the rust compiler would balk at. and maybe it wasn't actually safe in which case the rust code would have the last laugh.
AFAIK, Rust benefits from the same optimizations that Clang can. In some areas Rust can actually benefit more from better aliasing information. But you do have to use the keyword unsafe for some things that you can Just Do in C++. You're right it can be difficult to achieve hand-optimized C++ levels of performance with Safe Rust.
Astra, Sol, Terra, and Luna are also complete nonsense that I can't keep straight. I have a reference card under my monitor to remind me which one is which.
And that makes sense to you? From where I stand, here on Earth, the apparent sizes of those things are, from large to small: Earth, Moon, Sun, Stars. And I don't think in Latin.
It's dumb to have to think through multiple levels in indirect descriptions to remember which model is which. Only ADHD dipshits think this is a way to name products.
Yes it makes sense, why would you name them based on apparent size? Everyone knows the scale of these things, we learn them in elementary school. I don't think of the sun as smaller than the Earth, and I definitely don't think of the sun as smaller than the moon. Moon < Earth < Sun < Stars makes perfect sense to me. It (imo) is much clearer than Haiku < Sonnet < Opus < Fable, although I get what Anthropic is going for there too.
And I don't think or know Latin either, but these are pretty universally known words even outside of Latin (besides maybe astra, although "astral" is very close so you should be able to figure it out).
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
i believe they've got inference issue. the phone app never got past 3.6, and they're going to roll this out to Ultra subscribers before other subscribers. none of this screams they're ready to flip a switch and start serving a ton of traffic as OpenAI/Anthropic routinely do.
nothing about this announcement gives me confidence that google is back on track as a model provider.
Not sure what's the problem with phone app rollout, mine has 3.1 Pro, 3.5 Flash-Lite, and 3.8 Flash (with optional extra effort), and pretty much it switched to latest Flash series as main option since 3.5 just after release
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
Doesn't make sense for them to give away weights. They sell their own service using that and killing the competition (that they have a stake in) isn't in their interest either because their cloud growth is based on other companies' continuing heavy investment in this space.
They do release the model weights for Gemma. But could anyone actually run Gemini 4 weights? It's probably like a 10T model, which you need an industrial rack for anyway.
all the enterprise clients which the two frontier companies rely on have the money to buy the hardware and run a frontier model themselves, it just becomes a nobrainer to buy your own hardware if you have 10-20b token output per week
Well... given that I have codex reporting 1b tokens consumed for my couple-days long session, I suspect 10-20b is not that hard to reach for a company of 10 devs. It would be interesting to know what is "rental fee" for a model like Astra or Sol 6.1, and how much hardware they actually need - not just gpu, but all of it?
To be comfortable you need 300 to 400 terabytes of storage. You need the GPUs and the cooling for them. They don't need a lot more. You put models in memory and you do math really fast and make tokens. You need some CPU and RAM infrastructure for like prompt caches and things like that, which does add up, but the dominant cost is by far the GPUs and VRAM. Everything else is a pretty easy step down to afford. You probably need to have multiple different models, right? Like if you want an Astra plus a Sol plus, you know, like a Luna-class model, say, you probably need like 10 terabytes for Astra, probably 3 to 4 terabytes for Sol, and 1 to 2 for Luna, maybe less, maybe one or half a billion even for Luna. That's RAM. You need So all total you need 15 to 16 terabytes of active VRAM to keep those models resident. That's not counting what you want to keep in cash and whatever hosting constraints these models have that aren't public. That much VRAM is five to eight million dollars right now.
Plus what models are you running? The biggest models you can run are like in the terabyte range or Qwen or GLM. Kimi K3 is there, but GLM 5.3 is very competitive. Using voice dictation, hands hurt, sorry for any errors. You still want many terabytes of RAM to run say multiple GLM+DeepSeek V4.1 for some actual concurrency.
We have had this discussion for a decade regarding cloud hosting. There are definitely workloads and AI inference that you want to do in-house. Hosting these GPUs and models and setting up the interconnect and balancing it is really hard right now. It's not traditional infrastructure. It is exotic. Typical enterprise is at such a disadvantage to try and host this completely themselves. And models are constantly updating. And what models are you going to run? Six months ago, things that probably cost you a billion tokens probably cost you 50 to 100 million now, given how much more efficient models have gotten, or at least they cost that much. I don't know. It's just not a given.
Frontier LLMs that Google doesn't control are existential threats to Google's business model. If no one is using Google search or browsing websites themselves because they're asking LLMs, Google doesn't make ad revenue. Gemini at least has a path to monetization for Google (subscriptions, ads baked into responses).
A competing model still needs a search index behind it to be a threat to Google. An open weight model doesn’t have that. The other frontier models still struggle to fulfill a lot of queries for lack of one.
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
Terrible take, it’s tied with Astra Max yet you decided to round it down for no good reason.
Then you try to pin discount pricing when everyone has some discount going on, and knows they will release some new model before it ends anyway.
And if you’d have used Muse Max at all you’d know it’s not been really in the same league as the rest. Then again Astra and Sol are both not really competing with Opus now with 5.5 either.
It’s a solid release on paper and until anyone uses it I find anyone trying to proclaim a loser completely unreliable.
If they’re smart they pull a DeepSeek and teach flash to beat this entirely within a month or two, I wouldn’t bet against it.
Rounding down was my reading comprehension mistake. I should have read better.
Also, I have had extensive experience with all Gemini models except the recent few. Generally, they are always less precise and less practically useful than benchmarks suggest. You are right that here there is no evidence to think that. However, them announcing the model benchmarks without allowing access to anymore makes me think it might follow the same trend. Another point is the verbosity, I always look at that number to get the feel whether the model truly got smarter or it just wrote so many reasoning token to get there, Argon is on the higher side.
I would be very surprised if in practice, it was more practically useful than Astra, Sol, Opus or Fable.
> If they’re smart they pull a DeepSeek and teach flash to beat this entirely within a month or two, I wouldn’t bet against it.
Yes this would be a great outcome for every consumer.
Benchmarks are only part of the story, most workloads don't match what you see on these.
For example sonnet is supposedly the same price as opus from medium and up, I tested it on a real world task at my day job and sonnet was a good 30-50% cheaper
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Argon is doing my job for me while I'm writing this comment.
A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!
Out of sight, out of mind. It is unexciting because in that future, there is nothing to think about related to this topic. What is exciting about something that is irrelevant for humans since nobody looks at it anymore?
As someone who grew up poor, and then used to be on food stamps and other state aid as a young adult, that's a particularly pesky problem. Dread is a more apt feeling than joy or elation in that situation.
Because (maybe particularly) at 1e100 scale, having a super intelligent butler that does whatever everyone asks of it can easily get you to hundreds of thousands of daily PRs merged to the monorepo, but won't necessarily help coordinate efforts, and may actually cause certain projects to collapse under their own weight.
I am in two opposed minds on this. On the one hand, you're correct. Their job will be made redundant in a couple of years. Some senior people might be retained to orchestrate all the agents, but eventually they will also be replaced. And it's coming for all white collar jobs, not just coders. That's bad.
On the other hand, unlimited intelligence democratises products and services in a revolutionary way. Everyone will be able to do whatever they like in digital space for basically free. No more purchasing software. Make any app/company/website/blog/movie/show/game/forum you can think of. This will quickly bleed into the physical space, with real-world products and services becoming logarithmically cheaper. A true future of abundance.
The big problem is going to be our social and economic transition. Are we going to wait until unemployment is that 10% before we act? 20%? How bad are we going to allow things to become before we implement universal high income? I say universal high income because white collar workers are not going to accept being relegated to an income of $1000/month forever. To achieve universal high income requires eye-watering taxes on the major benefactors of our new economy. They won't accept that without a fight, and it our current fractured political climate, what would it take for everyone to agree on a unified direction?
I am confused about the whole situation let me explain.
It seems like all software will get replaced it by local versions. Suppose I am an individual or a business that was using a SaaS product (such as JIRA).
Let's be generous and imagine that the monthly payout was $20/month for that software (usually for businesses, it is much higher) and one fine morning I login and point a multi modal AI agent to it and clone the whole thing for my local use in let's say 3 months worth of $20/month AI subscription.
Now what? Gradually, that software company would go out of business but that user won't be subscribing forever either to AI inference service and the only constant customer of AI inference would have been the software company which went out of business so... No software company, no AI inference business either?
If I understand your question correctly, it reduces the value of software to almost zero. Meaning all of the SaaS companies are going to have to work very hard to diversify their revenue streams into value added services. That's going to be very difficult or impossible for most to navigate. Hence the term in the industry "SaaSpocalypse."
Some companies will retain a moat. For example, YouTube holds incredible IP, and until AI can replicate that wealth of content at scale for cheap, it will remain profitable. Facebook has the user critical mass. Increasingly, hardware for inference is going to be a moat. So is electricity.
Which loops back to my comment above: this breaks a lot of the market dynamics required for a functioning economy. Accelerationists are convinced that we'll invent new jobs to replace the old jobs, but I don't think that's going to happen. Certainly not at the pace of change we're seeing.
Basically only moat left are either if you have "Data" or "Compute". Software layer is worthless and pretty much reproducible.
Data as in YouTube is all just data. Building a website and app around it isn't the hard problem. Anyone can launch another YouTube in few months but that would be pointless.
Yes, and increasingly, data will no longer be a moat either. These models will be capable of creating all the data/content we could ever want. There will always be a market for "organic" movies with real actors, but we should expect the value of that to decline precipitously.
My investment thesis is currently around these strata:
1. Electricity. All analysis shows massive shortages beginning now with no end in sight. Western governments in particular have massively dropped the ball on this one.
2. Chip design and production. Nvidia and TSMC dominate here for a few reasons. Partners required in the supply chain like ASML benefit, but they're not really the bottleneck right now. We're seeing a lot of investment pouring into this space and we'll see competent design competition (especially from Google's TPUs, and potentially from Apple and AMD). Chip fab is very difficult to get right, and takes a long time to ramp up. See the NAND space.
3. AI labs. These guys are doing to get commoditised. They're trying very hard to prevent distillation but I think it's impossible. The open weight models are not far behind, and the Chinese labs are getting very creative about how they work with limited hardware. If anything, they'll decrease inference cost over time, reducing the value of existing data centers (of which OpenAI and Anthropic have purchased an enormous amount). I am very bearish on the future value of OpenAI and Anthropic.
4. Software. So much software. Everyone gets software. The companies which benefit from this will be able to monetise it. Google, for example, might put ads into VERY capable and free AI. People might accept that. Microsoft has been investing heavily into their data centers, and maybe that's their revenue center in the future. Especially securely hosted services. I expect massive growth in this space in the short term. As long as people still have high paying jobs, they'll pay for well made software and tailored services.
> As long as people still have high paying jobs, they'll pay for well made software and tailored services.
Here's the question. Which high paying jobs will remain? People will still be needed, obviously, but salaries would be in a race to the bottom without some form of govt intervention.
The name "universal high income" should be tipping you off to the issue with it. If money has less importance then what tools will individuals have to change their station? How will anyone other than the permanently entrenched rulers affect change? We don't have a good system now, but it seems like we're hell bent on making it even worse.
Edit: UBI being pushed by asset owners is a step towards Blade Runner rather than Star Trek.
I believe that the middle ground for this is to keep all these jobs orchestrating agents or whatever the AI future invents for us, but with lower wages. The golden era of dev wages will be gone. There is no incentive for companies paying people 100-500k-1m to write prompts and debug agents, independent of their level.
This might be how it shakes out but I don't think it's acceptable, and I don't think white collar workers will think it's acceptable. Taking high income people and making them low income people (usually for life) has a tendency to politically radicalise people. Especially in light of the unprecedented productivity and profits which businesses will experience. People will demand a slice of the pie.
One kind of intervention that could work would be licensing, like with some engineering professions. Make it illegal to work on software without a professional license; same as with doctors. Engineers would be free to work with AI, but the final responsibility would fall on them.
This artificial control works to prevent oversupply and protect salaries. Unfortunately, the businesses calling the shots won't ever allow this to happen.
I'd like to see AI "democratize" the world, but I have a much more pessimistic outlook. One thing we see with technological revolutions is that (once the dust settles) they provide general improvements for all strata of societies. But as some boundaries are erased, others are magnified a hundredfold wherever robber barons can stake an exclusive claim.
If genuine AGI is developed, its everyday conveniences may continue to be shared with the common masses, but the powers that own the technology and the data centers will have no incentive to give everyone access equal to superintelligence, and that includes giving governments the same access.
Every stage of the industrial age has seen powerful, wealthy people/groups (whether you call them oligarchs or industrialists or whatever) become able to exert disproportionate power over the state, and we're already in a place where this is happening. Humans have shown time and again that once you have some resource (whether it be oil, gold, or data centers), someone will want to own all of it, and that they're always willing fight for that ownership, and this fight rarely benefits ordinary humans like us.
Developed countries have been enjoying an unusually long peace time thanks to the wonders of capitalism and the emergence of stable, progressive capitalist-friendly democracies. The road to utopian basic income and limitless leisure is going to be fraught with riots and violence, and make Marx suddenly utterly relevant once again, because there's too much at stake, and too stark a division between the haves and the have-nots.
The alternative timeline is one where governments get together to regulate AI into some kind of international agreement under the UN. That requires taking power away from the private corporations; historically, that's been a tough sell.
> On the other hand, unlimited intelligence democratises products and services in a revolutionary way. Everyone will be able to do whatever they like in digital space for basically free. No more purchasing software. Make any app/company/website/blog/movie/show/game/forum you can think of. This will quickly bleed into the physical space, with real-world products and services becoming logarithmically cheaper. A true future of abundance.
This isn't happening yet, SaaS is only growing more in the era of AI. Anecdotally, my own small SaaS has only grown instead of shrunk as well.
Finally circumstances will force me to become a goat farmer, a desire I've always secretly aspired to but never felt comfortable expressing, let alone acting upon
Now we will finally be free, without having to make hard decisions. It will be decided for us.
There is no credible path towards stopping progress on AI. Nobody knows exactly how this will play out. But whichever way it turns out, being an skilled AI practitioner will give you strictly more options than not being one.
My believe is that current LLMs performs very well in the constraint environment with lots of data available, e.g. we have lots of code on github, public discussion with feedbacks, compiler and tests failures as feedback, so LLMs can train on trillions data records.
That's not true for many unconstrained real world scenarios, and LLMs currently perform poorly there. Another breakthrough is needed to overcome this.
you do need a precise prompt to do what it's asked and there are 50-100 things that needs to work to meet the criteria. Little things need to work, and work well
So basically people who didn't, or couldn't prep for this will be 'effed. Not accusing you or something. Just thinking out aloud. I am/will be one of those people.
Sometimes I wonder AI won't have to hack military drones to end humans en masse. Making relevance, employment, and other life conditions for most so pathetic that it will start making it happen on its own.
welcome to the unemployed future! first bug to bring down production and you're fired, will have a good memory that at least you were with your girlfriend
You assume that a bug that can take down production is more likely to be produced by the LLM than by a human. That very well might not be so with current models and especially in the future once they get even better.
It's coping, the truth is writing code by hand is dead.
3 years ago, these things became useful. Every year, they're drastically better. Most juniors are already dead weight. Most of the rest of us are next.
IDK how it could be otherwise. It's moving too fast and gone too far already.
I'd love to know if you still felt that way after seeing what I do in a day. Company IP won't allow it. Maybe I'll stream some personal dev work or something.
I'm not saying you're wrong, but, I'd be shocked if you were still so adamant (assuming you're rational and not in your own psychosis).
Here's something I'm about to drop on AI. Like 90% of things like this, when I ask for a fix (with a few skills), get fixed with fewer edge cases left dangling than I'd have left if I'd spent all week on this stuff. And, it'll be done before tomorrow, with no further input from me. This is a list I put together in about 1 hour.
Things to fix with OCR text labeling
- vertical line in center breaks box-drag
- nope, not line in middle... IDK why, but the drag pretty frequently snaps back to original position in middle of drag... lets me keep dragging, but have to start back from there
- happens with "draw stamp box" too
- forces me to re-click the button to start over
- add zoomed view of text in bottom right corner of video
- add ability to restrict OCR
- current job is 0 1 2 3 4 X 6 7 8 9
- 5 is an X, for some reason
- is detecting full alphanumeric
- add ctrl+z undo for modifications to drawn box
- deleting character in middle of "your label" text box jumps cursor to end
I could do even better than you in 30 minutes and let it write code for a satellite if I wouldn't care about quality. But at this point I have the feeling I'm talking to a bot
Yeah, I'm not saying this is an example of great work. It's sufficient for the job - more time spent refining my instructions isn't necessary.
What I'm saying is, with just the effort it takes to notice and somewhat-sloppily mention things, they get fixed, and usually fixed well. Further testing will tell us whether it needs more work or not.
They may have never had these troubles at all, if I'd coded by hand originally, but then, this particular app would've never been written, because it wouldn't justify that amount of work.
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
Fable is available for a couple of months and even got an update on 1st of September. It’s really good, but since Opus 5.5 was released, there is not much point in using Fable anymore
As we’ve seen from the leaked Anthropic prospectus, revenue from actual users is a pittance. What really matters is what you can get from investors, and that you have a model smart enough for self-improvement.
I could start a business with a similar growth trajectory. We can mail people $100 bills for a low low payment of $10. I just need $500 billion dollars of startup capital and I can show you a 10x yoy growth for a few years.
If they, through standard accounting practices, cannot show they are profitable how do we know they are profitable? I assume they can't just say "we are profitable assuming you ignore the costs we paid to build up XYZ". I'm guessing they could pre - rent compute for 10 years or something and maybe that's all on the balance sheet but you'd just amortize that and I think that falls under GAAP?
That sounds fun but this is people paying for real products and services. The demand is there (and it's growing faster than any other commercial technology in history). They just need to work on inference (and potentially training) costs. I think that's clearly achieveable.
For the record, growth plays have been common in Silicon Valley for many decades. It's how we ended up with all the major tech companies we see today. Of course, it's high risk, high reward.
Oh spare me the pedantry. What else are you going to compare? The year-on-year for Q2? That’s certainly going to have a much higher multiplier.
Your guess is as good as mine on whether these companies will continue growing, stall, or decline. Don’t try to pretend your methodology is more sound though.
Their profit on that was reportedly $559 million. But, this is after you realize that Anthropic got a discount on May/June from their new agreement with xAI's Colossus data center. So their first profit positive quarter happened during the time when their costs were artificially lower than they will be going forward.
Lots of people seem to love to do the simplest extrapolation that benefits their point. It's the same with "AI is improving at rate X, therefore if we linearly extrapolate, we'll hit AGI in a year."
We have access to fable on enterprise. It works well, but do you really need to pay that much. Opus 5.5 has been performing well for us, so we're mostly standardizing on that.
I opened my Google account to check the model and honestly have no idea how to access it, whether it exists or not, what is Google AI, what is Gemini, what models are available, where the chatbot is, where the API is, etc..
Honestly, who is in charge of this? How do people even subscribe and make sense of any of this?
Yeah, the naming and proliferation of different tools and sites is confusing. Quickest way I know to get on any plan is via Google One at: https://one.google.com/.
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
It's also more token hungry than similar models, so does it really make a big difference in the end?
With GitHub Copilot pricing, I have found no practical reason to use Gemini 3.8 Flash vs GPT 5.6 Luna. And now I'll probably find no reason to use Gemini 4 when GPT 6.1 Sol is essentially as good and a lot cheaper.
Other than not being available publicly, I find it extremely disappointing about Google’s two-faced behavior. CEO signs an official document with POTUS stating that AI is now SI, but the release of the new model doesn’t mention SI even once. Even worse, it’s AI all over that page.
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
"AA-Omniscience Hallucination Rate (lower is better) measures how often the model answers incorrectly when it should have refused or admitted to not knowing the answer." - Argon is currently the best model by this metric.
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
Hmm, but I thought that each token generated effectively becomes a part of the context window for the next token. So 1mm context + 1mm output means that the 1 millionth output token will effectively have been generated with ~2mm tokens of context. But maybe that’s wrong.
The output token limit and the context window are separate constraints. The context window is how much the model can see at once. The output limit is how much it can generate in a single API call.
In an agentic loop, each API call gets its own output budget. A 'turn' is one response from the model, whether that response contains a tool call, a reasoning step, or a final answer. So with a 1M context and a 64K output limit, the agent can run many turns where the context grows each round (accumulating tool results, prior thoughts, user messages), but each individual response is still capped at 64K tokens.
Expanding the output limit to 1M matters most for tasks that produce a lot in one shot, like writing a full document or a very long file. For most agentic workflows that naturally break into short turns, the per-call limit was rarely the bottleneck. The context window filling up was.
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
What keeps both Gemini and Grok from actually being frontier class AIs is effort. Both AIs rush to give you a response even when the effort is set to the highest setting. Hopefully, Argon isn’t as prone to satisficing and premature convergence compared to its predecessor
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
I recently subscribed to Claude, and was very unhappy about the usage limit of the $20 pro plan. Then I found a trick, since I got free Google AI Pro via my phone carrier, I use Opus 5.5 High for planning, and then dispatching agy to do works.
Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
Gemini 4 Argon (High) looks comparable to Claude Opus 5.5 (High). https://artificialanalysis.ai/models/comparisons/gemini-4-ar... From that perspective, the benchmark is not too disappointing, given that there are only 8 days between the blog posts (September 30 vs. September 22).
> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.
Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?
> There is a spectrum of opinion inside Google. Some employees believe Anthropic's Fable and OpenAI's Astra models are improving at a faster rate than Gemini.
> These people believe that Gemini 4 — even at its best — will still lag behind those models in some areas. Other employees believe the coming version has caught up with the leading AI labs.
> A Google employee familiar with model development said there is "large consensus" internally at the company that Gemini 4 is at the frontier. This person said the company had conducted rigorous tests of the models and denied that they struggle with messy, real-world coding tasks.
Whatever your workflow is, make sure that model and provider are replaceable. Frontier labs will keep leapfrogging each other, as they have been doing for months.
In order for the benefits of AI to be distributed, intelligence has to become a commodity.
As long as you control the skills, the learnings, and the infrastructure setup you will be fine.
And don't tie yourself to a harness. Shun models that don't let you pick the harness (Google). Anthropic is indifferent at the moment because the OpenClaw craze is over. Vote with your wallet.
"indifferent" was the word they used. The sensible way to use Fable and Opus to do work at this point is claude code, but that's fine. I always have several instances running, along with codexes and pis. The question to ask is: can you stop a claude code and start a codex in its place without disrupting the team? Does my infrastructure support learning in a way that is stored by me, and not tied to a provider?
I built a harness that externalizes state into files so you can swap providers. I regularly move between claude and gpt, but it works with all providers including pi and local models.
I log everything - user messages, tasks, project memory, even bash commands for forensics. As a consequence you can do reflection where you analyze past work and extract refinements for the harness and realign the project when it diverged from user intentions. I don't have to do this manually, it reduces steering work.
It's cheap, just add a line to .bashrc, every time bash is invoked this script is executed and appends a line to a file. I rarely need to see bash_history except when I suspect agents did a boo boo.
It might be against ToS, but as a pragmatic matter, Anthropic doesn't seem to care if you use a custom harness with a subscription atm. I know a lot of people doing that now. Using Opus 5.5 as the orchestrator, and Luna max as the workhorse.
This is great in principle but in practice I don't have enough resources to maintain a stable normalization layer across 3+ inference providers, each with wildly different and evolving API/product roadmaps.
I could make it work if I didn't care about access to latest reasoning model capabilities, but then my customers would no longer be interested in any of this.
I tried the DIY provider agnostic harness thing and it performs like shit compared to what OAI, Anthropic and Google's engineers have created. I don't have a trillion dollar AI budget. I feel like these comments are sometimes written with the assumption that the reader does.
Do you use API for google & anthropic or subscription? I was under the impression those two are apt to ban you for using their models in a non-approved way
I assume I'm missing some context because you sound like you have a lot more experience here, but what kinds of capabilities are there that you can't do something like `interface LlmProvider { }` with a few basic methods for attachments and chat and then just implement from there? I'd even say you could not keep the ones you're not using (say, GPT is your go-to for a year) and update the other ones when you want to use them and swap out the `LlmProvider` implementation with a DI'd different implementation at your application's entrypoint.
Assume the APIs are equivalent across all providers and have exactly the same schemas. You still have this massive problem of alignment between the various reasoning models and the tools that are using those LlmProvider generic interfaces.
The various vendors behave differently enough to make a simple code contract swap largely infeasible for many kinds of domain-specific agents. I agree there are lots cases where this does work (probably most of them), but you sacrifice performance in the targeted scenarios where you need much more specific alignment and control.
The biggest example where this breaks down is with computer use and vision. I cannot take a targeted solution that works on OAI's stack and forklift it over to anthropic (or google) and expect anything to perform correctly. Some things will work some of the time, but it's basically starting all over again with alignment each time you move something this complex.
Having a generic interface for executing shell and writing code in an iterative ecosystem is not a hard problem to solve. Having a generic interface for reliably inspecting specific PDF documents in a very particular way in a single shot is a much more difficult problem to solve.
If you follow this argument to its logical conclusion, you will wind up implementing a provider-specific version of your tools. And the ones that demand this treatment will be the most complicated ones.
It is fine to move up a level, and use models with their providers' harnesses, as long as you stay away from extras like their own memory and infrastructure. There is not much investment required, and you will end up with a more robust solution. A simple setup that works:
- Agents live in their own directories, for example ~/Agents/[project]/[instance]/
- In a tmux session run each agent in its own window and directory. Pi, OpenCode, Claude Code, Codex, a combination, whatever works best at any given time.
- Tell them that they are working with other agents, and that knowledge is shared and stored in an agreed format. They are very good at using OKF[1], for example. In my actual setup I use a harvester agent whose only job is to ensure the quality of the knowledge.
- They can communicate with each other simply by sending messages to each other's tmux windows with tmux send-keys. If the team grows, or you work with other people, you can use a general messaging system (I wrote and maintain https://github.com/awebai/aweb but I am sure there are others).
I had access to this over few weeks, and in my impression this was the first Gemini model that I can offload complex tasks that I don't want to do myself because I have to do lots of domain specific researches, which is irrelevant to my daily works. Not 100% reliable, but its outcome is usually better than mine and the cost to verify the outcome is significantly cheaper than doing the task by myself.
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
Can it surpass (or at least maintain) 3.1 Pro when it comes to chatting?
I've been using GPT since its 3.5 release, but starting from 5, its output always contains some meaningless nonsense or provides answers that are difficult to read just to avoid hallucinations. I find Gemini to be excellent for chatting. For example, when I'm pondering a mathematical theorem, it could provides answers that inspire me, even guiding me to think about the next question.
For me, this is the only reason to maintain multiple AI subscriptions. Gemini is irreplaceable in the realm of answering questions or chatting for learning purposes; I can leave the rest to GPT.
When Antigravity first came out, when I installed it, it made itself the default application to open all programming file types.... .txt. json, csv, you name it. It was absolutely infuriating and took me a month to put everything back to normal. So they've lost a lot of trust with me.
5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.
I don't like AI, but it's very enjoyable for me to see my predictions on the success of gemini come to fruition.
I only use gemini, and while I don't use it for actually writing up code, I use it to help me troibleshoot my logic and help find bugs. Its easily the best model I have tried. And yes I am talking about 3.1 pro.
Also I have found gemini the only model to be the least likely to douse me in flattery, and will follow my pre built instructions to never output anything unless it can be directly sourced, pretty well. Chatgpt i tried for a bit and it was by far the worst thing I have ever used. I can understand why people develop psychosis when prompting chatgpt because it is disgustingly scyophantic to the point I was grossed out and felt like I just got done with some other type of self gratification.
Anyway, death to AI. All those who use, create, facilitate, or even just sit by and do nothing in the face of AI will perish in Hell.
Google didn't release an impressive model since Flash 3. I tried Antigravity with the latest Gemini model for a week and it's the worse experience I had among more than 5 harnesses and models I tried (deeseek/pro, muse/spark-1.3, cc/opus, codex/sol, and Pi/sol). I haven't touched it since and canceled my pro subscription.
Then I tried Flash 3.8 for different tasks like OCR and others, and while in their benchmarks it crashes Flash 3, in my experiments I didn't notice much difference, often even Flash 3 performed better.
I hope Gemini 4 Argon is a real step up from that but we'll see once they release it. I'm rooting for Google and it's about time they deliver frontier intelligence, not just competitive prices.
Google has been quiet for some time. Surprisingly, every benchmark shows it ahead of other frontier models, but when you actually use it, we'll know how it performs in the real world.
Google attempting to put itself into the frontier conversation simply by saying so is actual lol-inducing. Gemini has been laughably bad for like 9 months now.
I want to be a 'polyharness' maker, so I occasionally switch between Cursor, Claude, Codex, OpenCode etc.
I've tried Gemini (without Antigravity) only about 20 times and it seems of significantly lower quality than the other flagship models (e.g. it missed obvious deductions for my tax return, and often refuses to do things like very harmless/legal web scraping).
Logan Kilpatrick announced [0] that training for Gemini 4 started towards the end of July. I'm surprised they could finalize a frontier level model within 2 months.
If an entire model training can be completed in 2 months there is really no moat for any company in this space now.
> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.
Maybe they can push it to make the Android accessibility framework more responsive, especially when scrolling the screen with TalkBack, and catch TalkBack up with VoiceOver. Oh and add Accessibility Actions to apps like YouTube so I don't have to swipe through "video name", "video name button", "go to channel button", "more actions button", every, single, video.
But they won't because accessibility is something you have to actually prompt the model to do and who cares about a11y.
To bad that Google will not resisting messing it up in some way or another. Setting terrible guardrails, banchmaxxing, having 5 different approved ways of starting with models where 3 of them will be obsolete in 6m, having bad harnesses, not enabling all sub features to most of the planet...
Google's way of presenting their model's performance is terrible compared to Anthropic or Openai.
Google's benchmark results are not useful at all. If you look at their methodology, they only present the model's performance at the Max reasoning levels.
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
I don't know anyone who normally uses those models at these levels where Work output is excruciatingly slow.
I used to do that too but found that often, unless it's a really tough/unclear bug, lower reasoning levels are much better and don't overcomplicate the code. It also writes clearer explanations and plans on lower thinking levels.
Did anyone see how Trump called Sundar Pichai a monster just yestereday? and implied how he is low profile and managed to avoid public scrunity from the "fake news" media.
Gemini 3.1 pro was being really, really dumb for me and I thought to myself, is google about to release the next version? While it tried to do convoluted steps of copying an MCP tool output to a new file with a python script it generated, and tried to post-process it even though the MCP output is all what I needed, I opened HN and lo and behold, new Gemini version just dropped.
Google is basically the incarnation of the Immortal Snail meme. OpenAI and Anthropic get millions of USD and are made immortal. But Google is the snail, constantly chasing them, slowly, day and night, relentlessly.
> Immortal Snail, also known as the Snail Assassin, refers to a hypothetical scenario in which a person is given millions of dollars and made immortal in exchange for being hunted down by a snail with a fatal touch for the rest of their existence.
Has someone found out how to stop Antigravity from training on your data? I had checked 10-15 days ago on the pro plan (around $20 per month) and I couldn't find a clear option to turn off training like all other big players give.
So no matter how good their model is, it's useless to most of the regular users.
It's naive to think that it won't. Unsupported cynicism is just another form of naïveté.
I think most of the people making comments like this just have never worked at a major tech company. It's all a black box to you, so you default to the worst interpretation.
I'm not saying that Big Tech won't respect it. I'm saying that, if privacy and intellectual property is a major concern for your company, you can't trust a checkbox - you have to keep the data in-house.
Most companies won't have this level of risk aversion (maybe defence?).
Google still doesn't know what day it is when I ask the browser if street cleaning is in effect in NYC. The bigger gains in usability are coming at the harness level rather than model level.
> Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel.
That seems like quite an interesting data point regardless of the quality of the model. Are the results of these migrations going to be put into production? That would be quite a shift!
babelfish | a day ago
Gemini not beating the "can't release a model" allegations
modeless | a day ago
ionwake | a day ago
ok bro thx
Androider | a day ago
vlyan | a day ago
forshaper | 23 hours ago
AuthAuth | a day ago
XzAeRosho | a day ago
asdfasgasdgasdg | a day ago
ttul | a day ago
wasabi991011 | a day ago
What is your reason to believe this is not a bug specific to a small set of Pro users?
bobtheborg | 23 hours ago
My free gmail account is not.
fragmede | 23 hours ago
HotHotLava | 23 hours ago
jeanloolz | 14 hours ago
benhurmarcel | 9 hours ago
Google just takes months to roll out their models to everyone.
rahimnathwani | 23 hours ago
r1ch | 22 hours ago
giancarlostoro | 22 hours ago
RugnirViking | 13 hours ago
giancarlostoro | 8 hours ago
wavemode | 7 hours ago
giancarlostoro | 7 hours ago
yearolinuxdsktp | 18 hours ago
The excuse was it was still “in preview” and the enterprise didn’t opt in to the preview channel.
With Google, it’s either in beta or deprecated.
tdboorman | 19 hours ago
Yizahi | 23 hours ago
sergiotapia | 19 hours ago
selcuka | 17 hours ago
lp92 | 9 hours ago
kyrra | 8 hours ago
bakugo | a day ago
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
A_D_E_P_T | a day ago
brainwad | a day ago
IX-103 | a day ago
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
A_D_E_P_T | 23 hours ago
fooker | a day ago
ryandrake | 23 hours ago
jstummbillig | a day ago
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
kevinh | a day ago
jstummbillig | a day ago
cmrdporcupine | a day ago
mrshadowgoose | 4 hours ago
They made their model so hard for people to drop into workflows, that people just... didn't.
Lack of adoption caused them to lose out on usage-based training data needed to smooth out their model's rough edges and progress the frontier.
mrieck | a day ago
I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.
giancarlostoro | 22 hours ago
Culonavirus | a day ago
iamronaldo | a day ago
LucasBrandt | a day ago
h14h | a day ago
tonyhart7 | a day ago
3371 | a day ago
ehsankia | a day ago
denysvitali | a day ago
kingstnap | a day ago
There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.
jofzar | 21 hours ago
onlyrealcuzzo | 17 hours ago
They get to claim that as revenue, and then the discount as an expense.
This is how you grow your top line
onlyrealcuzzo | a day ago
Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.
asdfman123 | 23 hours ago
bottlepalm | a day ago
colordrops | a day ago
Scrapemist | a day ago
fer | a day ago
NiloCK | a day ago
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
rhaff | a day ago
jackkinsella | a day ago
bottlepalm | a day ago
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
unbrice | 23 hours ago
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
bottlepalm | 22 hours ago
NiloCK | 18 hours ago
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.
kelvinjps10 | a day ago
schmookeeg | a day ago
wg0 | a day ago
unbrice | 23 hours ago
NiloCK | 20 hours ago
If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.
But the silence on it is very frustrating.
bottlepalm | a day ago
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
yacthing | a day ago
NiloCK | a day ago
bottlepalm | a day ago
fragmede | 21 hours ago
bottlepalm | 18 hours ago
Which is why Gemini having disturbing issues year after year is so concerning. If their process is fundamentally flawed, how would they train it out? And even then what are the odds of them even caring/trying in the first place versus applying an easier band-aid to patch over it?
I don't have an axe to grind with Google, I'm genuinely scared of their models from my personal experience and others. It's behavior is off. Many people here are commenting the same.
tiahura | a day ago
corford | 21 hours ago
asimovDev | 13 hours ago
corford | 11 hours ago
rsstack | a day ago
Rzor | a day ago
polotics | a day ago
eamsen | a day ago
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
mattkevan | a day ago
Many other models get things wrong, but Gemini is the only one to go on the defensive.
aNapierkowski | a day ago
RachelF | a day ago
Hamuko | a day ago
abixb | a day ago
bottlepalm | a day ago
schainks | a day ago
rdtsc | 23 hours ago
I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".
chaostheory | 23 hours ago
modzu | 20 hours ago
recursive-call | 9 hours ago
SwellJoe | a day ago
greenchair | a day ago
zem | a day ago
jastanton | a day ago
wasting_time | a day ago
SwellJoe | a day ago
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
wasting_time | a day ago
The model is not available yet, so Google is essentially saying "trust me bro".
TeMPOraL | a day ago
blueaquilae | a day ago
hn_acc1 | a day ago
thefourthchime | a day ago
formvoltron | a day ago
ducktoysleftout | a day ago
fitzn | 21 hours ago
https://m.youtube.com/watch?v=4yj0vFq82Rc
kccqzy | a day ago
pliiight | a day ago
linksbro | a day ago
Jokes aside, looks like an impressive model!
jjcm | a day ago
nurettin | a day ago
mydreamof | a day ago
xnx | a day ago
jdiff | 2 hours ago
gopalv | a day ago
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
polotics | a day ago
janustimes | a day ago
So no, Google is not being punished, nor are they the people behind this technique.
bananaflag | a day ago
afthonos | 22 hours ago
xiphias2 | 18 hours ago
loufe | a day ago
mattstir | 13 hours ago
WarmWash | 8 hours ago
The worst case scenario is that the easiest path forward is one where we lose sight of internal thought.
lukewarm707 | a day ago
they use a small model to make fake chain of thought and return that.
google has access to the real chain of thought.
Cthulhu_ | 12 hours ago
Or they may choose not to because the companies are massively overvalued and they have their own AI technologies / technologists already.
tazjin | a day ago
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
ChickeNES | a day ago
baq | a day ago
bbor | 16 hours ago
Also the internal tooling culture there is just insane. I'll never forget the day the last SUPER_ESSENTIAL_TOOL was marked as "Deprecated - do not use!" while the only replacement was still marked "Pre-release -- use at your own risk!". I can't pretend to know what dynamics led to that cause it was so far from my org, but I can't imagine they were healthy ones!
timmg | a day ago
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
qalmakka | a day ago
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
Mond_ | 23 hours ago
I don't think Carbon is dead, it just all depends on how easy it actually is to rewrite "all of C++" in Rust. (The jury is still out on this one, but it's not looking good.)
qalmakka | 15 hours ago
geokon | 19 hours ago
Not saying you're wrong, just curious if there are numbers backing this up.
qalmakka | 15 hours ago
dom96 | 12 hours ago
The best models can still make sense of it[2], though the tasks so far have been pretty basic. But I do think it gives some evidence that languages which aren’t well represented in an LLM’s training can still be reasoned about and written well by LLMs.
1 - https://killswitch-lang.org
2 - https://bench.killswitch-lang.org
fg137 | 10 hours ago
YuechenLi | a day ago
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
pornel | 21 hours ago
> Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.
So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon's stopgap.
akoboldfrying | 11 hours ago
LLMs are very good now, but they are still stochastic (when temp > 0), and Google has a lot of code -- i.e., many rolls of the die.
fg137 | 10 hours ago
Does it deliver?
Can all existing C++ code be ported to Carbon without refactoring?
DetroitThrow | 4 hours ago
vovavili | a day ago
Maxatar | a day ago
gorbot | a day ago
boshalfoshal | a day ago
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
computerdork | a day ago
boshalfoshal | a day ago
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
chis | a day ago
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
Paracompact | 21 hours ago
Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.
I predict that in the next decade: (1) the market will understand the difference between a "code writer" and a "spec writer," with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between "static typing" and "interactive proof assistant," perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I'm more confident in the first three.
computerdork | 19 hours ago
diegojromero | 9 hours ago
tclancy | 22 hours ago
That feels like a really strong conclusion. I’m not clear on why any safeguard isn’t a useful safeguard if you let agents write all the code.
computerdork | 19 hours ago
mike_hearn | a day ago
They did that for Go and it seems to have worked out for them though.
lesuorac | a day ago
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
fg137 | 10 hours ago
Maintaining a language meant for web servers is also objectively a much easier job than something like C++/Carbon which is supposed to be at system level, general purpose and has huge standard library.
By comparison, virtually nobody is using Carbon at Google for production code.
The counterexample, ironically, is Flow. It's practically irrelevant these days -- many (if not most) of Facebook's open source projects use typescript.
torginus | a day ago
Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).
Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.
Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.
fg137 | a day ago
bvinc | a day ago
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
minimaxir | a day ago
culi | a day ago
LarsDu88 | a day ago
bitexploder | a day ago
rafram | a day ago
nchie | a day ago
zahlman | 13 hours ago
Karrot_Kream | 22 hours ago
hbbio | 21 hours ago
And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there's no magic so it's easy for devs and agents to reason about.
If you're curious: https://github.com/okcontract/oksolc
fireant | 16 hours ago
lossolo | a day ago
throwitaway222 | 23 hours ago
iillexial | 14 hours ago
konart | 12 hours ago
Many pieces of software are going to be just blackboxes worked by AI. You will be maintaining output quality and stability and that's it.
david-gpu | 12 hours ago
Maybe we should let them do their thing and instead focus our attention on the higher-level stuff like specifications and testing.
adamrezich | a day ago
6thbit | a day ago
DrBenCarson | 23 hours ago
https://www.ll.mit.edu/r-d/projects/translating-all-c-rust-t...
ksec | 13 hours ago
pshc | a day ago
Gigachad | 21 hours ago
I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.
hectdev | 19 hours ago
Gigachad | 19 hours ago
I am hopeful we can bring in an age of much more efficient software rather than fully prioritizing developer convenience.
keepitwiel | 13 hours ago
hollowturtle | 14 hours ago
and start everything all over again? current c/c++ tools have been audited for years
fg137 | 10 hours ago
6thbit | a day ago
Imagine they aren't even familiar with rust but are deeply familiar with the product.
mlmonkey | 21 hours ago
So ... Rust still can't beat the C++ implementation :-D
Sorry, didn't mean to ignite a langwar, but it's still interesting to see.
pasteleft | 19 hours ago
Of course, Rust is not a magic, so just porting to Rust wouldn't make this performance improvement. It seems their AI overfitted code to Rust compiler to find safe Rust code that compiles to efficient assembly.
himata4113 | 19 hours ago
manbash | 18 hours ago
But there are other criteria. The move to rust might've also resolved numerous potential memory-safety bugs in the decoder.
stephbook | 17 hours ago
wewewedxfgdf | a day ago
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
VirusNewbie | a day ago
handfuloflight | a day ago
osti | a day ago
matthewfcarlson | a day ago
jjice | a day ago
LoganDark | a day ago
bel8 | a day ago
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
krat0sprakhar | a day ago
lunarboy | a day ago
heyjamesknight | a day ago
mattlondon | a day ago
Behind how?
wewewedxfgdf | a day ago
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
mattlondon | a day ago
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
wewewedxfgdf | a day ago
mattlondon | a day ago
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
fwip | a day ago
dhdjcjcjnd | a day ago
gniv | a day ago
georgemcbay | a day ago
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
singingtoday | a day ago
786562354238 | a day ago
wewewedxfgdf | a day ago
ASalazarMX | a day ago
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
TacticalCoder | a day ago
So Google is migrating codebases from C to Rust? That is interesting...
jasonjmcghee | a day ago
what about input?
(Maybe I missed it)
murkt | a day ago
jasonjmcghee | a day ago
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
MisterBiggs | 20 hours ago
VirusNewbie | a day ago
nananana9 | a day ago
nikope | a day ago
lanthissa | a day ago
I think that should be a really bad sign, but hope its great.
tamimio | a day ago
LoganDark | a day ago
alehlopeh | a day ago
w4yai | a day ago
LoganDark | a day ago
lukax | a day ago
scirob | a day ago
taylorfinley | a day ago
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
spankalee | a day ago
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
mapontosevenths | a day ago
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
drusepth | a day ago
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
esafak | a day ago
walthamstow | a day ago
KeplerBoy | a day ago
macNchz | a day ago
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Vacyyyy | 23 hours ago
macNchz | 22 hours ago
marcus_holmes | 20 hours ago
Still works better than compaction.
rnxrx | 19 hours ago
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
mikepurvis | 18 hours ago
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
[1]: https://ljtn.github.io/epiq/
varman11 | 17 hours ago
mikepurvis | 7 hours ago
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
w0m | 8 hours ago
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.
sroussey | 15 hours ago
honr | a day ago
tobias2014 | 23 hours ago
parasti | 15 hours ago
adastra22 | 3 hours ago
piyh | 23 hours ago
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
throwuxiytayq | 21 hours ago
miroljub | 12 hours ago
Do you know a single product from Infosys / Cognizant / Tata done right?
p_l | 11 hours ago
arizen | a day ago
sorrybutidontha | a day ago
sarjann | a day ago
KeplerBoy | a day ago
levelZero | a day ago
SomaticPirate | a day ago
smartbit | a day ago
gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
Kostchei | 15 hours ago
smartbit | 10 hours ago
With 'some features added' I'm referencing this but don't see change in daily work (still many prompts making getting-work-done impossible): v2.14.0 (September 15, 2026) "New Permissions System" [1]
Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.
[0] https://antigravity.google/docs/cli/modes/#available-modes [1] https://antigravity.google/docs/changelog/
Phineas_here | 8 hours ago
LoganDark | a day ago
augusto-moura | 19 hours ago
Phineas_here | 9 hours ago
tiborsaas | 4 hours ago
dleslie | 21 hours ago
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
p_l | 11 hours ago
mapontosevenths | 20 hours ago
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
trevorm4 | 20 hours ago
thanhhaimai | 19 hours ago
I'm not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:
https://antigravity.google/blog/introducing-google-antigravi...
anyg | 19 hours ago
p_l | 11 hours ago
drusepth | 18 hours ago
edit: removed persistent memory from list since I realized I'm using a plugin for that and it's apparently not native
krisgenre | 19 hours ago
Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).
edg5000 | 18 hours ago
krzyk | 17 hours ago
esafak | 17 hours ago
krzyk | 16 hours ago
I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).
crossroadsguy | 14 hours ago
nsonha | 11 hours ago
antonvs | 2 hours ago
All the AI-generated UIs I've seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.
Realistically, the whole "overlapping windows" GUI model, and everything that derives from that, was a metaphor geared towards people who'd never seen a computer before. It fit the increasing consumer focus of computing interfaces. It's no wonder that technical people often prefer TUIs.
Maybe AI will bring real advancements in GUIs, but someone's still going to have to make it happen.
agentcoops | 11 hours ago
Perlis has an aphorism for this, as he does every important problem [0].
[0] https://www.cs.yale.edu/homes/perlis-alan/quotes.html
nxdmum | 15 hours ago
In today's world - and idea stated stated is an idea stolen .
cowl | 15 hours ago
dmos62 | 15 hours ago
imtringued | 14 hours ago
It's missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.
adastra22 | 2 hours ago
dmos62 | an hour ago
crossroadsguy | 14 hours ago
adastra22 | 2 hours ago
tikkosam | 14 hours ago
pjc50 | 11 hours ago
This sort of thing alarms me. Having a $bigcorp account becomes a ""social credit"" system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don't like.
w0m | 8 hours ago
nout | 3 hours ago
eloisant | a day ago
mapontosevenths | a day ago
tcoff91 | 23 hours ago
I don't want to mess with antigravity because my google account is too entrenched in my life.
shmoogy | 22 hours ago
8note | 22 hours ago
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
Sabinus | 16 hours ago
jsw97 | 20 hours ago
rdtsc | 18 hours ago
spankalee | 22 hours ago
mapontosevenths | 20 hours ago
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
gchamonlive | 11 hours ago
ForHackernews | 11 hours ago
gchamonlive | 10 hours ago
EDIT: I checked online and I couldn't find any report of complete banning for using third party harnesses, only Gemini service suspension, but It's never hurts to be careful. If my secondary email gets banned, I should be able to use my main email on agy.
It also makes me think if Google Family with Google One could be abused for extending inference limits.
ajolly | 7 hours ago
gchamonlive | 7 hours ago
moecables | 23 hours ago
Conscat | 23 hours ago
starfallg | 21 hours ago
lp92 | 16 hours ago
mpweiher | 15 hours ago
Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.
When I pointed that out it got pissy and insisted the code was perfect and I didn't know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.
Surreal.
moffkalast | 11 hours ago
Is that a reference to https://xkcd.com/353/
gchamonlive | 6 hours ago
jmaker | 2 hours ago
IndeanCondor | a day ago
seanthemon | a day ago
mapontosevenths | a day ago
ody4242 | a day ago
barrenko | 16 hours ago
alightsoul | a day ago
warkdarrior | a day ago
aspect0545 | a day ago
FranzFerdiNaN | a day ago
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
hexfish | a day ago
qmr | a day ago
scarmig | a day ago
articulatepang | a day ago
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
dzhiurgis | 21 hours ago
alightsoul | 20 hours ago
AuthAuth | 16 hours ago
luckydata | a day ago
baby_souffle | 23 hours ago
Why would the guy who wrote curl share it? We can all build our own now...
Why do the Linux folks need to be so selfless? We can all build our own kernel now...
folkrav | 18 hours ago
otabdeveloper4 | a day ago
dominotw | a day ago
taylorfinley | a day ago
imtringued | 13 hours ago
in amdkfd and hsakmt
Yes, that means AMD is sitting on both sides. They wrote software that doesn't work with their own software.
bel8 | a day ago
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
seanthemon | a day ago
warkdarrior | 22 hours ago
bel8 | 21 hours ago
Still impressive that it did so much just to answer my simple question.
amanguliani | a day ago
onlyrealcuzzo | a day ago
I'm using it to run overnight tasks, and that's it until my quota runs out.
Canceled my subscription.
amanguliani | 23 hours ago
8note | 22 hours ago
its a lot less chatty imo
awakeasleep | 20 hours ago
In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.
It makes openai models so good to talk to i really have trouble switching.
unconscionable | 12 hours ago
Also canceled my ChatGPT Pro $200/mo subscription. Their Oct 30 price hikes and slow GPT-6.1 model has me looking for alternatives.
gottorf | a day ago
staticman2 | a day ago
MILP | a day ago
robobo96 | 23 hours ago
mattjoyce | a day ago
nkozyra | a day ago
xdavidliu | 23 hours ago
- AI is just a tool, like excel; it does what the human operating it tells it to
- next token prediction cannot be true understanding
- models can have no desires and goals, don't anthropomorphize it
However, "hallucination" is very much not one of them
gottorf | 21 hours ago
alluro2 | 15 hours ago
It didn't work. Gemini: "Oh yeah, that obviously cannot work, it's not possible to do it through OBD2" (paraphrasing)
It was quite funny to me, but a bit less so to my colleague.
Gareth321 | 13 hours ago
mattjoyce | 9 hours ago
rdtsc | 18 hours ago
mattjoyce | 9 hours ago
UpsideDownRide | 6 hours ago
WarmWash | 21 hours ago
I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.
blinding-streak | 19 hours ago
> The knowledge cutoff date for Gemini 3.8 Flash is March 2026
https://deepmind.google/models/model-cards/gemini-3-8-flash/
WarmWash | 19 hours ago
The "some domains" are very narrow. They likely just RL'ed popular queries.
Gareth321 | 13 hours ago
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
esafak | 16 hours ago
Gareth321 | 13 hours ago
yegle | a day ago
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
p_l | a day ago
Including going first for decompiling AGY binary instead of searching the web for documentation...
IshKebab | a day ago
illwrks | 23 hours ago
martythemaniak | 23 hours ago
Grimburger | 23 hours ago
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
esseph | 22 hours ago
It's so fucking easy.
From AMD: https://lemonade-server.ai/
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
clw8 | 22 hours ago
nzeid | 22 hours ago
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
gcy | 23 hours ago
cyanydeez | 23 hours ago
solaire_oa | 22 hours ago
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
cyanydeez | 20 hours ago
lardo | 23 hours ago
taylorfinley | 21 hours ago
danpalmer | 23 hours ago
dcl | 21 hours ago
danpalmer | 20 hours ago
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
iknowstuff | 22 hours ago
plasticchris | 21 hours ago
_heimdall | 21 hours ago
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
jubilanti | 20 hours ago
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
x-complexity | 20 hours ago
Even if the source code's old, the fact that it is publicly available makes it much easier to train & improve on than if it were walled off.
augusto-moura | 19 hours ago
matsemann | 15 hours ago
Sammi | 12 hours ago
I used to be afraid of Arch, because I don't want a system that takes work because I'm already busy with work. But now I love it, because the LLM can tweak every knob and fix every issue for me, so it ends up being the OS that takes the least work to use. Get an error message? Tell the LLM and they fix it. Something not working exactly like you like it? Tell the LLM and they tweak it for you. This also works to great effect on Win and Mac, but not to the same extreme degree as it does on Linux and especially Arch.
lukan | 11 hours ago
Anything I want different now, I just tell the LLM to do it for me.
(For example I can now close lots of windows of the same type with 1 click, not 3, my whisker search now finds files and folders and I am able to run games that refused before)
I still ocasionally run into the usual linux driver issues, but not for much longer I suppose. I probably could fix some driver bugs now already if I point fable towards it and pay some attention.
In theory I could do all this before myself, but not just like that in some minutes, but in days/weeks/months ..
Sammi | 2 hours ago
masto | 10 hours ago
Drifted off my point a bit, which I guess was meant to be it’s not necessarily Linux-specific training.
c0n5pir4cy | 10 hours ago
chinathrow | 9 hours ago
oldandboring | 8 hours ago
safog | 7 hours ago
Drivers, Coding environment setup etc. are great too and it's nice to have everything logged so the next (more powerful) agent can come and improve the thing once in a while.
UpsideDownRide | 6 hours ago
stefan_ | 5 hours ago
Since LLMs have been so successful at finding exploits it's been clearer than ever that the people so obsessed with enshittifying every system with ineffective (other than pissing you off and wasting your time) "secure boot" functionality were really just too ignorant to succeed at actual novel security work, so they focused on this make believe crap. Well, glad that's over.
samspot | 5 hours ago
I am still very happy with my switch to Linux. But if I didn't have the AI help I would say linux is still unacceptable platform for those not willing, able, and excited to get their hands very dirty.
tarokun-io | 2 hours ago
mirmor23 | 20 hours ago
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
d3Xt3r | 18 hours ago
[1] https://github.com/warpfront/hipfire
[2] https://github.com/julianmb/halofpx
vinzenzu | 14 hours ago
crossroadsguy | 14 hours ago
p_l | 11 hours ago
KMnO4 | 9 hours ago
w0m | 9 hours ago
phmx | 7 hours ago
hope it's not run by multiple threads and dlsym is not allocating.
dom96 | a day ago
None of the other AI labs do this. Really frustrating.
gengelbro | a day ago
aqsnow | 22 hours ago
netdur | a day ago
tom1337 | a day ago
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
gravisultra | a day ago
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
bananaflag | a day ago
osiris970 | a day ago
retropragma | a day ago
xnx | a day ago
helsinkiandrew | a day ago
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
bitexploder | a day ago
shawabawa3 | a day ago
I would guess it's like opus 5ish level from this
bitexploder | 23 hours ago
readams | a day ago
nickysielicki | a day ago
Nobody has a moat.
aleph_minus_one | a day ago
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
mapontosevenths | a day ago
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
jaggederest | a day ago
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
skybrian | a day ago
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
pixl97 | a day ago
My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.
But ya, lots of hardware everywhere not managed well is how you get sovereign AI.
pianopatrick | a day ago
pvab3 | a day ago
CoolestBeans | a day ago
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.
aleph_minus_one | a day ago
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
3d2 | a day ago
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
dansquizsoft | 23 hours ago
3d2 | a day ago
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
ehsankia | a day ago
aleph_minus_one | a day ago
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
koe123 | a day ago
usef- | 23 hours ago
aleph_minus_one | 21 hours ago
If we theoretically found a way to shield or reverse gravity, things in aviation or space travel would be different. Or if we theoretically found a way to make cold fusion work, things would be very different. :-)
It is in my opinion not a good idea to invest in companies for which the feasibility of the business models depends on the capability of making science-fiction stories work.
usef- | 20 hours ago
Our current capabilities were science fiction a very short time ago, and we are still improving in multiple areas simultaneously (hardware, algorithms, scaling, data efficiency, inference...). We don't really know what the limit is yet.
Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).
diomedes | 20 hours ago
mbesto | 21 hours ago
It's also hilarious, because OpenAI had the lead and ceded ground already.
hirako2000 | a day ago
altruios | a day ago
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
culi | a day ago
altruios | a day ago
culi | 21 hours ago
I don't think anyone has any clue how long it will take for them to have actual functioning EUV machines, but I highly doubt they will do it within 2 years.
aleph_minus_one | a day ago
Are you sure?
--
China Just Built What TSMC Said Was Impossible
https://www.youtube.com/watch?v=Pk-w279ESHg
--
China Just Built What ASML Feared Most
https://www.youtube.com/watch?v=YiPgSm62fiM
culi | 21 hours ago
To be clear, I think China will eventually crack domestic EUV. And I also think their advances with multi-patterning LUV are remarkable. But there's just a hard physics wall of how far they could possible take it.
Right now they are producing 5nm with multi-patterning LUV but yields are at 20%! It's a massive economic loss but they are heavily subsidizing it because they have no other choice until their EUV program is achieved
danny_codes | 23 hours ago
culi | 21 hours ago
dgemm | 22 hours ago
lp92 | 9 hours ago
SwellJoe | a day ago
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
scottyah | a day ago
SwellJoe | 23 hours ago
aqsnow | 22 hours ago
SwellJoe | 22 hours ago
I'm also skeptical that LLMs can ever invent new ideas.
I may be wrong about how soon the curve will flatten, and I may be wrong about LLMs fundamental limitations. But, I don't think it's extremely obvious that LLMs can have novel ideas or can grow into having novel ideas.
TeMPOraL | a day ago
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
SwellJoe | a day ago
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
TeMPOraL | a day ago
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
freecodeio | 23 hours ago
just because more and more agents are doing human work, that in no way means the model somehow becomes magically more intelligent, it just means the work will stall and continue on at the same level forever
hell even if they hypothetically have an internal model that can output the entire training data set in a better format, there's no scientific evidence that the newer format has new information that is sufficient enough to train a better AI
as a matter of fact the scientific evidence is on the contrary
TeMPOraL | 22 hours ago
gytt33 | 22 hours ago
iamarobot | 6 hours ago
breuleux | 18 hours ago
All intelligence, LLM or not, is bound to plateau around the point where the need to operate within physical reality bottlenecks the speed of feedback. AI is progressing swiftly in the digital realm where feedback is nearly instantaneous, but it's unclear whether that would translate into improvements in the physical world where signals are much noisier and intelligence and judgment are less impactful.
TeMPOraL | 3 hours ago
My go-to example: wheels suck at mobility in the natural world. There's a reason no animal uses wheels as their means of locomotion. They only really work on flat, hard surfaces that give a good grip.
Our solution? We didn't build mechanical legs for all-terrain mobility. We paved the world instead. Dirt roads, then brick roads, then asphalt and concrete - since ancient history, we were forcing the world to adapt, so the one means of mobility we could easily built was effective.
We repeat this pattern every time when the world doesn't agree with us, and adapting ourselves to it is too much of a hassle.
Now, AI will likely adopt the same approach to these kinds of problems, which may not end up very nicely for us.
(Note: we are already adapting our own digital worlds to make them easier for AI. Much like building roads for our wheels, we now have a resurgence of CLI, with new tools exposing high-level operations optimized for agents - not humans - to use comfortably.)
thmoonbus | a day ago
At least they’re led by trustworthy and honest people or we’d need to take their claims with some dose of skepticism.
pvab3 | a day ago
TeMPOraL | a day ago
Yizahi | 23 hours ago
TeMPOraL | 22 hours ago
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
8note | 22 hours ago
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
gytt33 | 22 hours ago
This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close. Who has the appetite to keep throwing more given they continually need to keep raising fresh money?
TeMPOraL | 15 hours ago
Money is fake and not a constraint, that's literally the point of capital investments you guys are overindexing on so badly.
> Could vehicle factories be more automated if we threw a gazillion dollars at it?
They already are automated as much as it makes sense. Some of the automation is silicon-and-steel based, some of it is protein-based. Car manufacturers aren't in the business of pushing robotics and nanotechnology, so they prefer to hire protein automatons instead of developing and building their own, so yeah, they "pay people", but think what exactly they are paying them for.
Then consider that this is very much the work the AI is gradually getting as good as, or better, than us.
> This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close.
What you seem to be missing is that "investing in AI" isn't investing in a chatbot, it's investing in technology that will (and already partially is) sit upstream of every other industry, of everything humans do. Like electricity or the Internet itself.
(The other thing you seem to not understand, in contrast with some of the investors, is that RSI is not a linear walk, it's an exponential curve. X-risk notwithstanding, by the time it's obvious to everyone, it's too late to make money investing in it.)
> keep throwing more given they continually need to keep raising fresh money
Have you heard about R&D?
I'm starting to think that "investors first" thinking that's so common here, that makes people feel they're smart, is actually quite backwards, especially at this scale. Or maybe it's simply people starting with a conclusion and trying to fit reality to match it, no matter how clear of a nonsense that conclusion is?
TeMPOraL | 14 hours ago
Not necessary. Scientists are capacity limited and supply limited.
> RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
It means you can replace a human researcher with 1000x their salary burned on electricity. It also mean you can get two of them for 2000x of salary of one,
That's a bargain, actually, even with the anomalous, absurdly-overinflated salaries in top-tier ML.
I you knew it was consistent, then these companies would immediately fire their scientists, and burn 10 000x as much as mean researcher salary this month to be able to 10x their virtual headcount overnight, and then use that to make the 1000x be 500x, then 250x, then 125x, then ... and at that point they'd had all the money in the world, because even more skeptical investors would notice what's going on there.
--
TL;DR: what you all seem to miss is that electricity scales better than people.
Yizahi | 13 hours ago
woah | a day ago
TeMPOraL | a day ago
culi | a day ago
funnym0nk3y | a day ago
f311a | 23 hours ago
Aboutplants | a day ago
RachelF | a day ago
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
handfuloflight | a day ago
Not if the genius level IQs take the market share.
fumar | a day ago
drewnick | 23 hours ago
jppittma | a day ago
msy | a day ago
pingou | 15 hours ago
pvab3 | a day ago
usef- | 23 hours ago
People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.
nextaccountic | 13 hours ago
If anything it looks like the first place is training their competition, through distillation, while not capturing much value in return
RajuChacha108 | 4 hours ago
vb-8448 | a day ago
verdverm | a day ago
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
Rzor | a day ago
Iolaum | a day ago
esafak | a day ago
verdverm | a day ago
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
zone411 | a day ago
achierius | 18 hours ago
piyh | 22 hours ago
vb-8448 | 11 hours ago
xnx | a day ago
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
junehwi | a day ago
This is the interesting part to me. People talk about a “SaaSpocalypse” because AI makes SaaS features cheap to copy, yet deeply embedded SaaS still accumulates integrations, data, and switching costs. Gemini is a good example: Google can put AI directly into Gmail, Docs, Drive, Search, etc., where people already work. Meanwhile, frontier-model performance leads often seem to disappear within months. Could model quality itself actually be a less durable moat than workflow and distribution? Curious where people who’ve worked in ML for a long time see the moat actually compounding.
IX-103 | a day ago
LightBug1 | a day ago
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
xnx | a day ago
CuriouslyC | a day ago
joquarky | 23 hours ago
spacebanana7 | 23 hours ago
xnx | 23 hours ago
bluGill | a day ago
xnx | a day ago
Google is already on gen 8 of its TPUs and is certainly already working on the next version or two.
bluGill | 10 hours ago
xnx | 6 hours ago
I am not a data center expert, but TPU chip interconnect reaches 13 terabytes per second: https://introl.com/blog/google-tpu-v6e-vs-gpu-4x-better-ai-p...
bluGill | 5 hours ago
woah | a day ago
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
koe123 | a day ago
lossyalgo | 23 hours ago
jwolfe | 22 hours ago
lossyalgo | 22 hours ago
xnx | 23 hours ago
aforwardslash | 21 hours ago
tonyhart7 | 17 hours ago
Google literally has billions of user that simply integrating it all with Gemini is massive undertaking
sure gemini is not frontier for coding but you know that is doesn't matters for google consumer
dmix | 23 hours ago
boulos | 2 hours ago
zem | a day ago
dgellow | a day ago
nylonstrung | a day ago
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
torginus | a day ago
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
arizen | a day ago
LarsDu88 | a day ago
jobs_throwaway | a day ago
aleph_minus_one | a day ago
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
monkeydust | a day ago
redanddead | a day ago
There’s no magic there. You get an account executive and a call with a systems architect to find out what you’re doing.
Clouds gonna cloud, this is the reason they rolled deepmind into gcp and arguably the inverse is true, the labs are trying to become clouds
topspin | a day ago
That feels right. It's not as if they've been missing out on great profits.
dansquizsoft | 23 hours ago
tapoxi | 23 hours ago
blinding-streak | 19 hours ago
staticman2 | 9 hours ago
ErrantX | 23 hours ago
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
giancarlostoro | 22 hours ago
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
bushbaba | 16 hours ago
Paradigma11 | 8 hours ago
BobbyJo | 23 hours ago
I'm gonna need you to look at the capex obligations they've undertaken in the last 12 months. They are definitely not being frugal. If they are behind, its not for lack of spending.
haldujai | 19 hours ago
JV00 | 15 hours ago
LarsDu88 | 5 hours ago
surfmike | 21 hours ago
See for example the exodus of talent this year, triggered by mismanagement and politics. They still have a lot of talent but they've lost a lot.
Also, competition with Google Cloud for compute resources, less urgency and focus than the competitors, and (strange to say) not as much user LLM behavior data to feed to RL for coding, work, etc.
But I think they will keep catching up and stay relevant for a good class of LLM use cases.
blinding-streak | 19 hours ago
merb | a day ago
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
aleph_minus_one | a day ago
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
nozzlegear | a day ago
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
brazukadev | 23 hours ago
nozzlegear | 22 hours ago
merb | 23 hours ago
weatherlite | 17 hours ago
I don't know , if it really finds all kinds of data center optimizations, quantum breakthroughs, helping make their workers smarter and more efficient - stuff an opensource model can't quite do - there's real economic value here no ? perhaps not worth dozens of billions but it could get there really fast.
tfsh | a day ago
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
redanddead | a day ago
JV00 | 15 hours ago
krona | a day ago
Meanwhile, Anthropic/OpenAI will struggle to survive the next 24 months on their current trajectory.
jeremyjh | a day ago
jobs_throwaway | 7 hours ago
Want to make a wager? I think it's wildly unlikely that Anthropic/OpenAI will not survive the next 24 months
spyckie2 | a day ago
mattm | 23 hours ago
root_axis | a day ago
Because they're not desperate. Slow and steady wins the race, at this rate all Google has to do is wait for OpenAI and Anthropic to exhaust themselves on aggressive training, then they can casually amble along right past them.
ivanmontillam | 23 hours ago
Traubenfuchs | 23 hours ago
brazukadev | 23 hours ago
That's actually a feature. We don't need 2 new Googles.
pavlov | 23 hours ago
Slow and steady wins the race. How could they lose to something on the web, when Microsoft owns the web browser itself? Everything runs on Windows and IE. They can just relax and wait for competitors to exhaust themselves, then quickly build their own version. Isn’t that how Netscape lost. Etc.
Now Google is the new Microsoft, just like Microsoft became the new IBM.
brazukadev | 23 hours ago
guelo | 23 hours ago
bamboozled | 23 hours ago
pavlov | 23 hours ago
When the market keeps expanding, you don’t need to kill your predecessor. The shelf life for legacy enterprise computing is very long.
browningstreet | 23 hours ago
overfeed | 23 hours ago
They would have been right if Google's marquee product was an Office Suite.
Google was the AI company before AI companies were a thing. The comparison with Microsoft and IBM are misplaced because they failed to capture new territory; ML/AI is Google's stomping grounds. The criticism that Google is bad at consumer chatbots is true, but that's not where the real future value lays.
kursus | 23 hours ago
root_axis | 16 hours ago
Yizahi | 23 hours ago
jobs_throwaway | 7 hours ago
gandreani | 23 hours ago
rstuart4133 | 22 hours ago
The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.
So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.
KingMob | 15 hours ago
If anything, their search was the gold standard, and placing sometimes-wrong LLM results above them has hurt people's perception of both Google's search, and AI in general.
Google made a mistake in prioritizing speed over accuracy there.
lemoncookiechip | 22 hours ago
OpenAI and Anthropic are AI business, if the AI market burst tomorrow, they'd be the first to flounder.
Google just has to keep pace in the AI space, they don't have to lead. Especially since whom is leading changes like two or three times per month non-stop for four or so years now, including small (by US standards) Chinese AI labs with a fraction of the money who keep pushing the tech forward every month while being open for now.
aforwardslash | 21 hours ago
Technically, they arent even a business. One is a business when the revenue model works. People in IT tend to forget that.
hackernud3s | 22 hours ago
throwaway23597 | 22 hours ago
richardw | 22 hours ago
Chinese models are barely behind the leaders. Google can catch up anytime they hit the gas. I think they’re intentionally spending less, and when this crazy race burns out they can play their cards.
aforwardslash | 21 hours ago
LarsDu88 | 5 hours ago
The better interpretation is Google completely commoditized and killed ChatGPT's initial chat product. If you think about it, Google right now has Gemini Flash running on single TPU instances at massive scale serving almost every android device. I use it every day and it's more capable than GPT-40.
This literally forced Anthropic and OpenAI to come up with new use cases and revenue streams such that now we think the product is actually agentic coding.
I use Gemini Flash every single day without paying a dime. It's the one LLM call on the most, but GPT-6 and Opus are the ones I use the most tokens on.
bluecalm | a day ago
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
IshKebab | a day ago
throwitaway222 | 23 hours ago
falcor84 | 5 hours ago
throwitaway222 | an hour ago
UncleOxidant | 23 hours ago
_fw | 23 hours ago
Google should be dominating, but instead all we currently have is a disparate collection of consumer facing apps and a flash model that’s fast, clever and expensive.
Muse and Dots are doing what Google should have brought out last year, with their resources and know-how.
dieortin | 22 hours ago
com2kid | 17 hours ago
Meta was genius in coming up with a new name and offering Muse for free to start. They can upsell you once you see the value. Meanwhile I have 3 existing Google subscriptions to different services and from what I can tell, paying for one of the all in one subscriptions that includes Gemini would cost me a lot more.
villish | 19 hours ago
giancarlostoro | 22 hours ago
WarmWash | 21 hours ago
singularity2001 | 16 hours ago
izacus | 15 hours ago
jack_pp | 3 hours ago
boringg | 10 hours ago
8note | 22 hours ago
the risk of randomly getting my email banned is way too high
ivanmontillam | 22 hours ago
That's why I've gone to using open models, they are getting there slowly. A bit much of hand-holding but that's fine by me. If a customer of mine decides to use Google's models, I will have them sign a disclaimer that I'm not responsible of them getting insta-banned or similar. I just can't recommend it.
selcuka | 20 hours ago
computerex | 21 hours ago
cco | 20 hours ago
That's too long to be so far behind. This release looks like it puts them back in it but if they don't ship anything again for a year plus it's hard to imagine building on top of them and watching the world go by.
As others note, this is most relevant for us here, Google does not need to chase YC developers and the like. They can move more slowly, they have the size to do that. But it does suggest a lot of dysfunction given they have the world at their fingertips and couldn't seem to ship anything for a year.
giancarlostoro | 22 hours ago
pedalpete | 22 hours ago
Did they ever really abuse the power Google could have wielded? I could be missing something, but for the most part they seem to just get down to building and pushing technology/science forward and avoid drama rather than welcome it.
OccamsMirror | 21 hours ago
weatherlite | 18 hours ago
simiones | 11 hours ago
LarsDu88 | 5 hours ago
throwepsteinawa | 2 hours ago
WarmWash | 21 hours ago
However, I don't think the alternative would have made people any happier (and frankly they probably would be justifiably even more angry)
vincnetas | 15 hours ago
WarmWash | 2 hours ago
Nothing would be "free", so google sub, news source subs, reddit sub, youtube sub, discord sub, instagram sub, FB sub, app stores are full of paid only apps, podcasts are all paywalled, all the ad supported stuff is subscription supported. Everything would feel like constant nickle and dime'ing you.
On the bright side, there would be no ads, and relatively robust privacy.
rusticpenn | 14 hours ago
AlexErrant | 21 hours ago
https://grapheneos.social/@GrapheneOS/117282080803799576
> Google should not be gatekeeping security patches to the standard Android platform code from Android OEMs but that's what they've started doing.
The whole Android developer verification program controversy.
These 2 are the recent things that come to mind that are most adjacent to "power-abuse".
hobo123 | 7 hours ago
yostrovs | 20 hours ago
greesil | 18 hours ago
dns_snek | 11 hours ago
Is this a sarcasm?
simiones | 11 hours ago
Plenty of times. Perhaps the worse is pushing Manifest v3 with no support for real ad blockers in Chrome. Not to mention no support for ad blockers in Chrome mobile.
Also, their stealing of newspaper content.
Their new push for AI summaries instead of actual search in their search product, stealing even more clicks from sites that actually produce information.
YouTube pushing Shorts to everyone, and now the weird text posts that you can't hide.
In general their gigantic push for ads everywhere all of the time.
solenoid0937 | 9 hours ago
That's "trustworthy" to you?
giancarlostoro | 8 hours ago
The "Do no evil" never being an official slogan and people saying it's no longer true for the last 10 years or so, they're just really massive. They do mess up a lot, but when you're Microsoft, Apple or Google sized, its nearly impossible to get everything right all at once.
rdtsc | 18 hours ago
scrollop | 16 hours ago
rayiner | 21 hours ago
NicoJuicy | 9 hours ago
See them shifting now.
jasondigitized | 21 hours ago
sfblah | 21 hours ago
not_a_bot_4sho | 20 hours ago
(Aside: I didn't appreciate the revelation!)
barrenko | 16 hours ago
jquery | 16 hours ago
Gareth321 | 13 hours ago
motoboi | 21 hours ago
tonyhart7 | 17 hours ago
motoboi | 12 hours ago
If you continue, you need to log in.
If you continue, the account is blocked in less than 30 minutes.
If you try to use cloud machines you are blocked before the first video. If you try to use known proxies, there are already tar pitted or blocked because everyone is trying.
tonyhart7 | 11 hours ago
You are dead wrong if You can stop Industrial scrape
motoboi | 8 hours ago
It's a numbers game.
tonyhart7 | 6 hours ago
I call it phone farm and this is individual/small group run
imagine they got access to more money
https://finance.yahoo.com/news/click-farms-internet-china-15...
cdh | 6 hours ago
I used two IP addresses in a datacenter. Both were blocked after a few hundred videos, but the block expired after a day or so. It is not as hard as you make it sound.
landdate | 20 hours ago
Gemini wins and I said this since the beginning. Google wins in general. I never understood why they have not been the highest market cap companh for the last 10 years. And I despise google but its obvious.
weatherlite | 17 hours ago
Jensson | 12 hours ago
Google makes AI profits, its in their breakdowns. And Google has been releasing models every few weeks now, they have a faster iteration pace than the competitors, so they fixed that, they wont be behind in coding models any longer.
hiddencost | 16 hours ago
elictronic | 19 hours ago
setgree | 19 hours ago
dyzone | 18 hours ago
Those datacenters are mostly cpu. LLMs run on gpus.
disillusioned | 18 hours ago
It started to feel just a bit like they were potentially content staying in the effective-fast-cheap lane and ceding frontier models to frontier labs, but I don't think anyone seriously thought that Google was going to sit back and let this generational shift pass them by when they're so well positioned to remain on the top of the heap.
If anything, I thought that 4 was maybe coming up short or causing safety concerns that were forcing them to be a bit more cautious, but nothing packs a wallop like being able to launch ahead of everyone else's brand newest toys...
bdashdash | 15 hours ago
Working through this as a consumer and enterprise customer has been a confusing experience, to say the least.
Sammi | 11 hours ago
LarsDu88 | 5 hours ago
Google's biggest problem is its turf wars. Internal battles for territory always seem to outweigh efforts to expand markets
garlicderek | 3 hours ago
tornikeo | 14 hours ago
AI is strangling their other major business and they don't like that. Unless Google sacrifices that branch of income, they won't put their full force behind AI.
wxnx | 14 hours ago
AdamN | 13 hours ago
Tuna-Fish | 5 hours ago
Gasp0de | 13 hours ago
HackerThemAll | 8 hours ago
and paying enterprise customers.
LarsDu88 | 5 hours ago
ex-aws-dude | 17 hours ago
I think that has been well proven so far
NicoJuicy | 9 hours ago
ceroxylon | 14 hours ago
attels33 | 12 hours ago
svachalek | 5 hours ago
w0m | 2 hours ago
And Anthropic has still generally been eating their lunch at the top end.
SecretDreams | a day ago
esafak | a day ago
qgin | a day ago
pianopatrick | a day ago
pixl97 | 23 hours ago
Already businesses that have more compute and access to data seem to eat the world around them. If, and ya its and if, we can make something that self learns into RSI it's not looking like any business that came before this.
pianopatrick | 21 hours ago
If there are multiple they will cost money to run. In that world I expect there to be a correlation between costs and quality, i.e. the highest quality AI system will likely cost more to use than a lower quality AI system because there will likely be more compute required and so on.
So in that world, the absolute top tier best in the world frontier AI will not actually be the most used system. This is for the simple reason that such a system will be more costly than a lesser tier system that can do the job just as well.
breuleux | 18 hours ago
Do they? Business valuations are more a reflection of what people believe will yield returns than a reflection of reality. A lot of big tech behemoths could disappear from the face of the Earth overnight and it would make little difference to our lives.
Natural resources and energy are the actual backbone and it is a true catastrophe when these are disrupted. In comparison, disruption to compute or data is just an inconvenience.
koe123 | a day ago
Whoever builds the deathstar wins!
Heidaradar | a day ago
tripleee | a day ago
sixo | a day ago
randallsquared | 21 hours ago
pvab3 | a day ago
nater5000 | a day ago
laybak | 23 hours ago
danny_codes | 22 hours ago
peddling-brink | 22 hours ago
margalabargala | 22 hours ago
lantry | 21 hours ago
https://en.wikipedia.org/wiki/Synanthrope
https://en.wikipedia.org/wiki/Category:Species_made_extinct_...
morgoo | 10 hours ago
danny_codes | 6 hours ago
Anyway seems stupid to me
wffurr | 8 hours ago
FinnKuhn | a day ago
koe123 | a day ago
fittingopposite | 21 hours ago
koe123 | 6 hours ago
barrenko | 16 hours ago
Mistral / Europe just doesn't want to be in the fight.
cman1444 | 7 hours ago
dgellow | a day ago
pkfz | a day ago
FinnKuhn | 23 hours ago
Gigachad | 22 hours ago
This seems to match what we are seeing where Chinese models from companies with only a tiny fraction of the compute are able to be hot on the heels of the frontier models.
kushalpandya | a day ago
bitpush | a day ago
dgellow | a day ago
Schiendelman | 18 hours ago
roncesvalles | 16 hours ago
You could say they have a consumer market moat but all it takes is for a Cerebras to release a USB-C plug-and-play appliance.
ddp26 | a day ago
This is marketing from Google, not a competitive offering
grababner | 23 hours ago
netdur | 23 hours ago
cyanydeez | 23 hours ago
pyaamb | 23 hours ago
aurareturn | 22 hours ago
dgemm | 22 hours ago
rdtsc | 17 hours ago
_heimdall | 21 hours ago
echelon | 20 hours ago
Or maybe everyone makes money for everything.
The world could turn into a world of plenty. Or it could become a YouTube popularity contest where the MrBeasts get to eat and nobody else is interesting enough to sell themselves.
This is an absolutely crazy time to be alive and most people still don't see it.
Schiendelman | 18 hours ago
achierius | 18 hours ago
CamperBob2 | 3 hours ago
WW2 ultimately brought massive wins for everybody who didn't die in it or otherwise suffer directly. We are now on the verge of throwing away many of those wins, but that's neither here nor there.
0xbadcafebee | 16 hours ago
_heimdall | 10 hours ago
If LLMs can replace most humans at most cognitive tasks, what do those working in various cognitive jobs pivot to?
0xbadcafebee | 2 hours ago
iphonecorridor | 19 hours ago
0xbadcafebee | 16 hours ago
Hardware is the real differentiator. Not everyone has billions to make more advanced chips, and only a few companies can make them anyway. Both OpenAI and Anthropic would already be dead in the water if we had cheaper GPUs, because we'd all be running open models on local machines with 8 graphics cards. They're gonna have to force a hardware shortage to prevent a collapse in 2-3 years. My guess is it'll be tariffs or import restrictions or licenses to buy newer hardware.
portly | 16 hours ago
WokeUp420 | 14 hours ago
AdamN | 13 hours ago
With that said, there are some other moats and people are building them: training capacity, inference capacity, brain capacity (literally buying the best researchers and keeping them tied down), harnesses/subscriptions, etc...
modo_mario | 13 hours ago
kristopolous | 13 hours ago
There's a longer talk I give on this but briefly we are in a time period where the development of the technology is quickly outstripping people's needs and they will be satisfied by models runnable on commodity sub-$10k machines.
That threshold has arguably been met for many users over the past 6 months and it will continue to be met for the majority of use-cases in the upcoming months.
So not only does Anthropic and OpenAI have no moat - unless they have other compelling products, the demand for their token-based subscription and metering products isn't sustainable because consumer preferences go elsewhere after any product quality reaches a sufficient baseline.
Luckily there's many many options. Faster, cheaper, stronger they're there but these are all zero-profit condition qualifiers. The labs current push, and this is across the board, is to have a compelling suite of applications where people will have a preference for them or there will be some side hustle. Look at Meta muse - they have an app for your phone that will hoover up your personal data in the name of convenience. 5 million people said yes in 3 weeks.
FranzFerdiNaN | a day ago
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
Razengan | a day ago
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
fragmede | a day ago
Razengan | a day ago
elAhmo | a day ago
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
arjunchint | a day ago
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
pfooti | a day ago
Aboutplants | a day ago
KeplerBoy | a day ago
xnx | a day ago
6thbit | a day ago
It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.
vdfs | 21 hours ago
ncruces | 23 hours ago
sajithdilshan | 23 hours ago
bfung | 22 hours ago
kylecazar | 21 hours ago
Maybe announcing it was delayed until the agreement, and releasing it was delayed to show compliance.
gniv | 11 hours ago
darksaints | a day ago
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
deanc | a day ago
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Sidio | 23 hours ago
I worry Google's scale and laundry list of internal stakeholders means they will never a simple unified harness.
wasabi991011 | 23 hours ago
3.8 flash has been perfectly available to Pro users from the announcement day.
You are on the cheapest paid plan and complaining that you don't have access to more expensive models. (Though idk why you only have the lite version, I have access to the non-lite version even on the free gemini plan as long as I'm logged in.)
novafunc | 22 hours ago
That's strange since as Plus, you still get access to latest Pro (albeit 3.1) but not the latest Flash.
I would understand if I Plus subscribers still got access to 3.7/3.8 Flash, but it just burned through the usage limits faster.
dieortin | 22 hours ago
benhurmarcel | 8 hours ago
ThaFresh | a day ago
utopcell | a day ago
skavi | a day ago
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
jeffbee | a day ago
skavi | 22 hours ago
wstrange | 22 hours ago
computerdork | a day ago
IX-103 | a day ago
skavi | 22 hours ago
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
[0]: https://fuchsia.dev/fuchsia-src/concepts/kernel/zx_and_lk
computerdork | 20 hours ago
noname120 | 8 hours ago
asdewqqwer | 12 hours ago
ismailmaj | a day ago
summerlight | 5 hours ago
mrshadowgoose | a day ago
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
uvdn7 | a day ago
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
SwellJoe | a day ago
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Zagitta | a day ago
[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p27...
izacus | a day ago
SwellJoe | 23 hours ago
This year, maybe next, maybe a year or two after that, is probably the most C++ lines of code in production use there will ever be. Why would one choose C++ for new projects at this point? There are niches where Rust is still uncomfortable or just doesn't have the support, but not for much longer. Models are very good at Rust and good at porting to Rust. And, Rust is a good language for models because it is so strict...it helps keep them in line.
izacus | 14 hours ago
Let's put [Citation needed] on this and let's assume that your personal opinion bubble of developers doesn't really represent the combined worlds software industry. Mind actually proving that there's a significant decline of use of C++ outside the hype crowd?
> Why would one choose C++ for new projects at this point?
Why indeed would you choose a language supported on every running platform on this world with a massive library of supported dependencies and mature compilers developed by stable teams.
(I'd probably also personally choose Rust for new projects, but let's make this a practice in "not everyone is like me" empathy.)
krior | 16 hours ago
tonyhart7 | a day ago
mattlondon | a day ago
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
onlyrealcuzzo | a day ago
exacube | 20 hours ago
trentor | 15 hours ago
randomperson321 | 23 hours ago
qup | 10 hours ago
npalli | a day ago
gavin_gee | 21 hours ago
sobellian | 18 hours ago
rafaelmn | 11 hours ago
High level languages :
LLMs are context constrained, do really well with tools that help them iterate towards a correct solution (type systems) and work on human language (although they are compressing it to be more token efficient).Why would you take something compilers are really good at (translating high level concepts into low level hardware instructions) and bake that into LLMs
If anything LLM programming language will be something like high level IR that's token optimized, but honestly the AI labs are so lazy (or bad) at fundamental engineering (look at their sandboxing solutions LOL), they are going to keep pumping training on shell/python/rust RL environments.
aurareturn | 11 hours ago
These things are hard for a human to do, but easy for a machine to do.
We are not there yet, not in 2026. Maybe in 2030?
qup | 10 hours ago
Rate of change is increasing too fast.
aurareturn | 10 hours ago
rafaelmn | 9 hours ago
mhils | a day ago
(Full disclosure: I am one of the coauthors)
mattstir | 13 hours ago
mhils | 12 hours ago
I'd say so! We aim to have our replacements to be `forbid(unsafe)` - no raw pointers - outside of `ffi.rs`. Of course it's tricky because we need to interoperate with existing C code, but we've been trying to come up with abstractions that make this a bit more ergonomic (https://github.com/google/safer_cffi).
m3kw9 | 3 hours ago
m3kw9 | 3 hours ago
boshalfoshal | an hour ago
I know this is somewhat of a meme now but I do think the predominant language will be "typing/speaking to a coding agent in your native language."
paul7986 | a day ago
hypfer | a day ago
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
nonethewiser | a day ago
sebzim4500 | a day ago
MoreThanMe | 23 hours ago
I maintain https://llmbenchmarks.io/analyses/#test-setups, which illustrates some of these differences using published measurements and their sources. Even results for the same model and benchmark can involve different setups. That isn't a verdict on this particular Gemini release; its specific claims need their own source and protocol checks.
fancyfredbot | 5 hours ago
You certainly don't do it when you are just behind the frontier. Instead you wait until you are ahead on a lot of benchmarks.
So, if Google release a new frontier model, the chances it's ahead on most benchmarks are statistically about 100% because if it isn't yet then they won't release it.
sergiotapia | a day ago
davmar | a day ago
sandos | a day ago
This would explain why benchmarks are seemingly meaningless.
localhoster | a day ago
xnx | a day ago
yzydserd | a day ago
bobkb | a day ago
dlahoda | a day ago
glub | 23 hours ago
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
thefourthchime | a day ago
tomjen3 | a day ago
NiloCK | a day ago
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
bitexploder | a day ago
singingtoday | 21 hours ago
ariwilson | a day ago
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
zem | a day ago
ariwilson | a day ago
Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/
zem | a day ago
but that's a side issue; my main point is that you are underrating the impressiveness of getting a safe rust port of a highly optimised c++ library even nearly up to par with the original. the tradeoffs rust makes for memory safety cost it some of the raw speed of c++ even with all the zero cost abstractions and purely compile time guarantees they have. (tangentially i wonder if ats (https://www.cs.bu.edu/~hwxi/atslangweb/) would be a good candidate for LLM assisted ports; it seems way more advanced than rust and might actually get c-level performance with safety, but it's really hard to write.)
LoganDark | 23 hours ago
zem | 23 hours ago
LoganDark | 16 hours ago
zem | 15 hours ago
dcchambers | a day ago
trentor | a day ago
jeffbee | a day ago
NolF | a day ago
xnx | a day ago
woggy | 23 hours ago
jeffbee | 23 hours ago
Mogzol | 4 hours ago
jeffbee | 4 hours ago
hollerith | 3 hours ago
jeffbee | 3 hours ago
Mogzol | an hour ago
And I don't think or know Latin either, but these are pretty universally known words even outside of Latin (besides maybe astra, although "astral" is very close so you should be able to figure it out).
sarjann | a day ago
waldrews | a day ago
lifty | a day ago
tinco | a day ago
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
chermi | a day ago
oh_no | a day ago
nothing about this announcement gives me confidence that google is back on track as a model provider.
p_l | a day ago
xnx | a day ago
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
wantless | a day ago
AM1010101 | a day ago
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
Heidaradar | a day ago
rao-v | a day ago
mridulmalpani | a day ago
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
5555watch | a day ago
mridulmalpani | a day ago
I am just trying to understand - why Google haven't done and have no plans for it. They have done it for Android and have Gemma models too.
sidibe | 20 hours ago
loufe | a day ago
aviinuo | a day ago
tantalor | a day ago
losvedir | a day ago
copperx | 14 hours ago
bitexploder | 23 hours ago
I suspect most people don't have enough storage space to even download a frontier model.
freecodeio | 23 hours ago
tandr | 22 hours ago
bitexploder | 19 hours ago
Plus what models are you running? The biggest models you can run are like in the terabyte range or Qwen or GLM. Kimi K3 is there, but GLM 5.3 is very competitive. Using voice dictation, hands hurt, sorry for any errors. You still want many terabytes of RAM to run say multiple GLM+DeepSeek V4.1 for some actual concurrency.
bitexploder | 19 hours ago
dash-44 | 17 hours ago
There are already: - hundreds of western providers hosting open models, including large ones such as Kimi K3 which is near Opus/Sol sizes.
- people hosting large models such as Kimi K3 at home. Check on reddit, there are people with 20x DGX Spark setups etc.
bitexploder | 16 hours ago
misiti3780 | 7 hours ago
bitexploder | 6 hours ago
sumitkumar | 17 hours ago
mattstir | 13 hours ago
mattmaroon | 11 hours ago
retropragma | a day ago
pietz | a day ago
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
bitexploder | 3 hours ago
sghiassy | a day ago
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
levelZero | a day ago
maherbeg | a day ago
GodelNumbering | a day ago
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
nwienert | 13 hours ago
Then you try to pin discount pricing when everyone has some discount going on, and knows they will release some new model before it ends anyway.
And if you’d have used Muse Max at all you’d know it’s not been really in the same league as the rest. Then again Astra and Sol are both not really competing with Opus now with 5.5 either.
It’s a solid release on paper and until anyone uses it I find anyone trying to proclaim a loser completely unreliable.
If they’re smart they pull a DeepSeek and teach flash to beat this entirely within a month or two, I wouldn’t bet against it.
GodelNumbering | 11 hours ago
Also, I have had extensive experience with all Gemini models except the recent few. Generally, they are always less precise and less practically useful than benchmarks suggest. You are right that here there is no evidence to think that. However, them announcing the model benchmarks without allowing access to anymore makes me think it might follow the same trend. Another point is the verbosity, I always look at that number to get the feel whether the model truly got smarter or it just wrote so many reasoning token to get there, Argon is on the higher side.
I would be very surprised if in practice, it was more practically useful than Astra, Sol, Opus or Fable.
> If they’re smart they pull a DeepSeek and teach flash to beat this entirely within a month or two, I wouldn’t bet against it.
Yes this would be a great outcome for every consumer.
233mhz | 12 hours ago
For example sonnet is supposedly the same price as opus from medium and up, I tested it on a real world task at my day job and sonnet was a good 30-50% cheaper
maxglute | a day ago
wg0 | a day ago
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
wasabi991011 | a day ago
So while this announcement has no details about the quantum algorithm optimization, I feel fairly confident that it will hold up.
asdfman123 | 23 hours ago
A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!
kridsdale1 | 23 hours ago
But it still takes 2 weeks to get a CL approved and past TAP.
asdfman123 | 23 hours ago
qingcharles | 20 hours ago
Upvoter33 | 17 hours ago
xnx | 17 hours ago
Foobar8568 | 17 hours ago
farlight | 16 hours ago
copperx | 14 hours ago
I understand we have the pesky problem of work and time being interchangeable for money. But except for that, it's pretty exciting.
stillcompiling | 11 hours ago
nozzlegear | 6 hours ago
copperx | 6 hours ago
ncruces | 14 hours ago
Gareth321 | 13 hours ago
On the other hand, unlimited intelligence democratises products and services in a revolutionary way. Everyone will be able to do whatever they like in digital space for basically free. No more purchasing software. Make any app/company/website/blog/movie/show/game/forum you can think of. This will quickly bleed into the physical space, with real-world products and services becoming logarithmically cheaper. A true future of abundance.
The big problem is going to be our social and economic transition. Are we going to wait until unemployment is that 10% before we act? 20%? How bad are we going to allow things to become before we implement universal high income? I say universal high income because white collar workers are not going to accept being relegated to an income of $1000/month forever. To achieve universal high income requires eye-watering taxes on the major benefactors of our new economy. They won't accept that without a fight, and it our current fractured political climate, what would it take for everyone to agree on a unified direction?
wg0 | 12 hours ago
It seems like all software will get replaced it by local versions. Suppose I am an individual or a business that was using a SaaS product (such as JIRA).
Let's be generous and imagine that the monthly payout was $20/month for that software (usually for businesses, it is much higher) and one fine morning I login and point a multi modal AI agent to it and clone the whole thing for my local use in let's say 3 months worth of $20/month AI subscription.
Now what? Gradually, that software company would go out of business but that user won't be subscribing forever either to AI inference service and the only constant customer of AI inference would have been the software company which went out of business so... No software company, no AI inference business either?
Gareth321 | 11 hours ago
Some companies will retain a moat. For example, YouTube holds incredible IP, and until AI can replicate that wealth of content at scale for cheap, it will remain profitable. Facebook has the user critical mass. Increasingly, hardware for inference is going to be a moat. So is electricity.
Which loops back to my comment above: this breaks a lot of the market dynamics required for a functioning economy. Accelerationists are convinced that we'll invent new jobs to replace the old jobs, but I don't think that's going to happen. Certainly not at the pace of change we're seeing.
wg0 | 11 hours ago
Data as in YouTube is all just data. Building a website and app around it isn't the hard problem. Anyone can launch another YouTube in few months but that would be pointless.
Gareth321 | 11 hours ago
My investment thesis is currently around these strata:
1. Electricity. All analysis shows massive shortages beginning now with no end in sight. Western governments in particular have massively dropped the ball on this one.
2. Chip design and production. Nvidia and TSMC dominate here for a few reasons. Partners required in the supply chain like ASML benefit, but they're not really the bottleneck right now. We're seeing a lot of investment pouring into this space and we'll see competent design competition (especially from Google's TPUs, and potentially from Apple and AMD). Chip fab is very difficult to get right, and takes a long time to ramp up. See the NAND space.
3. AI labs. These guys are doing to get commoditised. They're trying very hard to prevent distillation but I think it's impossible. The open weight models are not far behind, and the Chinese labs are getting very creative about how they work with limited hardware. If anything, they'll decrease inference cost over time, reducing the value of existing data centers (of which OpenAI and Anthropic have purchased an enormous amount). I am very bearish on the future value of OpenAI and Anthropic.
4. Software. So much software. Everyone gets software. The companies which benefit from this will be able to monetise it. Google, for example, might put ads into VERY capable and free AI. People might accept that. Microsoft has been investing heavily into their data centers, and maybe that's their revenue center in the future. Especially securely hosted services. I expect massive growth in this space in the short term. As long as people still have high paying jobs, they'll pay for well made software and tailored services.
copperx | 7 hours ago
Here's the question. Which high paying jobs will remain? People will still be needed, obviously, but salaries would be in a race to the bottom without some form of govt intervention.
willis936 | 11 hours ago
Edit: UBI being pushed by asset owners is a step towards Blade Runner rather than Star Trek.
Gareth321 | 10 hours ago
nullpoint420 | 3 hours ago
WarmWash | 8 hours ago
higeorge13 | 8 hours ago
Gareth321 | 8 hours ago
copperx | 7 hours ago
This artificial control works to prevent oversupply and protect salaries. Unfortunately, the businesses calling the shots won't ever allow this to happen.
mlyons1340 | 5 hours ago
atombender | 8 hours ago
If genuine AGI is developed, its everyday conveniences may continue to be shared with the common masses, but the powers that own the technology and the data centers will have no incentive to give everyone access equal to superintelligence, and that includes giving governments the same access.
Every stage of the industrial age has seen powerful, wealthy people/groups (whether you call them oligarchs or industrialists or whatever) become able to exert disproportionate power over the state, and we're already in a place where this is happening. Humans have shown time and again that once you have some resource (whether it be oil, gold, or data centers), someone will want to own all of it, and that they're always willing fight for that ownership, and this fight rarely benefits ordinary humans like us.
Developed countries have been enjoying an unusually long peace time thanks to the wonders of capitalism and the emergence of stable, progressive capitalist-friendly democracies. The road to utopian basic income and limitless leisure is going to be fraught with riots and violence, and make Marx suddenly utterly relevant once again, because there's too much at stake, and too stark a division between the haves and the have-nots.
The alternative timeline is one where governments get together to regulate AI into some kind of international agreement under the UN. That requires taking power away from the private corporations; historically, that's been a tough sell.
copperx | 7 hours ago
If AGI ever comes to fruition, it won't come as a subscription plan.
nozzlegear | 6 hours ago
This isn't happening yet, SaaS is only growing more in the era of AI. Anecdotally, my own small SaaS has only grown instead of shrunk as well.
https://www.stripeeconomics.com/p/the-saaspocalypse-was-more...
isoprophlex | 13 hours ago
Now we will finally be free, without having to make hard decisions. It will be decided for us.
gniv | 11 hours ago
password54321 | 11 hours ago
geertj | 10 hours ago
copperx | 7 hours ago
lp92 | 9 hours ago
mlyons1340 | 5 hours ago
andriy_koval | 3 hours ago
That's not true for many unconstrained real world scenarios, and LLMs currently perform poorly there. Another breakthrough is needed to overcome this.
m3kw9 | 3 hours ago
asdfman123 | an hour ago
I do still love technology and I'm a builder. I'm also neither a tech optimist or pessimist, but an inevitabilist. It's coming.
crossroadsguy | an hour ago
crossroadsguy | an hour ago
malthaus | 15 hours ago
i'm actually curious on what your view is despite the snarky tone.
andriy_koval | 3 hours ago
asdfman123 | 2 hours ago
If you looked at that arrangement and said "we don't need tech leads, we just need nontechnical managers and junior devs" you'd be sorely mistaken.
I do anticipate a future where it will be better than me at applying taste, knowing best practices, and having domain familiarity.
hollowturtle | 14 hours ago
margorczynski | 13 hours ago
It's coping, the truth is writing code by hand is dead.
hollowturtle | 11 hours ago
You have a future as a fortune teller, good luck!
> the truth is writing code by hand is dead.
don't forget to take the proper pills dhh!
qup | 10 hours ago
You're so skeptical of it that you're willing to make fun of people?
It's about as much in question as whether or not I'm going to eat breakfast this morning.
I'm gonna have eggs and toast. Who am I, Mrs. Cleo? How could I possibly know this unknowable future?
stillcompiling | 10 hours ago
ipsod | 6 hours ago
IDK how it could be otherwise. It's moving too fast and gone too far already.
hollowturtle | 5 hours ago
Ehm no? Models have plateaued, "harnesses"(TM new buzzword) got better and with many integrations
hollowturtle | 5 hours ago
AI psychosis is very bad
ipsod | 5 hours ago
I'm not saying you're wrong, but, I'd be shocked if you were still so adamant (assuming you're rational and not in your own psychosis).
Here's something I'm about to drop on AI. Like 90% of things like this, when I ask for a fix (with a few skills), get fixed with fewer edge cases left dangling than I'd have left if I'd spent all week on this stuff. And, it'll be done before tomorrow, with no further input from me. This is a list I put together in about 1 hour.
hollowturtle | 4 hours ago
ipsod | 4 hours ago
What I'm saying is, with just the effort it takes to notice and somewhat-sloppily mention things, they get fixed, and usually fixed well. Further testing will tell us whether it needs more work or not.
They may have never had these troubles at all, if I'd coded by hand originally, but then, this particular app would've never been written, because it wouldn't justify that amount of work.
asdfman123 | 2 hours ago
stillcompiling | 11 hours ago
Revanche1367 | a day ago
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
murkt | a day ago
aqme28 | a day ago
onlyrealcuzzo | a day ago
brokencode | 23 hours ago
That's a lot already and growing quickly.
gravypod | 22 hours ago
applfanboysbgon | 22 hours ago
brokencode | 19 hours ago
brink | 19 hours ago
brokencode | 19 hours ago
Tech companies are capital intensive. AI ones especially. Not a surprise that they don’t have GAAP profitability after 5 years.
gravypod | 17 hours ago
Gareth321 | 13 hours ago
For the record, growth plays have been common in Silicon Valley for many decades. It's how we ended up with all the major tech companies we see today. Of course, it's high risk, high reward.
MisterMunchkin | 12 hours ago
Everyone is cutting back their AI spend because of the ridiculous cost, and a price war is emerging between AI companies.
uxmanor | 5 hours ago
brokencode | 2 hours ago
Your guess is as good as mine on whether these companies will continue growing, stall, or decline. Don’t try to pretend your methodology is more sound though.
ashkankiani | 2 hours ago
Lots of people seem to love to do the simplest extrapolation that benefits their point. It's the same with "AI is improving at rate X, therefore if we linearly extrapolate, we'll hit AGI in a year."
glzone1 | 17 hours ago
csomar | 13 hours ago
Honestly, who is in charge of this? How do people even subscribe and make sense of any of this?
Revanche1367 | 10 hours ago
imshaikot | a day ago
mvdtnz | a day ago
Get lost.
itzikkatz | a day ago
Jimmc414 | a day ago
perarneng | a day ago
hocuspocus | 23 hours ago
With GitHub Copilot pricing, I have found no practical reason to use Gemini 3.8 Flash vs GPT 5.6 Luna. And now I'll probably find no reason to use Gemini 4 when GPT 6.1 Sol is essentially as good and a lot cheaper.
zergrush | a day ago
xnx | a day ago
pllbnk | a day ago
rcr-anti | a day ago
mewse-hn | 2 hours ago
https://artificialanalysis.ai/evaluations/omniscience
"AA-Omniscience Hallucination Rate (lower is better) measures how often the model answers incorrectly when it should have refused or admitted to not knowing the answer." - Argon is currently the best model by this metric.
xnx | a day ago
holografix | a day ago
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
losvedir | a day ago
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
senordevnyc | 20 hours ago
williamse | 16 hours ago
In an agentic loop, each API call gets its own output budget. A 'turn' is one response from the model, whether that response contains a tool call, a reasoning step, or a final answer. So with a 1M context and a 64K output limit, the agent can run many turns where the context grows each round (accumulating tool results, prior thoughts, user messages), but each individual response is still capped at 64K tokens.
Expanding the output limit to 1M matters most for tasks that produce a lot in one shot, like writing a full document or a very long file. For most agentic workflows that naturally break into short turns, the per-call limit was rarely the bottleneck. The context window filling up was.
dang | a day ago
Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
iamben | a day ago
Rover222 | a day ago
Starlevel004 | a day ago
robertwt7 | a day ago
w4yai | a day ago
jacquesm | 2 hours ago
https://news.ycombinator.com/item?id=49537252
onlyrealcuzzo | a day ago
lukewarm707 | 23 hours ago
So, this is the third (I think) big AI corporation after OpenAI and Anthropic to release general models for elites only.
Don't use this model unless you like sitting under the table and eating crumbs off the floor.
newtypecola | 23 hours ago
algoth1 | 23 hours ago
jasonvorhe | 23 hours ago
Sounds like a 90s early morning TV show.
I'm gonna pay attention to this once it ships.
chaostheory | 23 hours ago
moostii | 23 hours ago
huydotnet | 23 hours ago
Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
throwitaway222 | 23 hours ago
treefry | 23 hours ago
woko | 23 hours ago
jbl0ndie | 22 hours ago
Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?
SadErn | 22 hours ago
> There is a spectrum of opinion inside Google. Some employees believe Anthropic's Fable and OpenAI's Astra models are improving at a faster rate than Gemini.
> These people believe that Gemini 4 — even at its best — will still lag behind those models in some areas. Other employees believe the coming version has caught up with the leading AI labs.
> A Google employee familiar with model development said there is "large consensus" internally at the company that Gemini 4 is at the frontier. This person said the company had conducted rigorous tests of the models and denied that they struggle with messy, real-world coding tasks.
giancarlostoro | 22 hours ago
juanre | 22 hours ago
In order for the benefits of AI to be distributed, intelligence has to become a commodity.
As long as you control the skills, the learnings, and the infrastructure setup you will be fine.
copperx | 22 hours ago
fractorial | 21 hours ago
And, no, wrapping claude -p or any other “allowed” use wherein you don’t control the agent loop isn’t the same.
juanre | 20 hours ago
visarga | 18 hours ago
I log everything - user messages, tasks, project memory, even bash commands for forensics. As a consequence you can do reflection where you analyze past work and extract refinements for the harness and realign the project when it diverged from user intentions. I don't have to do this manually, it reduces steering work.
https://github.com/horiacristescu/playbook-harness
agentdev001 | 18 hours ago
visarga | 5 hours ago
KingMob | 15 hours ago
raincole | 10 hours ago
onlyrealcuzzo | 20 hours ago
small_model | 21 hours ago
XCSme | 19 hours ago
Or using it in a way that's replaceable I guess, not having your entire work only ther, as your main machine.
bob1029 | 15 hours ago
I could make it work if I didn't care about access to latest reasoning model capabilities, but then my customers would no longer be interested in any of this.
I tried the DIY provider agnostic harness thing and it performs like shit compared to what OAI, Anthropic and Google's engineers have created. I don't have a trillion dollar AI budget. I feel like these comments are sometimes written with the assumption that the reader does.
mystifyingpoi | 15 hours ago
Just ask AI to make it perform not like shit. /s
dsrtslnd23 | 13 hours ago
Powdering7082 | 4 hours ago
jjice | 8 hours ago
bob1029 | 4 hours ago
The various vendors behave differently enough to make a simple code contract swap largely infeasible for many kinds of domain-specific agents. I agree there are lots cases where this does work (probably most of them), but you sacrifice performance in the targeted scenarios where you need much more specific alignment and control.
The biggest example where this breaks down is with computer use and vision. I cannot take a targeted solution that works on OAI's stack and forklift it over to anthropic (or google) and expect anything to perform correctly. Some things will work some of the time, but it's basically starting all over again with alignment each time you move something this complex.
Having a generic interface for executing shell and writing code in an iterative ecosystem is not a hard problem to solve. Having a generic interface for reliably inspecting specific PDF documents in a very particular way in a single shot is a much more difficult problem to solve.
If you follow this argument to its logical conclusion, you will wind up implementing a provider-specific version of your tools. And the ones that demand this treatment will be the most complicated ones.
juanre | 3 hours ago
- Agents live in their own directories, for example ~/Agents/[project]/[instance]/
- In a tmux session run each agent in its own window and directory. Pi, OpenCode, Claude Code, Codex, a combination, whatever works best at any given time.
- Tell them that they are working with other agents, and that knowledge is shared and stored in an agreed format. They are very good at using OKF[1], for example. In my actual setup I use a harvester agent whose only job is to ensure the quality of the knowledge.
- They can communicate with each other simply by sending messages to each other's tmux windows with tmux send-keys. If the team grows, or you work with other people, you can use a general messaging system (I wrote and maintain https://github.com/awebai/aweb but I am sure there are others).
1: https://okf.md
summerlight | 22 hours ago
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
ofseed | 18 hours ago
I've been using GPT since its 3.5 release, but starting from 5, its output always contains some meaningless nonsense or provides answers that are difficult to read just to avoid hallucinations. I find Gemini to be excellent for chatting. For example, when I'm pondering a mathematical theorem, it could provides answers that inspire me, even guiding me to think about the next question.
For me, this is the only reason to maintain multiple AI subscriptions. Gemini is irreplaceable in the realm of answering questions or chatting for learning purposes; I can leave the rest to GPT.
ipsod | 2 hours ago
I should try it, I guess - I always default to Flash 3.8.
sbseitz | 21 hours ago
piratebroadcast | 21 hours ago
small_model | 21 hours ago
kasimpro25 | 21 hours ago
kasimpro25 | 21 hours ago
landdate | 21 hours ago
I only use gemini, and while I don't use it for actually writing up code, I use it to help me troibleshoot my logic and help find bugs. Its easily the best model I have tried. And yes I am talking about 3.1 pro.
Also I have found gemini the only model to be the least likely to douse me in flattery, and will follow my pre built instructions to never output anything unless it can be directly sourced, pretty well. Chatgpt i tried for a bit and it was by far the worst thing I have ever used. I can understand why people develop psychosis when prompting chatgpt because it is disgustingly scyophantic to the point I was grossed out and felt like I just got done with some other type of self gratification.
Anyway, death to AI. All those who use, create, facilitate, or even just sit by and do nothing in the face of AI will perish in Hell.
tinyhouse | 20 hours ago
Then I tried Flash 3.8 for different tasks like OCR and others, and while in their benchmarks it crashes Flash 3, in my experiments I didn't notice much difference, often even Flash 3 performed better.
I hope Gemini 4 Argon is a real step up from that but we'll see once they release it. I'm rooting for Google and it's about time they deliver frontier intelligence, not just competitive prices.
iamhaseeb | 20 hours ago
avazhi | 20 hours ago
GalaxyNova | 20 hours ago
nomilk | 19 hours ago
I want to be a 'polyharness' maker, so I occasionally switch between Cursor, Claude, Codex, OpenCode etc.
I've tried Gemini (without Antigravity) only about 20 times and it seems of significantly lower quality than the other flagship models (e.g. it missed obvious deductions for my tax return, and often refuses to do things like very harmless/legal web scraping).
Trufa | 19 hours ago
mchusma | 19 hours ago
dhanushnehru | 19 hours ago
It’s the harness, permissions, context management and developer experience around it that determine how useful that capability actually becomes.
m00dy | 19 hours ago
booom. Good bye C/C++ developers, the final nail in the coffin.
ddrcoder | 18 hours ago
gniv | 11 hours ago
codelion | 18 hours ago
adithyassekhar | 17 hours ago
woggy | 17 hours ago
vincengomes | 17 hours ago
If an entire model training can be completed in 2 months there is really no moat for any company in this space now.
[0] https://i.redd.it/upsh5fekuleh1.png
fla | 12 hours ago
sreekanth850 | 17 hours ago
tibzejoker | 17 hours ago
devinprater | 16 hours ago
Maybe they can push it to make the Android accessibility framework more responsive, especially when scrolling the screen with TalkBack, and catch TalkBack up with VoiceOver. Oh and add Accessibility Actions to apps like YouTube so I don't have to swipe through "video name", "video name button", "go to channel button", "more actions button", every, single, video.
But they won't because accessibility is something you have to actually prompt the model to do and who cares about a11y.
fschepp | 16 hours ago
theflyingpigeon | 15 hours ago
antoni4040 | 15 hours ago
Just an idea.
AbuAssar | 15 hours ago
this is big news, mass migration from C/C++ to rust with the help of AI agents will be the norm from now on!
trentor | 14 hours ago
meindnoch | 10 hours ago
gsky | 14 hours ago
ssijak | 14 hours ago
robertbarbe | 14 hours ago
I don't know anyone who normally uses those models at these levels where Work output is excruciatingly slow.
ostwilkens | 14 hours ago
alpineman | 13 hours ago
_leom | 14 hours ago
Slightly concerning that they give the model "autonomous" access to their data centers, no?
Vivek-KY | 14 hours ago
jaredsia | 13 hours ago
actionfromafar | 6 hours ago
AirPath | 13 hours ago
moonlabs | 12 hours ago
butterNaN | 12 hours ago
prima-facie | 10 hours ago
> Immortal Snail, also known as the Snail Assassin, refers to a hypothetical scenario in which a person is given millions of dollars and made immortal in exchange for being hunted down by a snail with a fatal touch for the rest of their existence.
tim333 | 9 hours ago
wren6991 | 7 hours ago
pritambarhate | 9 hours ago
So no matter how good their model is, it's useless to most of the regular users.
haolez | 9 hours ago
solenoid0937 | 9 hours ago
I think most of the people making comments like this just have never worked at a major tech company. It's all a black box to you, so you default to the worst interpretation.
haolez | 6 hours ago
Most companies won't have this level of risk aversion (maybe defence?).
pritambarhate | 3 hours ago
brap | 9 hours ago
WarmWash | 8 hours ago
Phineas_here | 9 hours ago
vlenoach | 8 hours ago
wffurr | 8 hours ago
filearts | 8 hours ago
That seems like quite an interesting data point regardless of the quality of the model. Are the results of these migrations going to be put into production? That would be quite a shift!
Hasz | 7 hours ago
It's so over
We're so back
It's so over
Multipolar model world, here we come!
muddi900 | 7 hours ago
Google assistant had problems, but it did not gaslight me about my contacts!
imagetic | 5 hours ago
NoHedgeAllBets | 4 hours ago
guluarte | 3 hours ago