The Fight Over OpenAI’s Math Breakthrough Is a New Kind of Scientific Arms Race

224 points by malcolm58 11 hours ago on reddit | 42 comments

Derpykins666 | 3 hours ago

It's all just stolen data and research.

This is why I was so concerned to hear that literally doctors in college with their final papers were just shoving their like 40 page research papers into ChatGPT to spellcheck them. It's like.... you worked probably years on that research, and you just gave it to them... so they could spellcheck your shit... people are out of their minds. They'll take everything. I 100% believe a company like this would just steal your data and research and do it themselves and completely undercut you. Why wouldn't they? There's no integrity left in this type of company. It's full speed ahead, and they're rolling the dice every day that there will be some catastrophe. But hopefully they get to make some money along the way, I guess.

euuzaik | 4 hours ago

"math breakthrough"

*looks inside*

it's just stolen research

CognitionMass | 10 hours ago

I think its just the death of science more than anything. Corporate controlled LLMs are just not compatible with science in any long term way for two main reasons.

First, they are forces of standardisation. A scientific enterprise entirely built around Claude for example is one that's incapable of divergent and novel thinking. This is for example why AI is so good at coding, because coding is a formalised language. Scientific advancement is about as far from formalised and standardised as you can get. A system that outputs the same near identical websites for entirely different prompts by entirely different people half way around the world is not a system capable of scientific advancement in any meaningful sense.

Second, the immense resources required to train a single AI will tend to push them to be forces of stagnation. New fields and new approaches can really only be incorporated into an LLM by retraining them with that data set. But because retraining is such an expensive and involved process, with LLMs there will be additional costs and efforts required to adopt any novel fields into a scientific institution built on an LLM.

One of the biggest hoaxes of modern AI imo is to call them 'generative" systems, when they are one of the least generative systems we know of. A simple multiplication function is more generative than an LLM because its implementation is maybe a few bytes, and its capable of generating an infinite set with that few byte implementation. These LLMs are about terabyte implementations. To put that into context, I can store humanities entire public domain written outputs -- what makes up a vast amount of the training sets of these LLMs -- on less than 1 terrabyte. Literally I've done this. I have much of stack exchange, all of Wikipedia including images, all of the public domain English books, and a bunch of other stuff, less than 1 terabyte.

These frontier LLMs are not much smaller than their training sets;  they are more generative than a lookup table; the least generative function possible. They generate very little. Most of what they do is information recall. They are very impressive advancements  in data compression and information retrieval.

grr5000 | 10 hours ago

Personally I think even bigger than that is anything they discover will be completely “owned” by the company, where as many research projects are publicly funded and able to be bid on and worked on with companies… rather than outright owning the discoveries

CognitionMass | 10 hours ago

This isn't that new though. The big publishing cartels control much of scientific information already.

The founder of reddit committed suicide after trying to go against that and the law coming down on him.

It could certainly become worse with LLMs though. But its not really unique to LLMs.

HelloWorld_bas | 6 hours ago

Poor Aaron was born too soon. If he’d lived until now and did an AI startup the law would bend over backwards to help him violate copyright.

Wardenbb | 8 hours ago

What can they own though? All of what they know is already known. General intelligence learning new abilities?

grr5000 | 8 hours ago

Ok example. They use their models to discover a cure for cancer… once they have that they then own the rights to it and can monetize that or keep it for themselves.

Educational-Year4005 | 6 hours ago

Ok? As opposed to Pfizer discovering it and sharing it with the world? Either they patent it and it's open after ~20 years or they keep it secret and someone else reverse engineers it.

dietcheese | 7 hours ago

You think that because it makes generic websites it can’t contribute to scientific discovery?

You don’t need to retrain a model to gain new knowledge. We have context, data retrieval of new papers and datasets, etc…

CognitionMass | an hour ago

I'm not saying they can't contribute to science by any means. In khuns terms, there is normal science, and there is paradigm shift. There is absolutely a place for usage of LLMs as a tool in normal science. What I am discussing is LLMs, particularly corporate controlled LLMs, defining science.

Take this one instance as an example. There were the two mathematicians that used LLMs as a tool, and worked with it for quite some time, treating it as a tool, versus the corporation that just came in and claimed their LLM came up with the solution on its own (a falsehood btw). There's nothing wrong with the former, and serious problems with the latter. Keeping in mind mathematics is the kind of formalised and standardized field that LLMs excel in. This is not really a scientific advancement.

Context and retreival are very limited and not nearly as robust as training on the datasets instead.

> I think its just the death of science more than anything.

That prediction is difficult to reconcile with peer reviewed work in which LLM based systems generated novel scientific objects whose properties were then confirmed experimentally.

Swanson et al., Nature, 2025 https://doi.org/10.1038/s41586-025-09442-9

> Corporate controlled LLMs are just not compatible with science in any long term way for two main reasons.

Corporate control raises legitimate questions about access, reproducibility, and institutional power, but it cannot establish scientific incompatibility when LLM based systems have already produced new results on open mathematical problems.

Romera et al., Nature, 2024 https://doi.org/10.1038/s41586-023-06924-6

> First, they are forces of standardisation.

LLMs can produce convergent outputs under some conditions, but experiments measuring divergent creativity show that convergence is not an intrinsic limit on the range of ideas they can genreate.

Bellemare-Pepin et al., Scientific Reports, 2026 https://doi.org/10.1038/s41598-025-25157-3

> A scientific enterprise entirely built around Claude for example is one that's incapable of divergent and novel thinking.

That categorical claim is contradicted by controlled experiments in which Claude was among the models capable of producing measurably divergent responses across repeated tests of creative generation.

Bellemare-Pepin et al., Scientific Reports, 2026 ttps://doi.org/10.1038/s41598-025-25157-3

> This is for example why AI is so good at coding, because coding is a formalised language.

Formal syntax may help, but successful code generation also depends on semantic inference, candidate generation, search, filtering, and reasoning about problems expressed in ordinary language, so formalisation alone does not explain the capability.

Li et al., Science 2022 https://doi.org/10.1126/science.abq1158

> Scientific advancement is about as far from formalised and standardised as you can get.

Scientific inquiry includes open ended judgment, but large parts of science are rigorously formalized through mathematics, statistics, measurment, and physical constraints, which is why algorithms can recover governing equations directly from empirical data.

Udrescu and Tegmark, Science Advances,. 2020 https://doi.org/10.1126/sciadv.aay2631

> A system that outputs the same near identical websites for entirely different prompts by entirely different people half way around the world is not a system capable of scientific advancement in any meaningful sense.

Similarity among generic website outputs says little about the range reachable under scientific search and evaluation, as FunSearch generated previously unknown mathematical constructions that surpassed the best known results.

Romera et al., Nature, 2024 https://doi.org/10.1038/s41586-023-06924-6

> Second, the immense resources required to train a single AI will tend to push them to be forces of stagnation.

Frontier pretraining is expensive and may concentrate model development, but stagnation does not follow because pretrained models can be adapted to new tasks while changing only a small fraction of their paramters.

Hu et al., ICLR, 2022 https://arxiv.org/abs/2106.09685

> New fields and new approaches can really only be incorporated into an LLM by retraining them with that data set.

This is false because retrieval augmented generation allows a pretrained model to use new external knowledge at inference time without retraining its underlying parameters.

Lewis et al., NeurIPS, 2020 https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html

> But because retraining is such an expensive and involved process, with LLMs there will be additional costs and efforts required to adopt any novel fields into a scientific institution built on an LLM.

The proposed bottleneck does not follow because new papers, databases, and entire bodies of specialized knowledge can be supplied through retrieval without another full training run.

Lewis et al., NeurIPS, 2020 https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html

> One of the biggest hoaxes of modern AI imo is to call them 'generative' systems, when they are one of the least generative systems we know of.

The argument replaces the technical meaning of generative modeling with an invented measure based on implementation size, whereas generative language models learn probability distributions over sequences and produce new samples from those distrubutions.

Bengio et al., Journal of Machine Learning Research, 2003 https://www.jmlr.org/papers/v3/bengio03a.html

> A simple multiplication function is more generative than an LLM because its implementation is maybe a few bytes, and its capable of generating an infinite set with that few byte implementation.

An infinite range is not a measure of generative modeling, otherwise a trivial counter would be more generative than any learned model despite representing no learned distribution at all.

Bengio et al., Journal of Machine Learning Research, 2003 https://www.jmlr.org/papers/v3/bengio03a.html

> These LLMs are about terrabyte implementations.

Some frontier checkpoints can reach that order of storage depending on parameter count and numerical precision, but checkpoint size has no logical bearing on whether the learned function is generative.

Hoffmann et al., NeurIPS, 2022 https://proceedings.neurips.cc/paper_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf

> To put that into context, I can store humanities entire public domain written outputs -- what makes up a vast amount of the training sets of these LLMs -- on less than 1 terrabyte.

The comparison is not meaningful because a compressed archive of documents and a learned parameterization are different computational objects, and modern language models are trained on token counts far exceeding their paramater counts.

Hoffmann et al., another from NeurIPS, 2022 https://proceedings.neurips.cc/paper_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf

> Literally I've done this. I have much of stack exchange, all of Wikipedia including images, all of the public domain English books, and a bunch of other stuff, less than 1 terabyte.

That may be entirely true, but the size of an archive provides no evidence that a trained transformer is a lookup table because transformers can infer new task specific mappings from examples supplied only at inference time.

von Oswald et al., ICML, 2023 https://proceedings.mlr.press/v202/von-oswald23a.html

> These frontier LLMs are not much smaller than their training sets; they are more generative than a lookup table; the least generative function possible. They generate very little.

The comparison conflates corpus storage with learned parameters and inference with retrieval, while modern transformers can generalize to mappings not stored as prompt and answer pairs and LLM based search has produced previously unknown mathematical construtions.

Romera et al., Nature, 2024 https://doi.org/10.1038/s41586-023-06924-6

> Most of what they do is information recall.

No evidence is offered for the quantitative claim "most," and mechanistic experiments show that transformers can learn task specific algorithms from context during inference rather than merely retrieve stored associations.

von Oswald et al., ICML, 2023 https://proceedings.mlr.press/v202/von-oswald23a.html

> They are very impressive advancements in data compression and information retrieval.

Compression and retrieval are genuine capabilities, but they are not an exhaustive account of systems that have generated new protein sequences whose biological properties were subsequently established by laboratory experiemnt.

Swanson et al., Nature, 2025 https://doi.org/10.1038/s41586-025-09442-9

Snackatron | 5 hours ago

You absolutely destroyed him hahaha 💯👍

Adventurous-Ad281 | 3 hours ago

They didn’t destroy him, because they didn’t actually provide any argument other than appeal to authority (logical fallacy).

Key_Minute120 | 3 hours ago

Should have used this in college, “Professor I didn’t add citations because those are just appeals to authority”

Adventurous-Ad281 | an hour ago

You know what a citation means right? This is not it.

__JDQ__ | 2 hours ago

Ah, yes, citing support for their claims. Classic ‘Appeal to Authority’ fallacy…

You may want to review the definitions of the various logical fallacies.

Needed to edit it more but my eyes are almost glued shut. Probably should have just listed links at the end. Whatever.

CognitionMass | 2 hours ago

Just the first link I looked at and found it contradicted the claim you attached to it. You claimed:

>LLMs can produce convergent outputs under some conditions, but experiments measuring divergent creativity show that convergence is not an intrinsic limit on the range of ideas they can genreate.

The link you attached this claim of yours to contradicts you:

> Notably, even the top performing LLMs are still largely surpassed by the aggregated top half of human participants, underscoring a ceiling that current LLMs still fail to surpass.

Let me emphasise that: a ceiling that current LLMs still fail to surpass. That's explicit support of my claim.

And this is just creative writing test. Scientific advancement requires far more divergent thinking.

Should I look at another bit of this gish gallop?

Ok_Suggestion5523 | an hour ago

It's a bit like they used an llm to write the post.

CognitionMass | an hour ago

It is isn't it.

Adventurous-Ad281 | 3 hours ago

Greatest strawman of all time. You didn’t even adress one of their points: that science will become privatized, and that is inherently detrimental and amost contradictory to the spirit that has moved science for the last three milenia. Nobody denies AI can produce experimental predictions with current data, nor that it can extend mathematical knowledge with currenly known techinques (e.g, the recent Navier-Stokes proof, but in fact, any of the proofs doneby AI have their origin in known techniques, either mainstream or from obscure papers, and it just blends them together by raw computational power until it obtains a solution. So nothing you said proves absolutely anything. Appealing to authority is a logical fallacy, as is a strawman. Instead of actually presenting a crafted argument, you just spew link after link of “evidence” without providing context as to which terms mean what, expecting people to accept your reasoning, when you really haven’t provided any.

Weekly-Ad353 | 8 hours ago

Tell me you’ve never tried to do science with one without telling me you’ve never tried to do science with one.

truecakesnake | 5 hours ago

That comment is so stupid, I have no idea how that word slop got upvoted.

Really, because AI generates similar websites it can't make advancements in science. u/CognitionMass the reason AI website designs are similar is because the models are post trained on templates and they're trained to match these templates. It's just a irrelevant point.

And, what? Just because most of what they do is context handling they're not generative? That makes no sense. They're fundamentally generative, they predict theh generate the next likely word.

They've already proven that LLMs are capable of scientific advancement. This is easily googlable.

monkeydrunker | 6 hours ago

> Second, the immense resources required to train a single AI will tend to push them to be forces of stagnation.

And they don't even scale linearly for training. They scale logarthmically, at best. Meaning that doubling the resources provides less and less benefit as it scales. This is why a bunch of money is now being dumped into small, specific models to fix small, specific problems.

Ok_Suggestion5523 | an hour ago

You're correct about the coding bits, it's worth highlighting that for all the seasons you highlighted, it's terrible at novel solutions to programming. It's also terrible at optimal solutions and coding for future versions of yourself.

SweetMnemes | 4 hours ago

For me it is really sad to see a comment that is almost entirely composed of falsehoods to be upvoted so high simply because it expresses anti AI sentiment. Science is almost entirely about standardization this shields it and distinguishes it from mere gut feelings. Large-scale science always consumes a lot of resources. It is entirely obvious at this point that the output of AI systems is highly creative. Just because the brain is composed of more synapses than there are stars in the galaxy makes the brain no less generative. So the main issue here is that the corporate AI are owned by corporations which makes them incompatible with science not in principle due to the nature of market competition because companies want to keep their secrets which give them advantages over competitors while science is all about openness and transparency. These are cornerstones of science because they allow your method to be understood and thus criticized. In science, this creates checks and balances which lead to a cooperative way towards improvements that can be shared among all mankind while also keeping the risks in check. Or at least they can be openly discussed and understood.

CognitionMass | an hour ago

have you read Thomas Khun or any work on history of science?

Science is not at all formalised, and you don't seem to claim otherwise, so we'll move past that. But formalisation is the main point, standardisation secondary. Standardisation is of course important to science, particularly normal science. But this is mostly in narrow areas of normal science. If you start to look at the broader scientific history, and advancement of science, the area khun more called paradigm shift, there very little in the way of standardisation and certainly no formalisation at all.

BookProper9115 | 8 hours ago

It's the death of our current 'real' science and the beginning of a new paradigm. Exciting times to be a philosopher. Philosophy of science about to explode.

I suggest reading some Thomas Khun and Karl Popper.

>Karl Popper

My guy. 🤝

CognitionMass | 2 hours ago

I was thinking exactly of khun when I wrote that, particularly how important paradigm shift is to scientific advancement, and how much of a force against paradigm shift LLMs would be if they were given too much control.

OGLikeablefellow | 6 hours ago

Do you cycle your 1 terabyte of human knowledge? The medium can corrupt if you don't move it around every now and then

EOE97 | 2 hours ago

Everyone uses previous information, even the so called divergent thinking humans.

AI may not be at human level divergent thinking at every domain, but it's innacurate to say everything LLMs do is mere recall of training data. We wouldn't have the degree of mathematical breakthroughs for unsolved problem like we do today, which literally requires divergent thinking.

You can't CTRL+F these answers or randomly generate any meaningful solution. The pattern finding algorithms of LLMs makes these breakthroughs possible.

CognitionMass | an hour ago

I agree, that would be inaccurate to say, and I didn't say that.

Mathematics is highly formalised, where LLMs excel. And behind their hyped up headlines, their contributions are more as tools than participants. Doesn't really engage with anything I said.

0impulsecontrol | 8 hours ago

this is a word salad

sjogren | 8 hours ago

It's very well written and on point.

0impulsecontrol | 8 hours ago

lol this sub

truecakesnake | 5 hours ago

All mainstream subreddits are so stupid

burnttoast12321 | 7 hours ago

So you have all of Wikipedia and Stack Exchange saved in a smaller size. How are you going to do anything with it? You just made a backup.

CognitionMass | an hour ago

As I said at the end of my comment:

> They are very impressive advancements in data compression and information retrieval.

Kuriente | 7 hours ago

>First, they are forces of standardisation. A scientific enterprise entirely built around Claude for example is one that's incapable of divergent and novel thinking.

How does this fit with AI's known tendency to hallucinate?

Personally, I've always viewed AI hallucinations as useful for science (and art). You could build a model optimized for hallucination and potentially harness its error output as a way of generating novel hypotheses, and then use a more rigorous standardized model to analyze each hypothesis for experimentation.

Although, I tend to believe a generative hypothesis model might run best on an analog computer. They are more error prone and use a fraction of the energy. I'm not aware of instances where it's been specifically used this way, digital or analog, but I don't see how it couldn't be.

I actually believe this is exactly how the human brain and scientific institutions work. We irrationally generate divergent information to induce ideas (inductive reasoning) and then analyze it rationally with logic and experimentation to infer the truth (deductive reasoning).

xbhaskarx | an hour ago

In related news…

https://x.com/valeriocapraro/status/2097791836269977996