I think it's not architectural. RLVR supplies a counterexample to the claim that post-training necessarily optimizes only the appearance of truth.
Suppose a model generates a formal proof and receives reward only when an independent proof checker verifies every inference. The successful output is selected because it satisfies an external, mechanically enforced truth condition.
Repeated reinforcement can therefore establish a non-accidental truth-tracking relation between the inference policy and a formal domain: invalid arguments receive no reward, however fluent they sound.
RLVR does not by itself establish belief, understanding, or Aristotelian knowledge, but it shows that statistical learning need not remain confined to plausible opinion; its outputs can be causally governed by verification rather than linguistic approval.
Good counterexample, but it shifts the burden from the LLM to the verifier. The model did not learn the proof; it learned to produce tokens that satisfy the checker. That is not accidental, but it is also not understanding — it is a more constrained form of δόξα. The truth-tracking relation is real, yet external to the model's own epistemic structure.
Token prediction is a task, it is not a measure of what is going on inside the model. You cant make a conclusion about what is going on inside the model just based on that goal.
If theodel cannot distinguish an external source of validation, and the external mechanical demand for truth distinguishes a kind of preference not an approach to truth, doesn't that just allow a move and not solution to the problem?
Yes, exactly. External verification does not give the model a path to truth. It only gives it a stricter preference. The model still does not know why the proof works. It only knows that the proof checker accepted it. The Gettier problem remains because the truth-condition is satisfied accidentally relative to the model's own reasoning. The approach changes the training signal, not the epistemic structure.
This was awesome, thanks for posting this. I did have a minor doubt about one bit:
> Furthermore, because modern LLMs are trained on vast web corpora that contain both valid scientific consensus and unverified speculation, model parameterization collapses distinct epistemic categories into single vector space.
While it sounds likely, you've asserted it with some authority, and I wonder what evidence or literature supports it. If this were wikipedia I might give it a "citation needed" tag 😁
So is there a citation which proves or demonstrates the collapse of valid and unverified claims into a single subspace? Or were you more trying to be "poetic" rather than rigorous?
I don't think this is controversial enough to demand specific citation, that llm training treats all written work as a singular linguistic expression is fundamental to the technology. The whole mechanism of transform coding and assessment that llms use supports the authors suggestion that it cannot distinguish categories of truthful scientific or pseudo or false information. In fact the very basic nature of the technology is based on frequency assessment so that if a lie or false belief is repeated more in the sampled language set it will be considered more valid in the model.
But a good example of how this problem is endemic if not integral, is the problems of alignment, threshold, and parameter setting in the technology that also make simplification decisions necessary, widely discussed in why strawberry is often described as having only 2 Rs by AI. In making tokens all knowledge in prress becomes informational tokens that are changed into a flattened statistical form. That form is not distinguishing science vs comic book vs Facebook comment. Maybe in an alternate work it could be. But as it is you are asking them to defend a description of the technology that isn't arguable?
You are right that the mechanism itself is not controversial. My "collapse into single vector space" was a conceptual compression, not a measured claim. The transformer does not literally merge science and rumor into one subspace; it assigns probabilities over tokens without epistemic labels. The point is that the architecture has no native way to mark this came from a peer-reviewed source versus this was repeated often. I should have said that more carefully.
Fair point. I was drawing on the general observation that web corpora mix reliable sources with misinformation, and models treat them as equally valid training signals. "Collapse into single vector space" was a conceptual claim, not a rigorously tested one. I should flag it as heuristic, not established.
My understanding is that LLMs tend to mirror input in such a way that if you provide them with a well-structured prompt employing domain keywords, then you tend to activate domain-knowledge weights. Of course, output veracity is best verified by experts.
As you argue, however, taking into account that the mechanism is probabilistic, that the vector space is subject to interpolation and that LLMs are compelled by training to return an output even in the face of insufficient training data to warrant justified belief, there is a tendency for LLMs to produce doxa.
Nevertheless, it should be noted that also humans tend to produce doxa because knowledge is limited--a more pressing shortcoming for an individual compared to an LLM given that the former cannot hold the same amount of information the latter can--and it is easy to fall prey of cognitive biases.
The fundamental flaw of LLMs, as I see it, is that (i) the world is not made of words, meaning that the emulated reasoning is based on an indirect and imperfect description of the world rather than a direct observation of it with little or no understanding of the underlying physical regularities; and (ii) they appear to be inherently incapable of conducting knowledge leaps outside the boundaries of the training material, that is to say, they may not be able to properly conduct genuine or emulated abductive reasoning.
Even with these flaws, however, LLMs generally yield output which may constitute a basis for further refinement after proper verification, which renders them a powerful tool, but whose limitations, admittedly, need to be better communicated to users who may think they are being served distilled episteme.
You may be interested by Bengio's "Scientist AI" aims (blogpost, paper).
> Much like ideal scientific theories, the Predictor aims to accurately and neutrally model the world as it actually is. And because the Predictor is never trained to act in the world or play any sort of conversational role, it is very much not an assistant model.
I personally agree with their goal, but I have doubts on their means.
The Predictor idea is interesting precisely because it removes the assistant role. My worry remains the same: even a world-model trained to predict is still optimizing for likelihood over tokens, not for being right in the world. The gap between good prediction and true understanding does not close automatically.
My understanding is that they do train on internet data, but they treat it as "probability that the next token is X in context Y". Then they also have specially annotated data of the form "probability that Q believes that R is true".
It contains true propositions, but not the living act of knowing. A person reading a true sentence may know it; the encyclopedia stores the proposition without possessing νοῦς. LLMs are closer to the encyclopedia than to the knower.
While I agree that many current LLMs suffer from a bias toward agreeableness and sycophancy, I don't think this is an inherent property of statistical reasoning models. Human brains, from our limited knowledge of how they operate, seem to use weighting (in the form of stronger or weaker connections between axons) to accomplish both "knowing facts" and "processing knowledge". In order to get a human brain to state true propositions "with proper causal justification grounded in reality", the brain in question must undergo rigorous training in that topic area, much of which involves the brain making statements which are corrected by a different brain that has already undergone such training (i.e. a teacher). I agree with you that the adversarial nature of this process tends to produce brains with internal weighting systems that consider logic, causal justification, and empirical justification as paramount over agreeableness. But I don't think these brains have "internal states corresponding to belief" that are separate and distinct from the weighted axonal connections between their neurons. What else would such internal states consist of?
If it's possible to train a brain to generate non-accidentally true statements of justified beliefs, I see no reason why the same can't be accomplished with LLMs. I think you are justified in your belief that capitalistic pressures are currently limiting our training practices and biasing them toward agreeableness. But if an LLM was rigorously trained and corrected by domain experts rather than "human evaluators working under time pressure", I see no reason to believe that they couldn't achieve epistemic fidelity in that domain. An appropriately trained LLM, which itself had experienced pushback against subtle misconceptions or false premises, would be fully capable of pushing back against future users in the same way.
The training methods most LLMs currently use create biases, but I don't think statistical learning itself inherently limits LLMs toward mere mimicry. If statistical inference does have concrete limits, I believe those limits are also present in humans.
The analogy is tempting, but it hides a difference. Neurons in a living brain are anchored to the world through perception, action, and error correction. A human learning geometry can draw the triangle, measure it, and be corrected by the figure itself. An LLM has only correlations among tokens. It can mimic the proof, but it never touches the triangle. Statistical learning may be part of human cognition, but νοῦς needs more than weights — it needs grounded intentionality.
A difference without a distinction, I think. Surely modern LLMs that can accept pictures as inputs have something verging on visual perception. Google's DeepMind has tools to perceive and model 3d reality.
If your claim is that LLMs that are only trained on and only process text-based information are incapable of having beliefs that are grounded in non-text-based information, I agree, but I think that's a trivial statement.
If you're making the stronger statement that LLM-style statistical inference can never ground any beliefs in non-text-based information because it is theoretically incapable of such grounding, I disagree. Humans have made huge leaps in scientific understanding based solely on data from instruments that measure things we cannot ourselves see or touch. There's no reason to think machine-based statistical inference models can't do this as well. Of course they may not currently do this, but I see no theoretical barrier given the right kinds of input and training.
Aristotle would separate εἰκάζειν (conjecture), πίστις (conviction), and ἐπιστήμη (demonstrative knowledge). LLMs operate mostly in the first two. Your "purpose-adapted information" sounds closer to τέχνη than to knowledge proper. The danger is mistaking fitness for a task with truth about the task.
You seem to be conflating how LLMs are actually trained with how they have to be trained. They don't have to be trained as you suggest. You are also assuming an internalist epistemology. It doesn't seem implausible that an LLM trained on a carefully curated dataset could have knowledge under an externalist account of knowledge
One of the contra-indications for me was a lot of the grammar seemed imperfect. Usually AI writes in a fairly generic way without these minor deviations from fairly standard english.
For example:
"From perspective of ..." rather than "From the perspective of ..."
"For statement to count as knowledge ..." rather than "For a statement to count as knowledge ..."
^ There were quite a few instances of this sort of thing which is common in writing by people who don't have English as a first language.
Ironically, your suspicion makes the exact point of the post. We now judge a text by how it seems, not by whether it tracks truth. The missing articles are not slop, they are Greek syntax showing through my English.
I don't think it was written by an LLM. To me it feels like the chain of thought that a person would follow, plus, some grammatical and vocabulary mistakes give away a non-native english speaker (or a very artful facewash of an LLM output, but I think this is less likely).
It's not slop though, which I think is an important distinction. There's been a lot of stuff posted here lately that is essentially a shower thought run through an LLM to give the veneer of philosophical argumentation. This is a coherent argument that builds on existing philosophy and applies it to a novel domain in a relevant way. If you're going to use an LLM, this is the way to do it.
yall_gotta_move | 16 hours ago
I think it's not architectural. RLVR supplies a counterexample to the claim that post-training necessarily optimizes only the appearance of truth.
Suppose a model generates a formal proof and receives reward only when an independent proof checker verifies every inference. The successful output is selected because it satisfies an external, mechanically enforced truth condition.
Repeated reinforcement can therefore establish a non-accidental truth-tracking relation between the inference policy and a formal domain: invalid arguments receive no reward, however fluent they sound.
RLVR does not by itself establish belief, understanding, or Aristotelian knowledge, but it shows that statistical learning need not remain confined to plausible opinion; its outputs can be causally governed by verification rather than linguistic approval.
[OP] vasilisvj | 6 hours ago
Good counterexample, but it shifts the burden from the LLM to the verifier. The model did not learn the proof; it learned to produce tokens that satisfy the checker. That is not accidental, but it is also not understanding — it is a more constrained form of δόξα. The truth-tracking relation is real, yet external to the model's own epistemic structure.
Aggressive-Share-363 | 4 hours ago
I think k you are making a common fallacy.
Token prediction is a task, it is not a measure of what is going on inside the model. You cant make a conclusion about what is going on inside the model just based on that goal.
Necessary_Cat_5662 | 8 hours ago
If theodel cannot distinguish an external source of validation, and the external mechanical demand for truth distinguishes a kind of preference not an approach to truth, doesn't that just allow a move and not solution to the problem?
[OP] vasilisvj | 6 hours ago
Yes, exactly. External verification does not give the model a path to truth. It only gives it a stricter preference. The model still does not know why the proof works. It only knows that the proof checker accepted it. The Gettier problem remains because the truth-condition is satisfied accidentally relative to the model's own reasoning. The approach changes the training signal, not the epistemic structure.
trinarynimbus | 15 hours ago
This was awesome, thanks for posting this. I did have a minor doubt about one bit:
> Furthermore, because modern LLMs are trained on vast web corpora that contain both valid scientific consensus and unverified speculation, model parameterization collapses distinct epistemic categories into single vector space.
While it sounds likely, you've asserted it with some authority, and I wonder what evidence or literature supports it. If this were wikipedia I might give it a "citation needed" tag 😁
So is there a citation which proves or demonstrates the collapse of valid and unverified claims into a single subspace? Or were you more trying to be "poetic" rather than rigorous?
Necessary_Cat_5662 | 8 hours ago
I don't think this is controversial enough to demand specific citation, that llm training treats all written work as a singular linguistic expression is fundamental to the technology. The whole mechanism of transform coding and assessment that llms use supports the authors suggestion that it cannot distinguish categories of truthful scientific or pseudo or false information. In fact the very basic nature of the technology is based on frequency assessment so that if a lie or false belief is repeated more in the sampled language set it will be considered more valid in the model.
But a good example of how this problem is endemic if not integral, is the problems of alignment, threshold, and parameter setting in the technology that also make simplification decisions necessary, widely discussed in why strawberry is often described as having only 2 Rs by AI. In making tokens all knowledge in prress becomes informational tokens that are changed into a flattened statistical form. That form is not distinguishing science vs comic book vs Facebook comment. Maybe in an alternate work it could be. But as it is you are asking them to defend a description of the technology that isn't arguable?
[OP] vasilisvj | 6 hours ago
You are right that the mechanism itself is not controversial. My "collapse into single vector space" was a conceptual compression, not a measured claim. The transformer does not literally merge science and rumor into one subspace; it assigns probabilities over tokens without epistemic labels. The point is that the architecture has no native way to mark this came from a peer-reviewed source versus this was repeated often. I should have said that more carefully.
[OP] vasilisvj | 6 hours ago
Fair point. I was drawing on the general observation that web corpora mix reliable sources with misinformation, and models treat them as equally valid training signals. "Collapse into single vector space" was a conceptual claim, not a rigorously tested one. I should flag it as heuristic, not established.
In_der_Tat | 2 hours ago
My understanding is that LLMs tend to mirror input in such a way that if you provide them with a well-structured prompt employing domain keywords, then you tend to activate domain-knowledge weights. Of course, output veracity is best verified by experts.
As you argue, however, taking into account that the mechanism is probabilistic, that the vector space is subject to interpolation and that LLMs are compelled by training to return an output even in the face of insufficient training data to warrant justified belief, there is a tendency for LLMs to produce doxa.
Nevertheless, it should be noted that also humans tend to produce doxa because knowledge is limited--a more pressing shortcoming for an individual compared to an LLM given that the former cannot hold the same amount of information the latter can--and it is easy to fall prey of cognitive biases.
The fundamental flaw of LLMs, as I see it, is that (i) the world is not made of words, meaning that the emulated reasoning is based on an indirect and imperfect description of the world rather than a direct observation of it with little or no understanding of the underlying physical regularities; and (ii) they appear to be inherently incapable of conducting knowledge leaps outside the boundaries of the training material, that is to say, they may not be able to properly conduct genuine or emulated abductive reasoning.
Even with these flaws, however, LLMs generally yield output which may constitute a basis for further refinement after proper verification, which renders them a powerful tool, but whose limitations, admittedly, need to be better communicated to users who may think they are being served distilled episteme.
RockManChristmas | 10 hours ago
You may be interested by Bengio's "Scientist AI" aims (blogpost, paper).
> Much like ideal scientific theories, the Predictor aims to accurately and neutrally model the world as it actually is. And because the Predictor is never trained to act in the world or play any sort of conversational role, it is very much not an assistant model.
I personally agree with their goal, but I have doubts on their means.
[OP] vasilisvj | 6 hours ago
The Predictor idea is interesting precisely because it removes the assistant role. My worry remains the same: even a world-model trained to predict is still optimizing for likelihood over tokens, not for being right in the world. The gap between good prediction and true understanding does not close automatically.
RockManChristmas | 5 hours ago
> is still optimizing for likelihood over tokens
My understanding is that they do train on internet data, but they treat it as "probability that the next token is X in context Y". Then they also have specially annotated data of the form "probability that Q believes that R is true".
fox-mcleod | 9 hours ago
Does an encyclopedia contain knowledge?
[OP] vasilisvj | 6 hours ago
It contains true propositions, but not the living act of knowing. A person reading a true sentence may know it; the encyclopedia stores the proposition without possessing νοῦς. LLMs are closer to the encyclopedia than to the knower.
wizkid123 | 8 hours ago
While I agree that many current LLMs suffer from a bias toward agreeableness and sycophancy, I don't think this is an inherent property of statistical reasoning models. Human brains, from our limited knowledge of how they operate, seem to use weighting (in the form of stronger or weaker connections between axons) to accomplish both "knowing facts" and "processing knowledge". In order to get a human brain to state true propositions "with proper causal justification grounded in reality", the brain in question must undergo rigorous training in that topic area, much of which involves the brain making statements which are corrected by a different brain that has already undergone such training (i.e. a teacher). I agree with you that the adversarial nature of this process tends to produce brains with internal weighting systems that consider logic, causal justification, and empirical justification as paramount over agreeableness. But I don't think these brains have "internal states corresponding to belief" that are separate and distinct from the weighted axonal connections between their neurons. What else would such internal states consist of?
If it's possible to train a brain to generate non-accidentally true statements of justified beliefs, I see no reason why the same can't be accomplished with LLMs. I think you are justified in your belief that capitalistic pressures are currently limiting our training practices and biasing them toward agreeableness. But if an LLM was rigorously trained and corrected by domain experts rather than "human evaluators working under time pressure", I see no reason to believe that they couldn't achieve epistemic fidelity in that domain. An appropriately trained LLM, which itself had experienced pushback against subtle misconceptions or false premises, would be fully capable of pushing back against future users in the same way.
The training methods most LLMs currently use create biases, but I don't think statistical learning itself inherently limits LLMs toward mere mimicry. If statistical inference does have concrete limits, I believe those limits are also present in humans.
[OP] vasilisvj | 6 hours ago
The analogy is tempting, but it hides a difference. Neurons in a living brain are anchored to the world through perception, action, and error correction. A human learning geometry can draw the triangle, measure it, and be corrected by the figure itself. An LLM has only correlations among tokens. It can mimic the proof, but it never touches the triangle. Statistical learning may be part of human cognition, but νοῦς needs more than weights — it needs grounded intentionality.
wizkid123 | 5 hours ago
A difference without a distinction, I think. Surely modern LLMs that can accept pictures as inputs have something verging on visual perception. Google's DeepMind has tools to perceive and model 3d reality.
If your claim is that LLMs that are only trained on and only process text-based information are incapable of having beliefs that are grounded in non-text-based information, I agree, but I think that's a trivial statement.
If you're making the stronger statement that LLM-style statistical inference can never ground any beliefs in non-text-based information because it is theoretically incapable of such grounding, I disagree. Humans have made huge leaps in scientific understanding based solely on data from instruments that measure things we cannot ourselves see or touch. There's no reason to think machine-based statistical inference models can't do this as well. Of course they may not currently do this, but I see no theoretical barrier given the right kinds of input and training.
fudge_mokey | 6 hours ago
>It produces δόξα (opinion or belief), specifically calibrated to look like justified true belief.
Can you explain more on how you justify a true belief?
I don't think there are any known methods to justify or validate an idea as true, or even probably true.
I think a better definition for knowledge is information adapted to a purpose.
[OP] vasilisvj | 6 hours ago
Aristotle would separate εἰκάζειν (conjecture), πίστις (conviction), and ἐπιστήμη (demonstrative knowledge). LLMs operate mostly in the first two. Your "purpose-adapted information" sounds closer to τέχνη than to knowledge proper. The danger is mistaking fitness for a task with truth about the task.
fudge_mokey | 6 hours ago
I think it's interesting to learn about how Aristotle might separate things, but I think we've made progress on the problem since then.
Are Newton's laws of motion "knowledge proper"? Or only τέχνη?
I think they are knowledge because they are information adapted to a purpose. They contain some truth, in addition to some errors or falsehoods.
But you didn't answer my most important question. If knowledge is justified, true belief, then how does one justify a belief as being true?
There are no known methods to justify or validate an idea as true.
Far_Course2496 | 4 hours ago
You seem to be conflating how LLMs are actually trained with how they have to be trained. They don't have to be trained as you suggest. You are also assuming an internalist epistemology. It doesn't seem implausible that an LLM trained on a carefully curated dataset could have knowledge under an externalist account of knowledge
bubibubibu | 15 hours ago
Irony is that this was written by llm
trinarynimbus | 15 hours ago
It didn't come across that way to me. What makes you think so?
bubibubibu | 15 hours ago
"pure Gettier case embedded", the cadence, paragraphs, long winded without saying anything...
trinarynimbus | 15 hours ago
Hmm, I can see that.
One of the contra-indications for me was a lot of the grammar seemed imperfect. Usually AI writes in a fairly generic way without these minor deviations from fairly standard english.
For example:
^ There were quite a few instances of this sort of thing which is common in writing by people who don't have English as a first language.
bubibubibu | 13 hours ago
Yeah sure, I might be wrong...it just gives me the same feeling as an ai wall of text slop
Attackoftheglobules | 9 hours ago
Nah i think this is just someone writing their thoughts.
[OP] vasilisvj | 6 hours ago
Ironically, your suspicion makes the exact point of the post. We now judge a text by how it seems, not by whether it tracks truth. The missing articles are not slop, they are Greek syntax showing through my English.
ellisftw | 15 hours ago
Go back and read it and try to find the word "THE".
trinarynimbus | 15 hours ago
Yes, exactly. It comes across as a writer who isn't a native English speaker. I'm not aware that LLMs never use "the"?
[OP] vasilisvj | 6 hours ago
That is the giveaway. An LLM does not forget articles; a non-native speaker does.
CGY97 | 9 hours ago
I don't think it was written by an LLM. To me it feels like the chain of thought that a person would follow, plus, some grammatical and vocabulary mistakes give away a non-native english speaker (or a very artful facewash of an LLM output, but I think this is less likely).
wizkid123 | 9 hours ago
It's not slop though, which I think is an important distinction. There's been a lot of stuff posted here lately that is essentially a shower thought run through an LLM to give the veneer of philosophical argumentation. This is a coherent argument that builds on existing philosophy and applies it to a novel domain in a relevant way. If you're going to use an LLM, this is the way to do it.
[OP] vasilisvj | 6 hours ago
Thanks for reading it charitably. I will take "coherent argument" over "native English" any day.