There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so.
This is fantastic, and in a healthier world this would be made law.
Excellent article, enough ideas to justify several, but I'm not complaining.
I have yet to hear from a single person who has said “yeah, we measured according to your methodology
Me neither, and I continue to wonder why. I feel like even I, with my limited scientific background, could run an experiment with reasonable, non-perfect controls that would be good enough for a corporate environment. I suppose the fact that I haven't means that either everyone else is sharing my reason for not having done so (overwhelming fatigue with the space), or it's just much harder than I am imagining.
Fatigue definitely has something to do with it. Also, many folks have already made up their minds. Suppose you're an AI promoter who's a genuine believer—you're not being paid to promote it or anything like that. From that person's point of view, whatever they've already observed must have looked like this, so what's the point of layering on additional scientific rigor? It will either confirm your existing beliefs, or you will think the experiment must be faulty in some way: maybe it's measuring the wrong thing, or maybe those non-perfect controls actually mattered. But your existing beliefs won't update.
Text editor analogy: a study showed editing text with keyboard+mouse is faster than keyboard alone, even though it feels slower. Except apparently that study was bogus? I dunno. The only thing I do know is I've never met anyone who read any of the research and then changed their behavior because of that: e.g., switched to/away from a modal editor.
(I do think we should do more science! I'm just pessimistic that it will change minds.)
This is a good article, but just getting LLMs to present a list of citations is not enough, because those citations will be cherry-picked. The only way to be sure you are getting a balanced view is to actually do your own research, at which point the LLM becomes useless.
This will obviously vary from person to person and task to task, but the question is: what is the acceptable level of accuracy?
It's never going to be 100% perfect. You could hire me, a bona fide human being with pure intentions, to do your research and provide you with a list of citations. But if I haven't had enough coffee, I might provide a citation that doesn't fully support what I think it does. And even if I was perfect, my source could be wrong.
In terms of rigorously maximizing accuracy, I think the closest we currently get is stuff like LLMs generating Lean proofs. There's still the matter of mathematicians verifying that the proofs prove what we think they prove, of course—I guess the bet there is that humans reading the proof will be easier than writing it from scratch.
kaig | 4 hours ago
I really like this article. In particular:
This is fantastic, and in a healthier world this would be made law.
elliotmorris | 3 hours ago
Excellent article, enough ideas to justify several, but I'm not complaining.
Me neither, and I continue to wonder why. I feel like even I, with my limited scientific background, could run an experiment with reasonable, non-perfect controls that would be good enough for a corporate environment. I suppose the fact that I haven't means that either everyone else is sharing my reason for not having done so (overwhelming fatigue with the space), or it's just much harder than I am imagining.
bitshift | 2 hours ago
Fatigue definitely has something to do with it. Also, many folks have already made up their minds. Suppose you're an AI promoter who's a genuine believer—you're not being paid to promote it or anything like that. From that person's point of view, whatever they've already observed must have looked like this, so what's the point of layering on additional scientific rigor? It will either confirm your existing beliefs, or you will think the experiment must be faulty in some way: maybe it's measuring the wrong thing, or maybe those non-perfect controls actually mattered. But your existing beliefs won't update.
Text editor analogy: a study showed editing text with keyboard+mouse is faster than keyboard alone, even though it feels slower. Except apparently that study was bogus? I dunno. The only thing I do know is I've never met anyone who read any of the research and then changed their behavior because of that: e.g., switched to/away from a modal editor.
(I do think we should do more science! I'm just pessimistic that it will change minds.)
neilmadden | 2 hours ago
This is a good article, but just getting LLMs to present a list of citations is not enough, because those citations will be cherry-picked. The only way to be sure you are getting a balanced view is to actually do your own research, at which point the LLM becomes useless.
bitshift | 57 minutes ago
This will obviously vary from person to person and task to task, but the question is: what is the acceptable level of accuracy?
It's never going to be 100% perfect. You could hire me, a bona fide human being with pure intentions, to do your research and provide you with a list of citations. But if I haven't had enough coffee, I might provide a citation that doesn't fully support what I think it does. And even if I was perfect, my source could be wrong.
In terms of rigorously maximizing accuracy, I think the closest we currently get is stuff like LLMs generating Lean proofs. There's still the matter of mathematicians verifying that the proofs prove what we think they prove, of course—I guess the bet there is that humans reading the proof will be easier than writing it from scratch.