This is measuring on scientific tasks though. I haven't used Fable lately but when I was playing with it when it was new its "safety" features made it practically impossible to use for biomedical science -- it once refused to work on a pipeline of mine that was analyzing pathogenicity islands in bacteria (presumably because it had guardrails to that effect to stop potential bioterrorists and the like)
IMHO it's roughly task depth (fable) vs breadth (opus). Fable is great at tracing and debugging sometimes, but otherwise shorthand for confabulation. It's persistent but ungovernable, struggles to switch contexts, and goes insane with too much freedom to explore. Don't point it at anything that looks like a "system" for actual work (but mapping or planning might be ok). Opus is maybe not as creative, but it's more stable and more trustworthy. Opus driving Fable could be awesome, but Fable unleashed/unsupervised on longer horizon tasks or things that require more methodical changes on lots of components seems like a disaster every time I try it.
Since this kind of thing is always down to harness, project-type, and other structural constraints, of course your mileage may vary. Fable is probably great for pen-testing, or as a decision-making kernel of other kinds of applications, and way better than Opus at those things. Probably fine for code-review or changing a codebase of a few thousand lines in any language. Actually building that codebase or changing an even bigger one? Woof.
How this fits in with science? IDK but I bet other existing causal reasoning benchmarks might tell the whole story and this is back to stability again. Sometimes having a smart idea is really important! But more often it's important to just not forget what you were doing. What was I talking about? Oh look a squirrel
Not surprised to see Claude significantly higher in scientific intelligence than Sol.
You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.
Don't get me wrong, codex is great at finding bugs and building games. It's great.
I wonder how long it's going to be before self improvement encompasses hardware and materials science, not just code. It's exciting, soon we'll be able to fully hand off scientific, mathematical, and technical progress over to the machines, and then we can fully lay back.
I got pretty good mileage out of context engineering, adding my personal coding heuristics to my AGENTS.md and referencing subdocuments on a "when doing X, consult Y" pattern. I assume others are doing similar things, but I was pretty surprised when I was able to get it to generate code that is pretty close to what I would do if I was doing it manually. I'm curious if scientists and mathematicians are doing things like that. "When I see X, I typically immediately check Y" or whatever their domain heuristics look like.
rubslopes | 4 hours ago
akshay_akula | 4 hours ago
vatsachak | 4 hours ago
Luna is good enough for me to give a parser spec and have it write one.
daveguy | an hour ago
vatsachak | an hour ago
mlmonkey | 3 hours ago
jerpint | 2 hours ago
From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
jhbadger | 2 hours ago
ianjbutler | 52 minutes ago
Since this kind of thing is always down to harness, project-type, and other structural constraints, of course your mileage may vary. Fable is probably great for pen-testing, or as a decision-making kernel of other kinds of applications, and way better than Opus at those things. Probably fine for code-review or changing a codebase of a few thousand lines in any language. Actually building that codebase or changing an even bigger one? Woof.
How this fits in with science? IDK but I bet other existing causal reasoning benchmarks might tell the whole story and this is back to stability again. Sometimes having a smart idea is really important! But more often it's important to just not forget what you were doing. What was I talking about? Oh look a squirrel
johnnyApplePRNG | 2 hours ago
You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.
Don't get me wrong, codex is great at finding bugs and building games. It's great.
a2ff6eeb0 | 2 hours ago
boorang | an hour ago
respectattentio | 58 minutes ago
AI should have started with science from the beginning, not after 4 years.
I am building on top of it with agents to improve scientific workflows.