All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.
For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).
Better than Astra and Fable? It looks quite pretty and even impressive at times if you squint, but look closer and it falls apart in terms of coherency. And I say that as somebody who mains Gemini 3.8.
I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
Its giraffe / grandfather clock one is pretty bad... (two necks? wearing a suit?)
Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.
Anyone else surprised the generations look so remarkably similar? All of these models have “independently” generalized that the moose should roughly be standing at the same position (left) or that the giraffe should have a certain color palette.
With respect to bikes (with or without pelicans) there's a strong natural bias because people displaying bikes tend to want to show off the side with the gears.
More-generally, I suspect an influence from how left-to-right languages (i.e. English) affect comic layouts. Overcoming that bias often means using vertical space to exploit the top-to-bottom habit instead. (Consider the rarity of an English-language comic panel where action is from bottom-right to top-left.)
steinvakt2 | 4 hours ago
samayashar | 4 hours ago
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
CamperBob2 | 3 hours ago
Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.
honeycrispy | 2 hours ago
That's pretty generous.
vova_hn2 | 3 hours ago
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
[0] https://en.wikipedia.org/wiki/Goodhart%27s_law
GaggiX | 2 hours ago
For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).
sajithdilshan | 3 hours ago
Jordan-117 | 2 hours ago
sceptic123 | 3 hours ago
A monkey, surely?
qiine | 3 hours ago
ormax3 | 2 hours ago
input_sh | an hour ago
The only other one I've spotted is animated is Gemini 3.0's 2025 run of an elephant.
svcrunch | 3 hours ago
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
[1] https://dorrit.pairsys.ai/
eddytrex_ | 3 hours ago
neilellis | 2 hours ago
BrokenCogs | 2 hours ago
GaggiX | 2 hours ago
pixelesque | an hour ago
Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.
villish | 2 hours ago
Qwen3.8 is very clearly distilled from Claude models.
dustfinger | 2 hours ago
dcreater | 2 hours ago
zaphar | an hour ago
mock-possum | an hour ago
kennywinker | an hour ago
outlore | 55 minutes ago
Terr_ | 52 minutes ago
More-generally, I suspect an influence from how left-to-right languages (i.e. English) affect comic layouts. Overcoming that bias often means using vertical space to exploit the top-to-bottom habit instead. (Consider the rarity of an English-language comic panel where action is from bottom-right to top-left.)