If I am relying on the model to do the writing without any context or learning on how I want it to write then yes. However if I build skills that have learnt how to write in the way I want them to then I find they write very well, or at the least how I want them to as opposed to how they do natively.
Accurate short answers / text are always harder than long answers, for human or AI. I know several authors and editors who write a lot longer at first, then spend a multiple of the initial time compressing it via a back and forth process to something dense. Sort of like weaving the initial threads.
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
If I had more time I'd write a shorter letter, a la Pascal.
Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.
I’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during training. This is just my prior though.
that and always putting the “current year” at the end of the search term (so the results are more recent, I guess?), except that “current year” consistently ends up being 2-3 years ago since I guess that’s what’s in the training data (even on a harness that injects the current date)
This behavior is a learned adaptation. It's most likely a feature not a bug.
Based on personal usage, I think it reflects search engine functionality degradation. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.
The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
Claude is still not perfect at reading and interpreting noisy graphical data (imagine something like an EKG or chromosomal microarray plot). Still better than an average person but makes mistakes, not sure if this fits your description.
It's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up.
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis).
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.
I've had a lot of trouble when it comes to sorting out UIs. I've tried with an iOS game and also a TypeScript app with UI elements from libraries like ReactFlow. The usual models can sometimes fix or change things based on screenshots but more often than not they just don't "get it" (e.g. certain shapes on a plane are overlapping, which I don't want, the models can't fix what they can't "see").
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
I have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated?
Or maybe just a lack of “imagination”
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
They aren't funny. The jokes they come up with are extremely lame and the sort of thing you would expect a company HR manager to tweet.
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
Editing a document without mixing edit instructions into the final document. Claude and ChatGPT do this all the time: I tell them to change X in a planning document or email draft, and instead of just changing X they also frequently add the edit instruction to “change X” into the document itself. They seem unable to take a step back and look at the document without “becoming” the document somehow. I do believe that dedicated subagents for editing may fix this but I am not sure.
blinkbat | an hour ago
Oh, you said simple. Speaking like a human
maxsavin | an hour ago
TZubiri | an hour ago
flippy_flops | an hour ago
veganmosfet | an hour ago
kanzure | an hour ago
NoPicklez | an hour ago
TZubiri | an hour ago
humanrebar | an hour ago
honr | an hour ago
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
FriedFishes | 50 minutes ago
Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.
dorianpruski | an hour ago
ghostpepper | an hour ago
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
astro1234 | an hour ago
mthoms | an hour ago
jedbrooke | 50 minutes ago
areoform | 50 minutes ago
Based on personal usage, I think it reflects search engine functionality degradation. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.
The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
bpodgursky | an hour ago
SubiculumCode | an hour ago
respectattentio | an hour ago
shoopadoop | an hour ago
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
sandcat_ | an hour ago
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
skeptic_ai | an hour ago
senectus1 | an hour ago
tartoran | an hour ago
TiccyRobby | an hour ago
lrvick | an hour ago
dhruv3006 | an hour ago
sghiassy | an hour ago
More of an image model than a LLM model tho
spike021 | an hour ago
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
eli | an hour ago
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
newsomix9xl | an hour ago
newsomix9xl | an hour ago
rufi | an hour ago
elliotto | an hour ago
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
alexandra_au | 52 minutes ago
alexeldeib | 51 minutes ago
znnajdla | 51 minutes ago