Wrong, not broken

4 points by sp6370 a day ago on lobsters | 15 comments

emk | a day ago

This week we launched Amazon CloudWatch Omni to observe agents, applications, and infrastructure together, and I wanted to give an inside look at the shift behind it and why it matters.

Hi, Claude. If Amazon can't be bothered to write their own release notes, I can't be bothered to read them.

(And, yes, Pangram, with its very low false positive rate agrees this is slop. But nobody who uses LLMs needs Pangram to know this, unfortunately. It's really amazing how jarringly easy this kind of slop becomes to notice after overexposure to LLMs. And it's sort of like wearing a clown nose in public.)

typesanitizer | a day ago

FWIW, I use LLMs frequently (but Claude much less often) and I did not realize that this was AI-generated on first reading. So thanks for doing this check and sharing.

janiczek | 23 hours ago

I got the ick right around:

Every dashboard is green. Nothing appears to be broken.

It’s wrong.

fragmentcity | 23 hours ago

Make a point of regularly doing focused reading of high-quality blogs and articles written before 2019. For that matter, the older, the better. Watch whatever YouTube says is peak C-SPAN. Find ways to immerse yourself in good speaking and writing that interests you. It's being produced now, too, but if you want to really understand the difference, you have to go pre-AI, pre-social media.

This is the sentence that indicated it to me:

Correctness now has to be measured at the level of the run.

dallen | a day ago

What a brave new world we live in. With no sound understanding of what a program is doing, you can't even properly assess if it's correct.

This creates a second problem: Quality has to become explicit enough to measure.

This is a pipe dream.

viraptor | a day ago

Yeah, I don't get why so many companies saw the agents and instead of "now we can easily create new, extended, strict support flows based on extra information" just went with "Yolo, let customers argue with the bot and hope for the best". How is the refund value randomly found in old documentation even a valid state?

thesnarky1 | 23 hours ago

Exactly, it IS broken if it can fetch the wrong document. The same as if your SQL statement would just guess which refund value to select from the database 5 years ago.

marti | a day ago

Back in the day, writing Unit test and integration tests paired with humans was enough acceptable correctness and you could get more/less correct by dialing it. Is it really worth the benefit to throwaway all of that for whatever this is? Who defines correctness? If it's still the human, why is this better than what we had?

fragmentcity | 23 hours ago

For Amazon publications to make no better an impression than the median slop post is pretty gobsmacking. Not that they were a pinnacle of the English language before, but whew.

tobin_baker | 17 hours ago

This is from a C-level executive so draw your own conclusions.

mxey | a day ago

Is this saying that before agents software couldn’t respond with HTTP 200 and invalid results? Weird phrasing.

Now that’s some crappy slop from a big company . Would you trust their AI solution sales pitches if they can’t use it superbly without looking dumb?

einacio | an hour ago

Dangerous topic to post from AWS. The last 4 companies I worked with did a move away from cloud to hosted because it was just a really bad fit (on latency, cost and options available)

Student | 20 hours ago

This is an argument for absolutely minimizing the use generative models on the hot path. They’re so slow that people should be doing that anyway.

tobin_baker | 17 hours ago

Comment removed by author