My team is quietly sliding into doing more and more programming work through Claude Code, while still mandating human reviews. I'm really struggling with this development.
When I review code written by a human, I do it via a sort of "technical empathy." If I had been given this task, how would I have approached it? What knowledge would I have needed? What mistakes would I have tried to avoid? I attempt to put myself in the other developer's shoes and understand their reasoning. Which is possible because we are both humans. When reviewing LLM-generated code, this method falls apart. There's no human reasoning to relate to. How the code was put together follows a totally different, inscrutable logic. Potential mistakes happened for other reasons and take different forms.
Even more infuriating, if I happen to find a mistake, there is no discussion as to where the mistake came from, what misunderstanding of the original problem or lack of technical knowledge led to it. There is an answer akin to "You're right. I made a mistake here -- let me rewrite that properly." There is no human exchange, understanding of each other's point of view, collaboration towards an acceptable solution, mutual growth. Just a cold, empty interaction with a black box.
There's something utterly disgusting in becoming a sort of biological appendage to systems that become more prevalent and more opaque with each passing day, and I have a hard time articulating exactly why.
There's something utterly disgusting in becoming a sort of biological appendage to systems that become more prevalent and more opaque with each passing day, and I have a hard time articulating exactly why.
Cory Doctorow has a term for this: reverse centaur. He's alarmist sometimes but he coins great terms.
Something to think about for that conversation: IMO the conversation we should be having isn't "human coding VS agentic coding". Instead we should be talking about how to incorporate agents in ways that make sense and don't suck for humans. There's no denying that, used well, agents can speed things up and soften pain points without lowering quality (too much) or driving people insane. The problem is that there's not yet conventional wisdom about what that should look like. We're still figuring out what coding agents are best used for, and what they aren't good at. And then re-figuring it out every model generation or two.
I think the answer for the time being is to be extremely intentional when incorporating agents. And expect a lot of experience and feedback driven iteration of the process. Scaffolding makes a huge difference and I've seen a lot of attempts that miss this entirely. Formalization is another big one but I won't digress. Most importantly: until things have settled and best practices start to emerge, every org should take the time to carefully create processes that work for them. Preferrably driven largely by senior engineers rather than management.
It sounds like maybe the process isn't particularly intentional at your company, so humans are left trying to adapt, rather than adapting agents to what works for humans. It doesn't have to be that way. Though of course I know nothing about your org, so I might have that wrong. And maybe you're at a scale where intentionally building processes is really hard. The principle is sound though.
An example of this problem I think is the devaluation of covering letters or recommendation letters from teachers. The value idea to be that if someone took time to write one of those, it was because they meant it. If they write themselves, they are less valuable.
The examples you gave are things that I have a hard relationship with. They are both things that I think AI has allowed me to do better. They are also both things that I assign no value other than the vestigial relationship we have with them and job searching. Every person I knew who wrote recommendation letters 10+ years ago just had a form letter they filled in and most job applicants had the same with a cover letter. Perhaps at the higher levels of a career these things become more important but even the HR reps I know didn't really value a cover letter. We're probably entering an end-phase of the war of job applications where HR has been using AI to sort applicants and applicants are using AI to try and get selected. Cynically, it is highlighting that the system was always kind of dumb.
I feel the same about references for most jobs. You should be able to find three people to lie for you. But that starts getting more off-topic.
And I really really dislike AI writing when I encounter it outside of documentation.
I really dislike AI writing of documentation.
Instead of one short page with clear explanations I got 3-7 (sometime 20) pages of dilutet text with repeated statements thats honestly hard to digest.
But of course all people that do not need to work with resulting documents are happy. You got proper big beautiful documents in just few minutes.
Which is to say that LLM prose folds in signals of authority, competence, rigor and insight without any of those things needing to be present in the content. A fluent human is often worth listening to while a fluent LLM is just the default, the fluency carries no information. I think that's part of the reason for AI psychosis, I think it effectively uses, now outdated, expectations to hijack perception.
I disagree with this line here. I'll give some counterexamples:
(1) I work with quite a few people who have English as a second language. AI has been really helpful for them in composing and managing emails - it reduces the cognitive load of having to write in English and they can express themselves easier and faster. The combination of their thought process and the AI composition lets them be more fluent in email then they could be otherwise.
(2) I have about 20 years of helping small nonprofits with finance and accounting issues. I needed to draft a policy for one I'm working with at the moment. I used Claude to draft the policy - and it did a pretty good job - it essentially hit the best practices. I needed to customize and redraft a few items to meet the organizations specific context, but overall, it was 85% complete. It saved me quite a bit of time. It was also better produced then about 75% of the nonprofit financial policies I've read in my lifetime. The raw AI output wasn't as good as if I had hand drafted it, but most nonprofits don't have someone like me in their corner. The fluency of the AI's policy is better then what is available to most nonprofits.
I used Claude to draft the policy - and it did a pretty good job - it essentially hit the best practices. [...] overall, it was 85% complete [...] It was also better produced then about 75% of the [...] policies I've read in my lifetime.
I've been thinking for a while now that Big Consulting is pretty much cooked as soon as their "illusion" fades away. (Big 4, WITCH, Deloitte, McKinsey, Accenture, Capgemini, Oliver Wyman...)
Yes - I broadly agree. A lot of the strategy work was already under pressure since so many of the strategy consulting tools and knowledge are well known. AI probably accelerates that except for the truly novel problems.
IBM's latest returns showed that the business process outsourcing work is really under pressure right now. AI agents are going to eat that portion of their business and automate what was probably automatable.
The ERP systems are probably quite busy with AI projects and attaching AI capabilities and use cases to the core systems.
I work in consultancy, though not as a prototypical consultant more of a secondment type of deal just doing your regular dev stuff as part of a client team. But, other departments of my company certainly fit the bill of what people think of. Anyway, reading OPs post I couldn't help but think that part of what they encounter is something I have been dealing with for years now from your stereotypical consultants.
Specifically this bit
Which is to say that LLM prose folds in signals of authority, competence, rigor and insight without any of those things needing to be present in the content. A fluent human is often worth listening to while a fluent LLM is just the default, the fluency carries no information.
Did actually prompt a jaded chuckle on my part as many consultants and consultancy firm have been doing exactly this for decades already. And historically people have been really bad at picking it up. Frankly speaking, depending on the firm (or department in the company I work for) I'd be more inclined to trust an LLM over many subjects as they, in theory, are at least building on a wide array of training data.
I've been thinking for a while now that Big Consulting is pretty much cooked as soon as their "illusion" fades away. (Big 4, WITCH, Deloitte, McKinsey, Accenture, Capgemini, Oliver Wyman...)
I am mostly seeing a pivot towards AI (as everywhere) with the angle being that they are of course the best bet in properly implementing AI solutions with all the safeguards, etc, etc. But yeah, I have heard rumblings that certain areas of these companies are struggling more already.
I've been thinking for a while now that Big Consulting is pretty much cooked as soon as their "illusion" fades away.
Ehhh... I don't think I agree. If the premise was that what consulting firms sell is information and experience, then yeah, I'd 1000% agree with you. They're finished. Almost all of what they do is draw upon a huge internal library of pretty standard info, customize it a bit, and deliver it as a product to a customer. The bigger customers get a bit of actual analysis and teams of people to do more in depth customization along with more white glove treatment, but it's mostly the same thing.
That's not actually the value proposition of management consulting firms though. If it was, they probably would have already been replaced by Google.
What they're really selling is authority. Its a arrow in an executive's quiver. If an executive wants to do something big, they need to convince someone of it. If it's a director, they need to convince the COO and CFO. If it's another C level, they need to convince the CEO. If it's the CEO, they need to convince the board.
They can't do that effectively without hard data and numbers. They go to a consulting firm because they almost always already have a conclusion in mind. The firms job is basically to justify that conclusion. If they can't justify it, the executive either gets a second opinion of someone who does, just does it live and conveniently leaves out the consultants report, or if they're lawful good, just drops the idea and comes up with another idea that doesn't suck.
The value that the firm brings isn't the actual info though, it's being able to go to the board saying "McKinsey said this was a good idea". That's why the big three are the big three. Everyone knows their name, and most people trust them. That name is their entire value proposition, and is the reason why they get six figures to put together some boilerplate PDFs.
Even if an AI tool gave you the exact same content, it would be worthless, because a CEO going to a board with an idea and defending it with "chatGPT said this was a good idea" would be laughed out of the room.
I don't see that changing any time soon (unfortunately)
Even though in modern tools, translation and text generation follow the exact same pipeline, to me, they're still fundementally different tasks that, to me, even LLMs handle fundementally differently.
For text generation, the goal is to replicate a convincing approximation of what a human being would write, given the prompt. For translation, the goal is to, as accurately as possible, move the exact ideas given from one language to another.
One of them is basically an exercise in trying to fool humans into thinking the writer is also human, the other one is an exercise in trying to fool the speakers of one language that a speaker of a different language speaks the same language as them.
The first one is much more blatant of a lie than the second, and requires way more outright fabricated bullshit. The second is being guided by the exact words and meaning behind them being given to them.
I don't know if LLMs go through different training pipelines between these tasks, but when I read AI translated text versus AI generated text, the former feels far more natural and human, and doesnt have that innate immediate offensive feeling that the latter does.
Most of the workflows that I've seen are less that they write in their native tongue and then translate, it's more that they have five bullet points and ask the AI to create paragraphs, or they write the full email and then ask the AI to clean up their grammar.
I guess I find it odd that anyone gets offended about ai generated text. I've read so much human generated crap that I don't get particularly excited about machine generated. Maybe my set point is higher than most people's or my one emotion is already permanently bruised.
If all they have to say can be summed up in five bullet points, I would much rather read an email with five bullet points than five paragraphs of nothing. There can be a lot of information lost going from bullet points -> AI paragraphs -> my own interpretation of those 5 points. If you just present the five points up front, there's a lot less going on, and less room for me to misinterpret your (literal) points.
Every time something is translated or rewritten (whether it be by a human or a tool), some information is lost or altered, even without accounting for the possibility that mistakes . Sometimes, it's acceptable or even good (e.g. changing "In my personal opinion, I think that it could be said that blah" to "I think blah"), and sometimes it's a very bad change. It's like playing a game of telephone, or using Google translate to go between languages a few times before you end up with unintelligible garbage, except with AI-assisted tools, it still seems coherent, so it's harder to tell that something's wrong.
For example, I ran the sentence "This sentence could not possibly be unfalse" through Hypertranslate, and the result is the complete opposite "This sentence cannot be false." Now, you may say that "unfalse" isn't a word, and it got translated to "false" along the way. And that's probably correct. But if a non-native speaker is trying to say something, and uses a word that isn't quite right, most English speakers can recognize what they mean, even if the word is wrong. Computers don't always have that nuance.
TL;DR:
If the message is just five key points, just use the key points
Every rewrite introduces the possibility of losing or changing meaning
AI-generated text can hide mistakes because it still sounds natural and coherent
Humans can usually infer from incorrect word usage but computers might not be able to
Why waste time say lot word when few word do trick?
Using five bullet points might be your (and honestly my) communication preference, but it's certainly not a universal norm. I've worked for bosses that prefer bullets and others that prefer paragraphs. Some corporate cultures - Amazon for instance - are notoriously text heavy. Others work via face to face discussion. Most people don't get to dictate those preferences all the time and have to adapt to the people around them.
Fair enough. I probably wouldn't work well in an environment where a boss wanted to dictate the style of internal informational emails, but I can fully accept that those bosses exist. Text-heavy emails just feels like another bullshit metric that some high-level exec decided was a good proxy for productivity, even though it's the opposite. People waste time writing lots of unnecessary words, and then (presumably) people waste time reading those words.
In places that do everything face to face, it's always good to send a follow up email summarizing the key points so that (1) there's a paper trail, and (2) if something was misinterpreted, there's an opportunity to correct it.
I don't know if LLMs go through different training pipelines between these tasks,
Likely there are different RF segments for these but the important thing is the prompt/context. "Translate this" results in inference that's going to try harder to stay close to the source material. "Write about this" is asking for fabrication, which is a free pass for all the tropes introduced in fine tuning to come out.
I think its similar to how many AI slop software we have. I already see AI writing everywhere, it's easy to recognize once you catch on the few phrases it often repeats.
I'm less worried about what I work with but more worried with society as a whole. A lot of non dev folks already have a hard time recognizing what is an AI slop website and what is a AI-assisted website. I'm honestly not sure what the solution is except for educating people and trying to improve the aggregate quality bar for both writing and websites.
kahuna_dumpster | 11 hours ago
My team is quietly sliding into doing more and more programming work through Claude Code, while still mandating human reviews. I'm really struggling with this development.
When I review code written by a human, I do it via a sort of "technical empathy." If I had been given this task, how would I have approached it? What knowledge would I have needed? What mistakes would I have tried to avoid? I attempt to put myself in the other developer's shoes and understand their reasoning. Which is possible because we are both humans. When reviewing LLM-generated code, this method falls apart. There's no human reasoning to relate to. How the code was put together follows a totally different, inscrutable logic. Potential mistakes happened for other reasons and take different forms.
Even more infuriating, if I happen to find a mistake, there is no discussion as to where the mistake came from, what misunderstanding of the original problem or lack of technical knowledge led to it. There is an answer akin to "You're right. I made a mistake here -- let me rewrite that properly." There is no human exchange, understanding of each other's point of view, collaboration towards an acceptable solution, mutual growth. Just a cold, empty interaction with a black box.
There's something utterly disgusting in becoming a sort of biological appendage to systems that become more prevalent and more opaque with each passing day, and I have a hard time articulating exactly why.
[OP] post_below | 10 hours ago
Cory Doctorow has a term for this: reverse centaur. He's alarmist sometimes but he coins great terms.
kahuna_dumpster | 9 hours ago
Thanks, I didn't know about that concept. Gives me more fuel for the upcoming discussion I'm gonna have to have with my colleagues.
[OP] post_below | an hour ago
Something to think about for that conversation: IMO the conversation we should be having isn't "human coding VS agentic coding". Instead we should be talking about how to incorporate agents in ways that make sense and don't suck for humans. There's no denying that, used well, agents can speed things up and soften pain points without lowering quality (too much) or driving people insane. The problem is that there's not yet conventional wisdom about what that should look like. We're still figuring out what coding agents are best used for, and what they aren't good at. And then re-figuring it out every model generation or two.
I think the answer for the time being is to be extremely intentional when incorporating agents. And expect a lot of experience and feedback driven iteration of the process. Scaffolding makes a huge difference and I've seen a lot of attempts that miss this entirely. Formalization is another big one but I won't digress. Most importantly: until things have settled and best practices start to emerge, every org should take the time to carefully create processes that work for them. Preferrably driven largely by senior engineers rather than management.
It sounds like maybe the process isn't particularly intentional at your company, so humans are left trying to adapt, rather than adapting agents to what works for humans. It doesn't have to be that way. Though of course I know nothing about your org, so I might have that wrong. And maybe you're at a scale where intentionally building processes is really hard. The principle is sound though.
Vito | 12 hours ago
An example of this problem I think is the devaluation of covering letters or recommendation letters from teachers. The value idea to be that if someone took time to write one of those, it was because they meant it. If they write themselves, they are less valuable.
Requirement | 2 hours ago
The examples you gave are things that I have a hard relationship with. They are both things that I think AI has allowed me to do better. They are also both things that I assign no value other than the vestigial relationship we have with them and job searching. Every person I knew who wrote recommendation letters 10+ years ago just had a form letter they filled in and most job applicants had the same with a cover letter. Perhaps at the higher levels of a career these things become more important but even the HR reps I know didn't really value a cover letter. We're probably entering an end-phase of the war of job applications where HR has been using AI to sort applicants and applicants are using AI to try and get selected. Cynically, it is highlighting that the system was always kind of dumb.
I feel the same about references for most jobs. You should be able to find three people to lie for you. But that starts getting more off-topic.
Deely | 7 hours ago
I really dislike AI writing of documentation.
Instead of one short page with clear explanations I got 3-7 (sometime 20) pages of dilutet text with repeated statements thats honestly hard to digest.
But of course all people that do not need to work with resulting documents are happy. You got proper big beautiful documents in just few minutes.
D_E_Solomon | 9 hours ago
I disagree with this line here. I'll give some counterexamples:
(1) I work with quite a few people who have English as a second language. AI has been really helpful for them in composing and managing emails - it reduces the cognitive load of having to write in English and they can express themselves easier and faster. The combination of their thought process and the AI composition lets them be more fluent in email then they could be otherwise.
(2) I have about 20 years of helping small nonprofits with finance and accounting issues. I needed to draft a policy for one I'm working with at the moment. I used Claude to draft the policy - and it did a pretty good job - it essentially hit the best practices. I needed to customize and redraft a few items to meet the organizations specific context, but overall, it was 85% complete. It saved me quite a bit of time. It was also better produced then about 75% of the nonprofit financial policies I've read in my lifetime. The raw AI output wasn't as good as if I had hand drafted it, but most nonprofits don't have someone like me in their corner. The fluency of the AI's policy is better then what is available to most nonprofits.
TaylorSwiftsPickles | 8 hours ago
I've been thinking for a while now that Big Consulting is pretty much cooked as soon as their "illusion" fades away. (Big 4, WITCH, Deloitte, McKinsey, Accenture, Capgemini, Oliver Wyman...)
D_E_Solomon | 8 hours ago
Yes - I broadly agree. A lot of the strategy work was already under pressure since so many of the strategy consulting tools and knowledge are well known. AI probably accelerates that except for the truly novel problems.
IBM's latest returns showed that the business process outsourcing work is really under pressure right now. AI agents are going to eat that portion of their business and automate what was probably automatable.
The ERP systems are probably quite busy with AI projects and attaching AI capabilities and use cases to the core systems.
creesch | 7 hours ago
I work in consultancy, though not as a prototypical consultant more of a secondment type of deal just doing your regular dev stuff as part of a client team. But, other departments of my company certainly fit the bill of what people think of. Anyway, reading OPs post I couldn't help but think that part of what they encounter is something I have been dealing with for years now from your stereotypical consultants.
Specifically this bit
Did actually prompt a jaded chuckle on my part as many consultants and consultancy firm have been doing exactly this for decades already. And historically people have been really bad at picking it up. Frankly speaking, depending on the firm (or department in the company I work for) I'd be more inclined to trust an LLM over many subjects as they, in theory, are at least building on a wide array of training data.
I am mostly seeing a pivot towards AI (as everywhere) with the angle being that they are of course the best bet in properly implementing AI solutions with all the safeguards, etc, etc. But yeah, I have heard rumblings that certain areas of these companies are struggling more already.
papasquat | 6 hours ago
Ehhh... I don't think I agree. If the premise was that what consulting firms sell is information and experience, then yeah, I'd 1000% agree with you. They're finished. Almost all of what they do is draw upon a huge internal library of pretty standard info, customize it a bit, and deliver it as a product to a customer. The bigger customers get a bit of actual analysis and teams of people to do more in depth customization along with more white glove treatment, but it's mostly the same thing.
That's not actually the value proposition of management consulting firms though. If it was, they probably would have already been replaced by Google.
What they're really selling is authority. Its a arrow in an executive's quiver. If an executive wants to do something big, they need to convince someone of it. If it's a director, they need to convince the COO and CFO. If it's another C level, they need to convince the CEO. If it's the CEO, they need to convince the board.
They can't do that effectively without hard data and numbers. They go to a consulting firm because they almost always already have a conclusion in mind. The firms job is basically to justify that conclusion. If they can't justify it, the executive either gets a second opinion of someone who does, just does it live and conveniently leaves out the consultants report, or if they're lawful good, just drops the idea and comes up with another idea that doesn't suck.
The value that the firm brings isn't the actual info though, it's being able to go to the board saying "McKinsey said this was a good idea". That's why the big three are the big three. Everyone knows their name, and most people trust them. That name is their entire value proposition, and is the reason why they get six figures to put together some boilerplate PDFs.
Even if an AI tool gave you the exact same content, it would be worthless, because a CEO going to a board with an idea and defending it with "chatGPT said this was a good idea" would be laughed out of the room.
I don't see that changing any time soon (unfortunately)
papasquat | 6 hours ago
Even though in modern tools, translation and text generation follow the exact same pipeline, to me, they're still fundementally different tasks that, to me, even LLMs handle fundementally differently.
For text generation, the goal is to replicate a convincing approximation of what a human being would write, given the prompt. For translation, the goal is to, as accurately as possible, move the exact ideas given from one language to another.
One of them is basically an exercise in trying to fool humans into thinking the writer is also human, the other one is an exercise in trying to fool the speakers of one language that a speaker of a different language speaks the same language as them.
The first one is much more blatant of a lie than the second, and requires way more outright fabricated bullshit. The second is being guided by the exact words and meaning behind them being given to them.
I don't know if LLMs go through different training pipelines between these tasks, but when I read AI translated text versus AI generated text, the former feels far more natural and human, and doesnt have that innate immediate offensive feeling that the latter does.
D_E_Solomon | 6 hours ago
Most of the workflows that I've seen are less that they write in their native tongue and then translate, it's more that they have five bullet points and ask the AI to create paragraphs, or they write the full email and then ask the AI to clean up their grammar.
I guess I find it odd that anyone gets offended about ai generated text. I've read so much human generated crap that I don't get particularly excited about machine generated. Maybe my set point is higher than most people's or my one emotion is already permanently bruised.
thecakeisalime | 5 hours ago
If all they have to say can be summed up in five bullet points, I would much rather read an email with five bullet points than five paragraphs of nothing. There can be a lot of information lost going from bullet points -> AI paragraphs -> my own interpretation of those 5 points. If you just present the five points up front, there's a lot less going on, and less room for me to misinterpret your (literal) points.
Every time something is translated or rewritten (whether it be by a human or a tool), some information is lost or altered, even without accounting for the possibility that mistakes . Sometimes, it's acceptable or even good (e.g. changing "In my personal opinion, I think that it could be said that blah" to "I think blah"), and sometimes it's a very bad change. It's like playing a game of telephone, or using Google translate to go between languages a few times before you end up with unintelligible garbage, except with AI-assisted tools, it still seems coherent, so it's harder to tell that something's wrong.
For example, I ran the sentence "This sentence could not possibly be unfalse" through Hypertranslate, and the result is the complete opposite "This sentence cannot be false." Now, you may say that "unfalse" isn't a word, and it got translated to "false" along the way. And that's probably correct. But if a non-native speaker is trying to say something, and uses a word that isn't quite right, most English speakers can recognize what they mean, even if the word is wrong. Computers don't always have that nuance.
TL;DR:
D_E_Solomon | 5 hours ago
Using five bullet points might be your (and honestly my) communication preference, but it's certainly not a universal norm. I've worked for bosses that prefer bullets and others that prefer paragraphs. Some corporate cultures - Amazon for instance - are notoriously text heavy. Others work via face to face discussion. Most people don't get to dictate those preferences all the time and have to adapt to the people around them.
thecakeisalime | 3 hours ago
Fair enough. I probably wouldn't work well in an environment where a boss wanted to dictate the style of internal informational emails, but I can fully accept that those bosses exist. Text-heavy emails just feels like another bullshit metric that some high-level exec decided was a good proxy for productivity, even though it's the opposite. People waste time writing lots of unnecessary words, and then (presumably) people waste time reading those words.
In places that do everything face to face, it's always good to send a follow up email summarizing the key points so that (1) there's a paper trail, and (2) if something was misinterpreted, there's an opportunity to correct it.
[OP] post_below | 59 minutes ago
Likely there are different RF segments for these but the important thing is the prompt/context. "Translate this" results in inference that's going to try harder to stay close to the source material. "Write about this" is asking for fabrication, which is a free pass for all the tropes introduced in fine tuning to come out.
IndieGamesCafe | 6 hours ago
I think its similar to how many AI slop software we have. I already see AI writing everywhere, it's easy to recognize once you catch on the few phrases it often repeats.
I'm less worried about what I work with but more worried with society as a whole. A lot of non dev folks already have a hard time recognizing what is an AI slop website and what is a AI-assisted website. I'm honestly not sure what the solution is except for educating people and trying to improve the aggregate quality bar for both writing and websites.
h3x | 12 hours ago
/offtopic
Yes, it's catching on!
[OP] post_below | 10 hours ago
It's got an old school dark children's story vibe to it
Deely | 8 hours ago
I personally prefer ...
But, of course, the more, the merrier
TaylorSwiftsPickles | 6 hours ago
Not including a tilde in it is criminal. Obviously it's Tilderiños, mierda!