The AI Pascal’s Wager

28 points by cgrinds 20 hours ago on lobsters | 43 comments

When you think about it, contributors who are completely turned away by the fact that they can’t use LLMs are probably contributors you don’t want in your project, even if you have a Debian-like “moderate use of AI” position.

If a project says, "No AI," then I absolutely respect that. Nothing which has been touched by AI in any way will be sent to that project. This is pretty clearly what a lot of people with that policy want. They ethically object to any use of LLMs. It's a matter of moral purity, not a practical question. I wouldn't offer them something that involved AI at any point any more than I would offer a vegan something that had touched butter.

Let me give you an example: I was working on a distributed word frequency estimator. Before I left for a meeting, I asked Claude to check the distributed counts and see if it could find anything interesting. Claude found some weirdness in the numbers, and—completely unasked—it started digging around in one of the open source libraries I was using. It found a nasty and very subtle bug in the open source library.

So now I have a problem: I know about a serious bug. But it was found entirely using AI. As lawyers like to say, it was "fruit of the poisoned tree." There is no way that I can retroactively remove AI from the process.

Now, that author had no policy against AI contributions. So I hand-wrote a nice bug report and offered to fix it. Fascinatingly, Claude made a subtle mistake in its proposed fix to the algorithm. But both I and the upstream author made exactly the same mistake Claude did the first times we tried to fix it. The only thing that actually fixed the bug was when I (or maybe Claude at my request?) wrote some proptests, which found a super tricky edge case. (I actually think the incomplete fix is present in some academic papers about the algorithm, IIRC. I'm still not even 100% sure that the final fix was correct, but it holds up under a quarter million randomly generated test cases checking various mathematical properties, which is a good start.)

So that's my dilemma: Sometimes a coding agent will find bugs in someone else's code. These may be subtle correctness bugs, or data corruption bugs, or serious security vulnerabilities. And if I discover one of these bugs, there's a decent chance AI will be involved in some way. This is because my first reflex when someone else's library does something weird is to ask the nearest coding agent to take a look. Life's too short to debug things like treesitter myself when an agent will almost always find the problem in 2-5 minutes. I may have to fix the problem myself, and of course I'll write my own commit messages and PRs.

But if you have a strict zero AI policy, then I'm going to assume that you're doing it on moral grounds. Which puts me in the awkward position of knowing that your code is computing incorrect answers, or corrupting user data, or leaving your users vulnerable to attack. And then I have to guess whether you want to know about a bug discovered using AI. If you have a strict no-AI policy, I'm not going to hide the AI use.

In practice, if it's a security bug, I'll probably verify the bug carefully by hand and disclose it to you via a private channel, if I can find one. For other bugs, I may switch to another library, or just quietly fork yours. Which feels antisocial. But given how many "no AI" policies feel like ritual purity laws (nothing wrong with ritual purity laws!), I don't want to inflict AI-discovered bugs on people who are strongly anti-AI.

On my personal projects, I actually have a mix of AI policies. I will normally close slop feature PRs without reading. Slop bug-fix PRs, well, I'll either read them myself, or if the PR is too annoying, and confused, I'll ask my own agent whether the "bug" is real. A few of my projects are "learning" projects, and these will forbid all AI-written code, because the point is learning. Or sometimes my policy is "trusted maintainers can use AI if they want, but we'll reject slop feature PRs from other people."

simonw | 8 hours ago

I have this exact problem at the moment. There's a bug in click-default-group which I opened 3 years ago and then failed to resolve without AI.

The other day I pointed Claude at it and got a simple one-line fix. Then I read the governing AI policy which made it clear that, with no exceptions, my solution that had been found with the help of AI was not welcome:

If your contribution looks like it was generated by such tools, or you attempt to obfuscate that, it will be closed and you are likely to be blocked.

I'm not going to obfuscate it - that would be unethical.

It's a one-line fix. I've now seen it, so my knowledge of the fix is polluted.

So I guess I leave that issue open and hope someone else stumbles across the same fix that I did!

Which puts me in the awkward position of knowing that your code is computing incorrect answers, or corrupting user data, or leaving your users vulnerable to attack. And then I have to guess whether you want to know about a bug discovered using AI. If you have a strict no-AI policy, I'm not going to hide the AI use.

So I guess I leave that issue open and hope someone else stumbles across the same fix that I did!

There is something quite odd about these comments.

As I understand it, the policies of an open source project do not prohibit individuals from publicising independent work. I suspect, irrespective of the policy of the open source project, if one were to find some mistake, one could fix it in a fork, one could write a blog post about it, one could comment on it (in a forum provided by the project… or even just as corroborating detail when making factual claims in other contexts on entirely other websites…) Surely there's no need to be so coy.

simonw | 7 hours ago

I filed an issue three years ago explaining the bug clearly enough that a human, or LLM, could find a solution.

I've now found the solution using an LLM, but the solution needs to be found by a human.

singpolyma | 7 hours ago

Yes. I think publishing a personal fork is a reasonable choice

simonw | 7 hours ago

If I publish a personal fork with my one-line fix am I now making it harder for that fix to make it into the upstream project? I don't know if there are alternative fixes, and my fix was AI-assisted.

I think everyone is getting extremely legalistic about other projects' AI policies on their behalf. No one wants the actual output of the AI, especially unchecked. But if you read the output and actually understand it - a pretty minimal bar for a contribution - and then fix the thing yourself, it has none of the associated issues with AI PRs.

And if there's only one possible fix, a one liner, then that's by definition the fix you'd write yourself once you understand the issue. Write it yourself and submit it. There's no "fruit of the poisonous tree," this isn't a courtroom.

singpolyma | 5 hours ago

For a one line easy fix I sort of agree, but I don't think this is true in general. I don't think if I write a patch with LLM help. Then put it away and recreate the exact same patch this time "without help" that I'm following the spirit of a ban.

For many maintainers I've talked to the point is to ban anyone who might be using LLM related tools from being a contributor.

FeepingCreature | 5 hours ago

I think everyone is getting extremely legalistic about other projects' AI policies on their behalf. No one wants the actual output of the AI, especially unchecked. But if you read the output and actually understand it - a pretty minimal bar for a contribution - and then fix the thing yourself, it has none of the associated issues with AI PRs.

Well great, then they can write that in their policy instead of "absolutely no AI"!

To avoid arguing with strawmen let's stick with the actual example Simon provided. The word "absolutely" doesn't appear anywhere. Here are some quotes:

Pallets is affected by a large increase in low quality contributions generated with LLM (large language model) or "AI" (artificial intelligence) tools.

There's the motivating problem behind the ban. It's certainly one we can all understand, right?

We strongly encourage everyone to participate and contribute without using LLM or AI tools.

That's a "strongly encourage," not an "absolutely no," which seems intentionally weakened to provide a tiny bit of space for exercising judgement.

This policy applies to all interactions with our projects and community, not only code and pull requests. This includes code, documentation, tests, PR descriptions, comments, reviews, answers, and all other interactions.

"Interactions with our projects and community" - note that they are specifying interactions between people here. You shouldn't subject other people to AI output under this policy. It says nothing about using them to help you understand something. The requirement to understand is still there, as it should be.

When you choose to use an LLM tool to generate output, you look like every other person who did so, rather than a unique contributor. Therefore, you inherit the distrust maintainers have based on every prior low quality interaction that looked exactly like yours.

So actually the vast majority of the policy is explaining the harm caused by generated output. We've all seen large chunks of generated output and know exactly what they mean.

In most cases a lot of work goes into crafting these policies, which are extremely controversial right now, and they are very descriptive, tell you what problems they're solving, and are nuanced. That effort seems wasted when people argue with a caricature of them instead of what they actually say.

FeepingCreature | 4 hours ago

I mean, tbh I wouldn't contribute to this project because this is a weird policy in the sense that it only actually mandates two things:

  • Do not use LLMs to translate or improve communication
  • "Carefully consider" ESQ issues before using them.

Does that mean I can make LLM PRs? Well they're telling me I am "likely to be blocked"... is that a no? Will they say, "fair cop, it's a good LLM PR but we'll block you now"? I think this policy is written so carefully and defensively that it doesn't actually tell the contributor what to do and what not to do. It's hard to have an opinion on what it's saying because it's not actually saying much.

They were IMO super clear about this exact point too:

If you are a new contributor, do not use LLM or AI tools to contribute at all; you do not have the trust or experience to take responsibility for what is generated.

So no, you can't. And this is a completely pragmatic stance for them that balances the harm they're receiving with the risk of not getting some potential quality contributions.

FeepingCreature | 59 minutes ago

That's the sort of sentence I'd have loved to see in the summary, not the explanation.

this isn't a courtroom

It is not.

In fact, perhaps one might even go so far as to suggest that where an open source project's policy disagrees with an individual's preferences, the necessary reaction does not seem to be wounded antagonism. Surely when the project authors set policy, they are aware that they may exclude some desirable contributions; that project may be disinterested in certain contributions, does not require one to make a show of withholding (completed) work under the pretense of following established policy.

The project is, as we all agree, free to set policy as it sees fit.

Irrespective of said policy, nothing enjoins an individual from, say, posting publicly—“here's a mistake I found”; “here's a one line-fix I found.” (In fact, as I recall, websites hosting independently maintained patches were quite common prior to wide-spread use and centralisation around version control tooling…)

In fact, I would suggest that if one were to disagree with a project's policies, and if one were to firmly believe that these policies were detrimental, then maintaining a site with all of the excluded contributions would be very strong corroborating evidence. I use an LLM frequently in my own work and studies, so I know that generating a Github-hosted <html><body><ul><li><a href="patches/some-fix.diff">Fix for … issue in some-project</a></li></ul></body></html> is well within their capabilities… (In fact, if you write it by hand, you can even skip some of the closing tags!)

Altogether, I think the original comments we are replying to are still quite odd.

reidrac | 4 hours ago

I guess it depends on the one-line fix and the project.

Publish it and make the project aware. Let them decide. If they reject it, you have the fix. Move on.

No offence but the way you're dealing with this doesn't feel completely honest.

tentacloids | 6 hours ago

with no exceptions, my solution that had been found with the help of AI was not welcome

I don't actually think that's the case. This policy focuses on generated contributions more than anything else. If you read a generated description of the bug and its fix, then implement it and write the PR yourself, that isn't explicitly banned. I don't think that's a shortcoming of the policy considering what I think is their primary focus:

We value contributions that are genuinely understood, created, and communicated by humans.

The overall message I get is: don't use LLMs to substitute your own personal development (and definitely don't use them to substitute actual communication).

I think the end of this paragraph is most relevant to this situation (emphasis mine):

If you are a new contributor, do not use LLM or AI tools to contribute at all; you do not have the trust or experience to take responsibility for what is generated. If you are someone who is familiar to the maintainers, and know that they already trust you to make high quality contributions, your contribution must still genuinely be understood and created personally by you, regardless of what tools you may use to assist you.

Comment removed by author

munificent | 3 hours ago

Your comment is a good one and I think does a good job of highlighting that current AI policies are drawing the boundary in a weird place.

I think the distinction many teams actually want to make is: was every line in this codebase read, understood, and accepted by a human or not?

There are valid arguments in both directions:

  • Codebases where the answer is "yes" mean that there is a human who is claiming some level of accountability and responsibility for the code. Not in a legal sense for most open source licenses, of course. But in a "can I trust that a person gave a shit about this?" sense. Humans bring a whole cloud of moral and social values with them and when a human accepts some code, it means we can more likely trust it in the large diffuse sense of the term "trust".

    It is more likely that the codebase will continue to be maintained in the future because people have already invested real effort into it.

  • Codebases where the answer is "no" can grow and add features much more rapidly. I don't think anyone knows what the long-term maintainability of vibe-coded programs is yet, but at least short-term it seems like AI can whip up a whole lot of mostly working code really fast.

So I think really where the policy line should be drawn is at code review. Does a human review every line of code or not?

The challenge is that if you have a policy that says "yes" but don't have a policy discouraging using AI to generate code, you will rapidly overload the reviewers. So the policy line gets drawn earlier at the point of authoring to try to avoid that. But I think then it runs into all of the problems that the parent comment here are talking about where it's not clear if a contribution or bug report is "tainted" by AI or not.

I suspect there is another approach one could take here that is a little more mechanical about the code. Maybe have a rule that says all PRs above a certain line count are auto-rejected unless a previous design discussion with maintainers has occurred.

Of course, if you object to the use of AI on moral/ethical grounds, I also think it is valid to have a policy that outright rejects all contributions tainted with AI in any way. But I think that's a somewhat different problem being solved than AI policies that are mostly trying to deal with not overloading human reviewers.

As someone who is very skeptical of “snake oil AI” but is nevertheless amazed by LLMs a s a technology, I think the discourse and the hype around it killed any possibility of a nuanced debate. The idea is not to filter out LLM-assisted contributions themselves, but contributors who believe LLMs are good at generating good code (or those who believe they can tame them with their awesome skills).

LLMs as of today are a Dunning-Kruger effect amplifier and I’m afraid this is not something that can be easily fixed. Also, there is more than enough evidence that cognitive bandwidth is a major factor when it comes to identifying good contributions as well as contributors (I can’t help but think of the XZ Utils supply chain takeover that led to a backdoor injection). If I were an open source maintainer, I’d be very wary of a sudden disproportionate surge of experts interested in my project.

Sanity | 10 hours ago

Yeah. I wish for a world in which the dust has settled. Imagine: After the crash, OpenAI and Anthropic are more or less bankrupt and aren't getting more VC injections, so hardware companies have started focusing on consumers again, but Azure etc. are still serving models and there are still loads of local ones and because there is consumer hardware, they now seem like a viable option for normal people. We're just past the trough of disillusionment, so companies have stopped making AI toilets, datacenter buildout has slowed to a managable pace, therapists have learnt how to effectively deal with AI psychosis (ok that one's a pipe dream) and a million consultancies focused on cleaning up startup-slop have formed. For bonus points: Image/video generators have been globally banned and their use leads to harsh punishments, while major companies serving LLMs as a service need to pay some % to authors.

k749gtnc9l3w | 8 hours ago

have been globally banned

We can't do it with simpler and more tangible things!

And honestly, upcoming EU solution «anything generated must not be distributed unmarked», if it grew a tax scheme on top, would sound like enough.

marcecoll | 10 hours ago

This article made a decision at the start and tried to hide it in a list of pros and cons sadly.

When you think about it, contributors who are completely turned away by the fact that they can’t use LLMs are probably contributors you don’t want in your project

When this is the premise, what did you expect to find with the exercise? I have pretty severe unmedicated ADHD, onboarding on a project is an incredibly hard task for me. Agents have vastly improved this for me by being able to ask questions about a codebase as a starting point and poke and prod at the system through natural language. It allows me to cross validate style, abstractions through the codebase and explore possible solutions that fit the codebase a lot quicker. I'm not going to vibe code a solution to this unless its for a private fork or something but I am vastly faster in understanding things because I can zoom in and out in a way I couldn't before.

increasingly dependent of on expensive and completely proprietary software over which you have no control

I don't think this follows for two reasons. The initial moat worry of 15 monts ago is mostly all gone. Anthropic and OpenAI pushed the envelope and are now being chased closely by open weights labs. I've done several projects with deepseek v4 family of models, an open weights model that is at most 2-3 months behind the sota and I think I've topped up $10 twice. So all in all around $20 of usage, or about 4 months of the smallest DO droplet. Second, why would it become dependent? At worst you are back to the levels of today.

I think there are moral reasons for choosing not to use LLMs, but I think those are leaking into the arguments after supposedly ignoring them at the start of the article.

tonyarkles | 8 hours ago

I chuckled a little bit at this from the article:

"I’m refusing to use this software because it was made by humans"

I can, with very little imagination required, come up with reasons why that could be a valid position in the near future. Especially extended to:

"I’m refusing to use this software because it was made by humans without AI support"

For a while, I was writing C++ code that had pretty serious reliability requirements. Not safety-critical but even short (like 10 seconds) downtime was quite expensive. I used a decent chunk of open source tools for doing static analysis, but also leaned pretty hard on a paid Sonar license to double-check my work. It did an excellent job and I wouldn’t hesitate to get a license again if I was working on similar stuff in the future. That cost money and it was money well-spent.

Same goes for Claude Code. It helps me find bugs. Like simonw above, that tool has helped me find multiple white whale bugs that have followed me around for years, a few of them in upstream RTOS code. I don’t use it for writing code much, other than throwaway analysis scripts and things like that. Not sure if I’m going to go through the effort of rewriting the bug analysis to put together PRs that are “human authored”. I’ve read through the couple line fixes, I’ve tested the crap out of them, and it’s unclear if my effort is actually wanted or not.

I really have a hard time understanding the comments on this post - it feels like deliberate misrepresentation to try to make projects with AI policies look absurd.

that tool has helped me find multiple white whale bugs that have followed me around for years

Ok, awesome, but

Not sure if I’m going to go through the effort of rewriting the bug analysis to put together PRs that are “human authored”

...for your "white whale," your nemesis over literal years, writing it up yourself is suddenly too much effort?

tonyarkles | 5 hours ago

Yes. If I have to go through an entire parallel construction effort to pretend that the bug wasn't actually found by Claude Code while I was out for a walk with my wife, when I have a perfectly good walkthrough report already that wasn't human authored but has been seriously human-reviewed... that's not effort that my employer would encourage me to spend.

And it was found by Claude Code, yes, but that was informed by both my human intuition and my large collection of not-fit-for-human-consumption Markdown notes that I've collected over the last 3 years about the issue. The bug was ultimately found when CC was working on something entirely unrelated to the bug and had something like this in the output:

FYI, when I flashed this build it seemed like the processor locked up. I reflashed it and it worked fine. There weren't any errors either time.

Which made me think "oh, shit, I think this might be that bug and I've never been able to repro it on this debug board... nice" and set in motion a whole other session to nail it down with all of the context I'd gathered over the years.

"Remember that bug we finally fixed last week boss? I want to submit it upstream but to do that I have to rewrite the report and patches by hand otherwise the upstream project will reject the fix." isn't an easy sell, unfortunately.

joshbuddy | 10 hours ago

Comment removed by author

sunshowers | 7 hours ago

Pascal's Wager is a fallacious argument?

Yogthos | 6 hours ago

Which is rather fitting here given that the argument is based on a straw man.

simonw | 9 hours ago

More importantly, rejecting AI will not alienate any users. Nobody in the world is saying "I’m refusing to use this software because it was made by humans".

I've seen enough now that I am less likely to trust the security posture of an "absolutely no AI allowed" project than one that uses the latest AI models for security review.

I've run AI security audits against my own code written pre-AI and turned up a whole bunch of obscure issues (and it's the obscure issues that get you). I thought I was pretty good at writing secure code; I was not.

spc476 | 5 hours ago

Have you run LLM-security audits against your LLM-assisted projects? Are they more secure overall?

Yogthos | 6 hours ago

I love how every one of these articles will have a straw man like:

Developing your project become increasingly dependent of on expensive and completely proprietary software over which you have no control.

I guess the position comes from the fact that anti LLM crowd doesn't actually use or understand how this tech works. So, they're just constructing arguments in the abstract based on how they think this tech works and what it can do.

The reality is that you can run a model like a Qwen 3.8 on a laptop, and it can do plenty of real world work without you being dependent on any expensive and completely proprietary software. You run a tool on your machine that you have full control over. But of course, if you just ignore the technology you're discussing, then you can come up with all sort of nonsense to base your argument on.

natkr | 6 hours ago

The reality is that you can run a model like a Qwen 3.8 on a laptop, and it can do plenty of real world work without you being dependent on any expensive and completely proprietary software.

Can you understand and modify that Qwen model? Can you retrain it yourself to accomodate an ever-changing world?

If not, yes, it is still proprietary. Whether or not it depends on some external website. You also wouldn't call Windows "open source" just because you can download an ISO from microsoft.com.

singpolyma | 5 hours ago

No one understands it. But yes we can modify it.

But also you're not reliant on it. That's the reason people still use this tools to assist in authoring code and not just tell the model to do things directly. The code is still code and humans are still part of the process of authoring it and if every LLM went away tomorrow the humans could keep going.

aaronm04 | an hour ago

The code is still code and humans are still part of the process of authoring it and if every LLM went away tomorrow the humans could keep going.

That's the case with current AI models but I wouldn't be surprised if future AI models are trained to generate subtly human-unfriendly code. AI labs have an incentive to do this.

singpolyma | an hour ago

You mean if using some big company's model? Maybe.

aaronm04 | 42 minutes ago

Yes, that's what I meant.

FeepingCreature | 5 hours ago

Not only can you retrain it, hundreds of people have already done so.

k749gtnc9l3w | 4 hours ago

Slightly horrifyingly, given that both internally and externally a typical LLM (within the same architecture) is fine-tuned not retrained from scratch, and yes it happens a lot externally, I am not even sure that «preferred form of the work for making modifications» (the definition of source code from GPL) is not yet the weights.

Microsoft does recompile Windows from source code regularly, though.

Yogthos | 4 hours ago

Yes, you absolutely can do that using services like runpod, which I, and many other people have done. You can easily create LoRAs to tune the model exactly to what you're doing, and there's a plethora of tuned models on huggingface because people do this all the time.

aaronm04 | an hour ago

The crux of the issue is that the line between tool and collaborator is becoming blurred by LLMs; we don't expect our collaborators to be open source.

simonw | 9 hours ago

While every study on the subject tends to conclude that using LLMs is actually slower than writing code yourself

Really?

People who don’t want to contribute without LLMs are probably not able to review and understand the code by themselves (or they will lose that ability soon enough)

Really?

alanmeira | 5 hours ago

The second one is true. Properly reviewing and understanding a project takes time, most people pushing quick AI PRs are doing so after the agent decided to use the project and "found a bug", which might be just a hallucination.

tonyarkles | 4 hours ago

People who don’t want to contribute without LLMs are probably not able to review and understand the code by themselves (or they will lose that ability soon enough)

most people pushing quick AI PRs

You're arguing against a specific form of this and I feel like that specific form is what's causing projects to adopt alienating policies. I've been doing this for 30 years, I am an AI-skeptic who has been exploring these tools for the last 6 months and am blown away by some of their capabilities. I don't yolo slop to prod, it gets carefully reviewed, refined, sometimes thrown away and redone. If I find a bug in a software package, ask an LLM to put together a map for me of how that feature works and where the likely issue would be, use that map to manually inspect the source and figure out what's going on and put together a patch from that... some projects' policies state that they would rather not have my patch because I used an LLM in the creation of it.

FeepingCreature | 5 hours ago

I love METR but that study has done massive damage to the discourse. (Through very little fault of METR, to be clear.)

aaronm04 | an hour ago

This seems like a facile and biased analysis.

I actually think point 2 (the Debian position) should not be dismissed so easily. The author states that the slop will just seep in inexorably without really arguing that position.

Maybe we should take time to figure out (and spell out for would-be contributors) what exactly responsible AI entails. Maybe it means that you can use AI to find bugs, but not actually write code, for example. Or instead it could mean only optional tools can have AI contributions. So for example if you have a programming language project, the language server could be written using AI because it's an optional tool.