rust-lang/rust is adopting an LLM policy

120 points by wezm 14 hours ago on lobsters | 34 comments

JulianSildenLanglo | 13 hours ago

It's fine to use LLMs to answer questions, analyze, distill, refine, check, suggest, review. But not to create.

It's nice to have such a clear and reasonable policy.

nemin | 12 hours ago

I wish this single-line guidance included "[...] review, assuming you clearly mark LLM output and exercise your own judgement." The section below clarifies things excellently, but I find this excerpt a little too vague on its own (or perhaps unintentionally giving a false impression).

JulianSildenLanglo | 12 hours ago

I think they mean that you're allowed to use LLMs to review your own work before sending in a PR, in which case there's nothing to mark.

nemin | 12 hours ago

But the "General Rules" section says this:

No one except the author is required to read LLM output unless they choose to: LLM output isn't allowed in public docs, PR descriptions, or Github comments unless it's clearly marked; reviewers aren't required to look at LLM PRs if they don't want to.

[...]

Disclosure is required for machine translation, "trivial" changes, discovering bugs, and reviewing other people's work using an LLM. We welcome messages posted in your native language; English translation is not required to contribute.

Don't get me wrong, I welcome this policy, I just wish the "one line zinger" was a little more detailed, because if you come away from the article without reading on (or, more realistically, if secondary sources pick up this quote without the surrounding context), one could easily get a wrong idea.

pyfisch | 11 hours ago

This appears to be only a summary. There are multiple carve-outs for generated code linked in the article. However the explanation is quite long and I am not sure I understand the details.

sanxiyn | 13 hours ago

We welcome messages posted in your native language; English translation is not required to contribute.

I very much welcome this. I will post all my comments in Korean to rust-lang/rust from now on.

Gaelan | 11 hours ago

I’d always rather someone post in their native language than machine-translated output in a language they don’t speak at all. That way, I can:

  • read it myself, if I speak the language to any extent
  • ask someone I know who speaks the language to read it for me
  • machine translate it, but if something looks weird, investigate using a dictionary or the above methods

If I just get machine-translated output, I can’t do any of those things - I just have to hope the translation was accurate.

Of course, I’d the author speaks some of my language, it becomes more of a balancing act, as most people probably will just machine translate it, and the author may be able to do better than the machine translation.

Commendable in spirit, but any kind of friction drives away your audience. In practice, I suspect that the overwhelming majority of readers would simply scroll past the comment rather than machine-translate it, even if a "translate only this message for me" were right there (and for a significant subset of users it currently won't be). After all, they likely don't know what they're missing.

As a poster, the incentive for engaging with an existing community is to use the language most of it clearly understands at least to some extent.

pektezol | 9 hours ago

I agree with this, it does seem baffling to me why Rust and Zig are adament on no machine-translated comments. A PR containing 20 different languages would straight up be chaos.

JulianSildenLanglo | 9 hours ago

I think it's at least partially to cut down on the "I didn't write this using AI, I just translated it using AI" excuses.

A PR containing 20 different languages would straight up be chaos.

Maybe at first, but if this culture is normalised, then it will likely become normal to just machine translate it yourself, assuming you care about participating in the discussion.

mitsuhiko | 3 hours ago

That’s what modern Twitter and Reddit is like and I can’t say I enjoy that.

I was going to joke that I also don't enjoy them, because they are twitter and reddit, but my real thought is that those spaces have a different mode of communication, especially twitter (ugh, I just remembered it).

I think your idea of having people post both the MT version and the original makes sense. that could also be encouraged as part of the culture.

mitsuhiko | 2 hours ago

I mentioned those two places because both are now auto translating which has lead got people speaking different languages engaging with each other in a way that was not taking place before.

shrike | an hour ago

People "punch up" their perfectly fine non-native english too lightly, which makes them sound like a sycopanthic ChatGPT - and that weirds people out.

The beauty of the english language is that you can mangle it pretty badly and still get your point across.

who is your "audience" here? the goal of an issue isn't to get as many people to read it as possible. in fact we often have the opposite problem, that there's too many comments and we can't tell what's important and what's not.

hoistbypetard | 10 hours ago

One feature that I'm really enjoying in vivaldi right now is the translation sidebar. I can highlight text and open the sidebar, and get a cromulent machine translation.

https://l.e42.us/QFLrFmdm

valdemar | 8 hours ago

Firefox have also added this kinda. If you enable the local translation you can select text and right click and select "translate section" for a pop-up

apropos | 2 hours ago

You can also enable local translation selectively while keeping the global AI kill switch checked.

mitsuhiko | 9 hours ago

I thought so but now that it’s becoming more common I changed my mind. It’s much preferable to get someone’s AI translation if it’s done well. My favorite comments now are coming from people who send me their AI translation but also leave the original text behind an expandable block. I have seen a few Chinese users do this now and I think it’s the optimal path at the moment.

Of course, I’d the author speaks some of my language, it becomes more of a balancing act,

Back in the days before decent machine translation, I sometimes needed to prepare business website text in French. Now, I'm fairly confident reading French. And I can write it... adequately enough for bare communication.

So instead of hiring a translator, I would write what I wanted in French, and then hire a French copywriter to actually make it literate. I could then check the accuracy myself, and I'd usually have a feel for the tone and style they used.

But:

  1. Operating in a second language imposes massive costs, even if it's strong enough to get admitted to a university that uses that language.
  2. Reading skills generally outstrip writing skills.

So if a native French programmer (who, in many subfields, probably reads technical English just fine, or did up until the invention of LLMs) wants to use machine translation in a group discussion and proof-read the output, then I'm willing to read some slop prose. If they read English well enough to proof-read, this will likely be the most accurate and the most effort-efficient way to communicate. And I'm painfully sensitive about always putting the heaviest linguistic burden on someone else, which is unfair as hell.

I mean, I don't want to read any more of Claude's prose than necessary, because I already do that more than enough, thank you. But I won't let my Claude fatigue impose unnecessarily heavy burdens on speakers of other languages, because I've been on the wrong side of those burdens and I know the cost. As a general rule of thumb, monolingual English speakers should look for some way of sharing the cost of efficient, accurate communication where they can. Because they routinely underestimate the costs being paid by non-native speakers of English, and they usually have no experience with the ettiquette that inevitably arises around such things. Requiring someone else to always speak your language is a lot like requiring someone to always cook dinner for you, with no way to pay them back. Even if they're happy to cook, it creates an imbalance that the receiving party should be looking to repay somehow. And this very much affects how I emotionally react to rules like, "Nobody should be required to read LLM output." Sometimes reading (edited) machine translation is both the most efficient way to communicate, and the one that prevents the non-native English speaker from always having to "cook dinner."

One reasonably optimal solution would be to encourage people to post in their native language, but then allow them to create (and edit!) an LLM translation into the dominant language of the group. Giving people with English reading skills the ability to proof-read translations is a net win for both accuracy and burden sharing. Design the UI to taste but keep both versions available.

(I mean, I get the anti-LLM position. I personally worry that we're closer to building SkyNet or some other grim dystopia than most people imagine. But that's a problem to be solved by difficult things like politics and international alliances, and not by focusing our frustration on non-native English speakers trying to minimize their portion of the linguistic burden.)

intelfx | 8 hours ago

I’d always rather someone post in their native language than machine-translated output in a language they don’t speak at all. That way, I can: read it myself, if I speak the language to any extent <...>

Sorry, but this does not compute.

Purely probabilistically, there's a much higher chance that the speaker of uncommon language $X also speaks English (the lingua franca of tech) to some extent, than the chance of an arbitrary English-speaking reader to also speak uncommon language $X.

From this it obviously follows that there's a much higher chance of the sender being able to (somewhat) vet a machine translation of a message he's sending, than of the reader being able to vet his own machine translation of a message he received.

And, of course, I don't even mention the fact that multiple readers of such a message will come up with their own machine translations, all of which will be unique, and some of which likely won't perfectly match each other, and the chaos will ensue if this rule actually gets exercised to any significant degree.

This (banning LLM translations on the sender side¹) is a completely illogical move, which makes no rational sense other than to plug the "I just translated it" loophole at a great expense in communication quality.


¹: OTOH, at least Rust had the common sense to restrict those rather than ban outright

travisgriggs | 5 hours ago

There’s a more general abstraction. My pet peeve is ai-washed person to person communication. I work with an anxious colleague, who originally felt empowered by ai tools. They could rough out a concern, and then let ai wash it up for me. I hated this. It lost all authenticity. And left me guessing what was really going on.

It wasn’t an entirely foreign language they were speaking, but even when we all speak the same language, we use it in subtly different ways. Text rendering of ideas and speech is already reductionist. And I found that the later ai was used in the process (I.e. use it at the consuming level tether than the production point) benefited these communications.

quintus | 28 minutes ago

My native language is not English (it is German). I sucked in English at school and did not see the point of learning it. Then I discovered tech and programming, back somewhere in the 2000s, and -- bummer: everything I wanted to know was only available in English. I forced it down my throat and actively subscribed to the Ruby-Talk mailinglist in order to properly learn how real programmers converse in English. I picked up most of my English tech vocabulary there (if not most of all my English vocabulary), and once I felt confident enough, I even tried writing there. It worked. Needless to say, the year after I went down programming my English grades went straight up in school. My English teacher was puzzled.

So, this is how I learned English proper. If LLMs would have been available, I would never have learned it in a respectable way. I fear that with the advent of LLMs, this incentive to learn English -- which I believe is all but rare in tech circles -- collapses. As a consequence, in-person conversations e.g. at conferences may become awkward in the long term as they are going to become dependent on the LLM translation tools as well, because the younger generation did not see the point of learning English, much as I did not see it before I was forced to learn it in order to accomplish my goal of learning how to program.

skade | 5 hours ago

When I was part of the Rust team - which is a debate-heavy project - I regularly had to remind first-language speakers that even if their peers can write English just fine, it's often an additional mental tax for them, so going lighter in discussions is helpful.

I'm quite fine with people writing down their thoughts in their native language and then translating it.

the code itself is the smallest and in some ways least important part of the change. We care much more about authors understanding what the code does, planning how it will change in the future, and deciding what it should look like. The code itself cannot help with any of those.

This struck me so hard I just wrote a post to my team about it. Yesterday I reviewed an LLM-generated PR from a teammate that did identify and fix the source of a bug, but did so in a ridiculous way that resembled a Rube Goldberg device. I don’t know whether the co-worker reviewed the change at all before creating the PR.

I did the same!

If I have to choose one or the other, I’d rather require the team to review the prompt and implementation plan together and let Fable write the code unreviewed, than have an individual run Fable and dump the code on the team for review. (But then I was a fan of design reviews decades before LLMs.)

gignico | 10 hours ago

This seems all very sensible and reasonable and I’m happy to see such a well thought take in a project like that. I hope more projects may take similare stances.

cryptocode | 6 hours ago

It's fine to use LLMs to answer questions, analyze, distill, refine, check, suggest, review. But not to create.

Those are all acts of creation though

dpc_pw | 2 hours ago

Please put summaries/TL;DR on comms like these. I don't really care if it was written by a clanker or not. I want to know how am I expected to behave without needing to sieve through pages and pages of text like I'm dealing with a government.

IMO. In generally every formal-ish communication should follow a progressive disclosure style, where it starts with a most concise summary, then progressively gets into more and more details that people can read on a per-need/interest basis.

Rust’s new LLM policy says:

  • Private use is fine: understanding code, brainstorming, summarizing, and reviewing your own work.
  • Public LLM-written prose is generally banned: issue text, PR descriptions, comments, docs, diagnostics, and substantive code comments.
  • LLM-assisted reviews and bug discovery are allowed with disclosure, but humans must verify findings and make decisions.
  • LLM-generated code PRs are allowed only by prior agreement, with disclosure, strong tests, human understanding, and usually only for non-critical code.
  • Contributors must disclose relevant LLM use, but reviewers should not publicly accuse people based only on writing style.

The goal is to protect reviewer time, accountability, and genuine technical understanding.

Seems reasonable to me. Disclosure: clanker generated that summary.

mitsuhiko | 2 hours ago

You need a TLDR on a one liner? The policy is literally just

It’s fine to use LLMs to answer questions, analyze, distill, refine, check, suggest, review. But not to create.

dpc_pw | 2 hours ago

The policy is literally just

That line is ... 3 screens down inside the article... . That's kind of my point. I don't have time to treasure hunt ATM.

Edit: And it is my general recommendation about effective communication. Assume a lot of people will read just the first part, and make sure you made a good use of it. You can be irritated about me being a lazy distracted idiot, but I assure I'm not the only one.

mitsuhiko | 2 hours ago

The article is now how people are going to find it, they are going to find it from the contributing guide that is linked when you make an issue. Which directly links to the LLM usage policy here: https://forge.rust-lang.org/policies/llm-usage.html

That is short enough, even for the tiktok generation ;)

clerno | an hour ago

I found this similar to the C3 LLM policy that we decided on about 4 months ago, after getting low quality PRs: https://github.com/c3lang/c3c/blob/master/CONTRIBUTING.md

It might be interesting to compare to.