I'm like as anti-AI as you can get, but I'm not sure how they think this will play out. They basically said "slop is banned, but we can't go after all LLM use that would be absurd"
Without clear rules that can and will be enforced, it will just become a game of who like or dislikes who...
I think people are a bit too focused on the enforcement aspect. I could be wrong, but I think the more important point is that it's an expression of their values and the type of community they wish to foster.
If you want to vibe code projects you can probably circumvent the rules and pretend you're in the community, but you might instead prefer to join a community that share your values.
It's their rules, and I won't tell them what to do, but I think the issue with "what type of community they wish to foster" is that this is ill defined. If the stance is "no LLM" that is very clearly defined. But beyond that, it's arbitrary and will lead to fights about it.
I think at the scale of Codeberg, it's fine if they figure that out as they go, it's new social stuff for a lot of developers anyway. It's nice to know that there are options.
Potential ForgeFed improvements might also mean that things aren't going to be too crazy.
Right, so read it as roughly "we can't define where the line should be precisely but we want to build a community that is interested in figuring out how close to maximum human interaction as is still practical".
The same can be said for the site guidelines here on Lobsters. The lines are blurry, a lot of offences slip through, and those the moderators catch are often caught only after most users that would see them already have. The moderators are also regularly accused of bias and favouritism.
All the same, I'm glad the guidelines exist. They set expectations and let the mods ban the worst offenders, such that the site is a much more pleasant place to be. I think that's what Codeberg is going for.
I think this is similar in intent to a "fair use policy". They probably just want a sign to tap when discovering projects that are overusing resources.
You don't think it will be acrimonious, the process of making accusations and then executing sentences? Even if all the first decisions are totally uncontroversial and target "the worst of the worst", how far do you take it?
All moderation decisions are acrimonious: people don't like hearing that they've broken a rule, after all. I still prefer this over an attempt to make "clear rules", which i feel are actually more likely be gamed by people trying to find a way to be 1 step behind the line.
Honestly, though? I don't expect this to be too big a deal. Why would a vibe coder willingly choose codeberg, a hosting site actively hostile to them, over github, a site actively courting them?
I generally think rules should be designed to be played like a game. The alternative is lawfare -- the idea that everyone is always breaking the rules and you just choose who you want to weaponize them against, probably based on politics
I'm concerned by how disconnected this blog post is from the text that's been added to Codeberg's terms of service.
The blog post makes several points about the costs that LLMs impose on society and ways that LLMs discourage collaboration. Of those points, I found this one to be the most convincing justification for banning LLM-generated code on Codeberg:
Although often well intentioned, sharing the result of an prompt and calling it "libre software" does not make the world a better place. Codeberg is not and does not want to be a place to dump such generated single-use software that no one else will ever look at. We are a place for people to collaborate and improve software together. Within this context, the recent votes can be understood as a reconfirmation of those principles: As we want to center on human collaboration, we will not actively support or engage in the creation of LLMs and will not put our limited resources to use for storing single-use software that would pollute our FLOSS commons.
But then I followed the link in the very last sentence of the blog post and read what they've added to the ToS:
You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as Claude, OpenAI Codex). Such projects having an unclear copyright status (see requirements § 2 (1) 1 and § 2 (1) 3) and furthermore have little safeguards to ensure that they do not include harmful code (c.f. § 2 (1) 5).
Those are entirely different points, and neither one is so much as mentioned by the blog post: the words "copyright" and "harm" do not appear in it. I suppose the copyright point is somewhat self evident, but the claim that LLM-generated code is more likely to be harmful is anything but.
I'm also surprised that the ToS entry seems not to have been proofread. It should say "Such projects have" instead of "Such projects having", "'generative AI'-tools" shouldn't be hyphenated, and it's inconsistent to write "OpenAI Codex" but not "Anthropic Claude". The text is fine as a first draft, but it makes me wary of Codeberg's governance that it went to a vote of over 400 people and got added to their terms of service without being revised.
I'm also surprised that the ToS entry seems not to have been proofread. It should say "Such projects have" instead of "Such projects having", "'generative AI'-tools" shouldn't be hyphenated,
Maybe they should just phrase their ToS in German, matching their residence, to reduce grammatical complaints. (OTOH there is a defined-and-kinda-binding Standard German, while the same doesn't really exist for English 🤷)
I've been a Codeberg member for several years; most of the things I've voted on have been in German with an English translation listed secondarily. Not sure why this one is different; maybe because it's the TOS vs the Codeberg eV bylaws?
If your work fits into these cases, it is unlikely that you are affected at all:
Projects who have an active community that cares about and maintains the software
What the authors of this post don't seem to realize is that it is possible to care about software while maintaining and developing it with the help of LLMs.
Case in point is MapLibre. We ported one project to Codeberg, but the maintainer makes heavy use of LLMs for development. So we have to close that repo again. I am also thinking of a project like Ghostty. Another serious project that wants to leave GitHub, but cannot move to Codeberg now.
This move by Codeberg is an effective ban on all LLM usage, because the rule is very vague and even the guidelines in this blog post are contradictory. No open source community with even a slightly liberal policy on LLM use is going to risk having their community booted off Codeberg.
Since the rationale is "Many LLM-based projects use an extraordinary amount of resources, especially compared to the developers and users involved", a good rule of thumb seems to be that they'll kick you off if the "heavy use of LLMs for development" involves spawning hundreds of CI builds on their gear, or sending gigabytes of artifacts back and forth, in a way that is atypical for human-centric development.
Also, Codeberg is people: Getting the boot there isn't the automatic anti-abuse abuse we got used to with the IT megacorps. I'd expect "Uhm, your project might not be a good fit" as a discussion starter, potentially followed by an offboarding plan to get your project set up elsewhere plus a permanent limited presence on codeberg to forward folks.
I don't see that "rule of thumb" mentioned anywhere in the blog post. This rule targets all projects that use LLMs. Whether they use significant resources or not.
"Uhm, your project might not be a good fit"
Which effectively means "Your project is not welcome on Codeberg". I understand that it's possible to self-host Forgejo. In understand that there are alternatives. But I had high hopes that Codeberg could be a welcoming community for various kinds of open source projects developed in good faith, with different stances on LLMs. I have not given up hope yet that Codeberg can be that. My member application was actually already pending.
I don't see that "rule of thumb" mentioned anywhere in the blog post
I derived that concept from
To us, it seems ridiculous to see projects with a single developer and virtually no users consuming as much or even more resources than some of the largest community projects on Codeberg, which operate frugal with CI/CD and storage resources.
and
We will also not spend significant amount of time and resources to automatically scan content on Codeberg.
Given
So while the following use cases are discouraged (similar to private repositories), they are likely to be tolerated in practice: … Specific tools and custom scripts that would be unlikely to find a community anyway, even if they were not LLM-generated
it seems quite clear that they're looking for projects where people interact to create code (see also the paragraph titled "The development team of none").
"You and your AI monkey" is a bit of a mismatch to that goal of theirs (just like my solo-developer project, really), but they also don't seem to be interested in kicking off accounts as long as their presence doesn't make the system worse for everybody else.
These are deliberate choices made as much by MapLibre and Ghostty as by Codeberg. Codeberg changing their policy would just be letting someone else impose their views on the Codeberg community. In the words of Michael Bolton; “Why should I change? He's the one who sucks.”
No open source community with even a slightly liberal policy on LLM use is going to risk having their community booted off Codeberg.
We’ll see how it plays out, I don’t agree that policies have to be completely closed to interpretation, I don’t think that is even possible. I am maintaining a fork of an LLM-heavy project on Codeberg which has a no-LLM policy in place, but still contains LLM-generated code inherited from the upstream project before the fork. My interpretation leads me to think we are still OK, but if not I’ll respect that and will have to move elsewhere.
Also it's a weird idea of what a successful project is. Some good software will have a community. Some bad software will too. Some will have a single maintainer, no love, and barely anyone caring about it until it goes missing. We had the corejs drama already. Does tzdata have a "community" as such? I'm the only maintainer of a mostly vibecoded utility which allows some other projects (with communities) to compile - does that quality?
I can't force them to not be selective, but the line they try to draw is weird and feels anti-opensource to me.
Agreed. If it's about resources (CI use or storage) then just cap that for non-popular projects?
If it's about community, that seems to have little to know correlation to quality?
I think it's a real shame a lot of the open source world has came down extremely anti LLM. I have done so much tinkering with open source projects (and finally switched to desktop Linux, as it's trivial to fix annoying bugs and config options with an agent), yet what seems to be like a majority of open source projects have taken an aggressively anti LLM stance.
I do totally understand that a huge influx of AI slop is time consuming and the mostly volunteer community can do what they want - but it seems to me the right thing to do would sit down and work on better tooling (for example, auto rejecting pull requests from people with no open source "track record" on your project, with some way for people to bootstrap it) to help shift through it.
For a lot of us, open source is where we go to unwind and tinker and mess around in our spare time. We don't want to deal with any corporate bullshit. Pounding tickets and uninspiring work should be squarely left at $DAYJOB. We're in it for the fun and intellectual challenge. We like building things with our own hands. Suddenly loads of people turn up and ruin the fun by submitting slop PRs, and hammer our services with scrapers, turning our fun side project into an admin chore. What's the point then?
I totally get that. But for many people "new(er)" to open source, they are getting the fun and intellectual challenge from building and fixing stuff with LLMs.
I've fixed so many problems in Linux locally that were fun to have an agent fix (for example, reverse engineering the new Android Quick Share protocol which was broken on Pixel devices for sharing files to my Linux desktop). It was a fun evening testing it and working with the agent to figure it all out. I don't have the time to do a multiday reverse engineering project like that by hand, and now something that was bugging me daily works on perfectly on my desktop.
Is this LLM "slop"? Yes certainly. Is it valuable to me and others? I'd say yes, it took time to test with a real device that someone doesn't need to spend now. How do we square this circle?
But of course, I understand the admin problem, but my view is that the platforms need to solve this for open source project maintainers, like dealing with scrape protection and building in automoderation tools that reject nonsense contributions, while still giving some sort of path .
Obviously everyone's opinions are different on this and I'm not trying to tell people what to do. But it seems to me that open source has went from "we don't have enough contributors" to "we have too many contributions, ban them all" when I think there has to be a middle ground of some sort.
But it seems to me that open source has went from "we don't have enough contributors" to "we have too many contributions, ban them all"
Regardless of how much fun you're personally having with LLMs, the average quality of those contributions dropped to negative. Previously, someone who showed up to contribute to a project showed drive and curiosity, and they were likely to try to learn more and stick around. Some of these people would eventually become maintainers - this is how open source projects grow and survive, and was also an on ramp to programming for many of us. Now, the "contributor" is likely to be a bot with little to no human involvement and a very low likelihood of any involvement after.
What a massive loss for all of us this has been, and LLMs are the cause.
Unfortunately some projects are banning contributions entirely - it is unfortunate but I also understand the decision practically. Others are just trying to get back to their previous signal to noise ratio.
I've been a paying member of Codeberg for 2 years. I didn't even realize this vote was happening. I just checked my email and found this line buried in the middle of an 1,800-word email.
Discussion of accepted proposals & spontaneous topics
Make presidium discussions available to members
Statement about use of AI by Codeberg
(Placeholder for further proposals)
This was the only announcement I received that this vote was even happening. The email says there will be a ballot, but I never received one. I'm not sure if I'm an anomaly or if many other members did not receive notice about this vote.
When you became a paying member, you should have had the option of becoming an "active" member or a "supporting" member. If you're an active member, you should have received the invite to the annual assembly where all the proposals were discussed and debated, and then links to vote on each of the changes. If you're a supporting member, though, you don't get an invite or a vote (or a bazillion emails discussing pull requests to the org docs).
You should be able to change your membership to active whenever if you'd like.
If not, you can also keep an eye on the pull requests to the org documents: https://codeberg.org/Codeberg/org to keep up to date on possible changes.
I believe I'm an active member. I was invited to the annual assembly, although it's possible I switched to supporting at some point because I got exhausted by all the noisy changes to the org docs.
They've also declared that they're blocking cryptocurrency related projects. They have explicitly said "this is not a neutral space, Codeberg is political" on the subject.
Not to concern troll, my day job is in crypto. But still. This is your backup service, to save your work in case your computer fails; this is your public identity as an open source community project. Why would you select, for these tasks, a service that advertises that tomorrow it may not like your face, due to the shifting winds du jour, and your backups and identity are gone? You want your infra to be neutral; that is the job of infra.
I see your argument, but my gut says no. I actually like that they have the balls to apply some ethics to what they're doing. Ethics seems sorely missing from computing.
Whether you agree with their particular values or not, perhaps you can agree that the modern enshittified ad-infested user-hostile tech hellscape is a direct result of organisations not publicly stating any values and standing by them. Google silently dropping their "don't be evil" motto is a culturally significant symbolic milestone IMO (even if they had already jumped the shark long before).
I don't agree with that, no. Enshittification is a business strategy that usually implies being owned by someone who would prefer you not have values, but I don't think the causality runs in the other direction, where values prevent enshittification (the abstract concept of values, obviously specific engineering values contradict it but those aren't the ones Codeberg is expressing). I also don't think GitHub (the obvious comparison here rather than Google) qualifies: the core service, despite the constant outages, is basically unchanged and just as free. If you're migrating from GitHub to a neoforge, it's probably related either to outages, Microsoft's existence, or the aesthetic of Copilot "intrusion" (which I suppose is a better counterargument to what I said: this is their audience).
I mean, obviously specific values will not prevent enshittification (as an extreme example: profit and growth overriding all other concerns). So not having any clearly stated values should be a huge red flag, as you have no idea where they stand. But you could counter that talk is cheap anyway, and companies have been known to completely reverse course.
Personally I've disliked GitHub for a longer time than the Microsoft takeover: their gamification (with badges, stars and whatnot) and focus on "social coding" was a huge turnoff for me. In retrospect it's an obvious (but arguably failed) attempt to make the site more "sticky" and addictive. We can guess what values underlie those changes. Also, their "macho bro culture" was well-known even back in the day when they just got started. I distinctly remember a picture of Scott Chacon wearing sunglasses, walking towards the camera with explosions in the background, whisky glass in hand.
So yes, the service is essentially unchanged but I'd argue many people haven't had a close look at GitHub's apparent values in the first place and just went with "the default".
The values of Codeberg specifically stated here don't prevent enshittification, for sure. But the fact they are a collectively owned nonprofit at least rules out the typical reason for enshittification (profit, duh) and makes it harder for a single "owner" to make unpopular decisions. Of course there are (AFAIK unstated) values underlying the decision to set up the organization like that.
I don't get this at all. If you're in cryptocurrency you can pay for your own hosting instead of leeching off a non-profit free forge. Like, you're in it to make a buck, taking resources from idealists is really low.
We must all make ethical choices every day. Having people disagreeing with yours is not a case of "tomorrow it may not like your face" (sic). The fact that you say that tells me that you perhaps haven't evaluated enough the ethics of make believe currencies to understand why others might not feel the same way you do.
I select them exactly because they are not afraid to take a stand AND they let me freely use the software which runs that very service ( that they have all the rights to manage as the community sees fit).
Wow this is quite shocking already. This can literally mean anything. If you have an anti-fascist project that alerts communities of ICE raids hosted on Codeberg, it could be seen as "harming the reputation of Codeberg", because the MAGA movement will not like it.
If your physics library starts being used by some war company, boom, your project can now deemed to be harming the reputation of Codeberg.
I really think Codeberg should see the ToS as a sort of constitution, requiring a three-fourths majority to change, because otherwise they will ban more and more content that half of the members don't care about or have prejudices against.
I suspect that forbidding projects which use LLMs to produce code will seem as strange in twenty years as forbidding projects which use high-level languages strikes us all now.
Back when compilers were invented, did folks point out that a compiler uses more energy than does an assembler, and an optimising compiler uses even more? Did they complain that using a high-level language enabled a single programmer to produce what used to take a team? Did they complain about a ‘digital divide’ between those who could afford procuring and running high-level languages and those stuck with machine code?
At a different inflection point for technology, I do recall that a lot of folks spent a lot of time believing that they could manage memory better than a garbage collector (pretty sure I did!), and that automated bounds-checking was too costly for performance (guilty!), and so forth. But even back then I would have been shocked if, say, SourceForge had refused to host projects in Lisp or Haskell.
It’s just so weird that a technology comes out which empowers programmers (which most of us here are) and software users (which almost everyone on Earth is) to such a frankly amazing extent, and to see folks so profoundly negative about it.
Nobody complained of a digital divide because compilers didn't drive up the costs of computing equipment due to an investment bubble. Nor did compiler developers abusively spider information services, making it more difficult and expensive for smaller organizations/individuals to keep their services online for everyone.
LLMs pose a qualitatively different set of issues than the analogies you suggest. It would be more constructive to engage those issues on their own terms, without strained analogies. The LLM industrial complex is unlike compilers, and unlike higher level languages, in every line of criticism explored in that blog post.
It's so weird that their proponents keep trying to emphasize what similarities there are rather than addressing the criticisms directly.
Amnesty International finds that standalone generative AI systems, based on unlawful web scraping, depend on mass invasions of privacy by design, and are fundamentally incompatible with IHRL.
Do Amnesty International have a position on why scraping the web to build a search engine isn't incompatible with International human rights law?
Well, GDPR guarantees the right to data erasure, and google will in limited circumstances let you have a "right to be forgotten", as in, they'll intentionally de-index pages on behalf to maintain your right to privacy. Can't do that with an LLM without totally retraining it (i think? not an expert). This is the only point in the doc (section 6.1) that seems directly related to normal search engines as an abstract concept.
That said, you can find docs on their website about google, predating LLMs, and equally here they view Google (as a broad organization that includes ads and adtech) as not practically respecting the right to privacy, either: https://www.amnesty.org/en/documents/pol30/1404/2019/en/
Unfortunately it’s really an argument against open weight models more than anything. Right to be forgotten is implemented via filtering on the way out on search engines, not by data deletion.
I’m afraid people might be reaching for the GDPR gun and end up potentially in a worse situation where open weight models are restricted.
If I’m not mistaken, the GDPR requires data holders to delete the data. Search engines instead should not list data that is not hosted by them.
So, if a LLM provider implements it by filtering alone, it is probably in breach of the GDPR (IANAL/TINLA/etc.). It makes no difference if the model is open or not.
Also, if open model maintains show a good track record of complying to the GDPR in that regard, that is even better than fake compliance by obscurity I believe.
The GDPR does not really regulate data storage but processing. You just need to stop processing that data. In practice a data deletion request is implemented as writing a row somewhere that data was deleted and will no longer be processed.
In particular that is needed because you cannot delete those rows from backups.
In theory the text says erasure but in practice we know what data protection agencies are okay with and that is just to ensure that production systems are seeing that data.
How would output (or input) filtering prevent information from leaking through second-order queries? "What is the former creator of Flask up to these days?" doesn't need to mention "mitsuhiko" or "Armin Ronacher" anywhere.
The core premise of LLMs is completely counter to the idea that they could ever meaningfully comply with basic privacy. So either they are incompliant or must be banned, or they are worthless and we'd lose nothing by banning them anyway.
It's unlikely that I could use GDPR to remove my Wikipedia page, so there are some limits there with regards to public information. But to answer the non strawman version of this question: labs are already doing this! Look into how refusal steering works. The "issue" is that on open weights models you have a lot of control over that and you can undo a lot of it. People have been doing this to get Chinese models to be less restrictive when it comes to certain topics that are trained on, but not politically acceptable in China.
Perhaps also the difficulty in understanding where information comes from. With a search engine, that's obvious, and you can go after both the search engine to de-list something and the originating web site to get the publication itself pulled. With LLMs, things pop up in the output and it can be impossible to tell where it came from.
Case in point: I've had "AI search results" which claimed something untrue (AFAICT) and pointed to sources which didn't actually contain these "facts". So either the LLM fabricated it out of whole cloth, or it appeared elsewhere in the training data. No way to know.
I don't know, but I can hazard a guess at one line of reasoning. Suppose on a website, someone has posted a picture of Alice wearing mirrorshades. Alice doesn't like the picture of herself wearing mirrorshades, but perhaps the site/operator in question isn't kind and won't take it down. She can ask the search engines to avoid linking to that site, and most could easily remove it from results. It's then also usually trivial to check if mirrorshade Alice is on the search engine at that point.
It's much harder to know if a picture might end up somewhere embedded in the weights. It's not clear what the training data of many of these systems is, and it's not clear if you are somehow embedded in the weights. If someone asked for a picture of Alice, they might get regular Alice or, if the training data is minimal, they might only have mirrorshade Alice to offer. More subltely, if they're asked for "cyberpunk Alice", that may yield mirrorshades Alice instead of, for example, hoodie Alice.
Asking the AI company to redact mirrorshade Alice from a (possibly open-weight) model sounds pretty complicated. They might remove mirrorshade Alice from the training data on request of Alice, but it might be hard to know she's in there to begin with unless the company makes their training data public. They might try to proactively remove Alice from results yielded by AI companies, but the capabilities for that don't seem mature. (with a search engine, it is often easy to know via site: style directives if something is included/linkable, and many index-based search engines can be intuited by humans).
OTOH, some search engines might be incompatible with human rights laws (depending on what part of the web you're scraping and how).
It's so weird that their proponents keep trying to emphasize what similarities there are rather than addressing the criticisms directly.
I guess one way of looking at this is that there's a history of AI proponents seeing arguments they think are overblown or made in bad faith by the anti-AI crowd, and set against benefits that are obvious to them, it's easy to start downplaying even the stronger criticisms.
You could make similar arguments on both sides about organized religion. Abuses by organized religious institutions would not, by themselves, justify excluding every project influenced by a religious person. I'm not trying to argue that AI and religion are the same, but I am pointing out that in both cases a complicated argument about institutions, individual responsibility, and actual harms can drift away from reality and towards identity pretty easily even when both sides claim to be arguing about the "facts".
Even if the alleged harms are real, they do not automatically justify excluding a particular LLM-assisted project. You still have to connect that project to those harms and show that excluding it would actually help, and isn't disproportionate.
(For what it's worth, I'm actually kind of fine with Codeberg saying the direct economic burden is enough reason here. I just doubt that was the main thing driving the votes.)
Most of the criticism against so-called “generative AI” has nothing to do with the technology itself but rather with the absolute insanity that has taken over the industry.
Just like people like myself have nothing against nuclear fission technology itself but strongly oppose to nuclear WMD.
did folks point out that a compiler uses more energy than does an assembler, and an optimising compiler uses even more?
No, but in the 50s, they did complain that using a "compiler" was a waste of perfectly good CPU time back when you leased the mainframe from IBM or the seven dwarfs and you paid plenty in electrical costs to keep it running. Computers were cheaper than programmers at that point.
Did they complain about a ‘digital divide’ between those who could afford procuring and running high-level languages and those stuck with machine code?
When I got into computers (mid-80s, as a teenager), the only language that was "free" was BASIC that came with the computer. Everything thing else, assemblers, compilers, linkers, applications, all cost money. I had to save money to get my first assembler. I paid for two C compilers (one for MS-DOS, one for the Amiga). And I only got the C compilers in the 90s, when compiler technology had advanced enough to make it generate "slightly less than horrible code".
and that automated bounds-checking was too costly for performance
Only for a certain class of programmer. There were programmers in the 70s that begged compiler writers to keep bounds-checking intact. Sadly, management drank the benchmark Kool-Aid and thus only now are we getting languages with bounds-checking.
It’s just so weird that a technology comes out which empowers programmers and software users
I grudgingly agree with "software users"; I hate using PHP, but I don't begrudge anyone who uses it to get their vision implemented. I just don't want to have to use or maintain such a thing. And I'm less sold on "empowering programmers." Sure, some programmers have been empowered. I haven't.
Back when compilers were invented, did folks point out that a compiler uses more energy than does an assembler, and an optimising compiler uses even more?
People pointed out that you're better off putting the source code into the code repository rather than the compiled result (and that also included heated debates about scripts generated by autoconf/automake, where the rough consensus these days seems to be that they belong in release tarballs but not in a repo).
So far I haven't seen repos that are useful while consisting only of prompts, so the comparison "compiler <-> LLM" seems to be stretched on that part alone.
Back when compilers were invented, did folks point out that a compiler uses more energy than does an assembler, and an optimising compiler uses even more?
In the 1950s von Neumann was employed as a consultant to IBM to review proposed and ongoing advanced technology projects. One day a week, von Neumann "held court" at 590 Madison Avenue, New York. On one of these occasions in 1954 he was confronted with the Fortran concept; John Backus remembered von Neumann being unimpressed and that he asked, "Why would you want more than machine language?" Frank Beckman, who was also present, recalled that von Neumann dismissed the whole development as "but an application of the idea of Turing's 'short code."' Donald Gilles, one of von Neumann's students at Princeton, and later a faculty member at the University of Illinois, recalled that the graduate students were being "used" to hand-assemble programs into binary for their early machine (probably the IAS machine). He took time out to build an assembler, but when von Neumann found out about it he was very angry, saying (paraphrased), "It is a waste of a valuable scientific computing instrument to use it to do clerical work."
I suspect that forbidding projects which use LLMs to produce code will seem as strange in twenty years as forbidding projects which use high-level languages strikes us all now.
This is not really matching the spirit of what you wrote, but I would love to have a place where I can find only projects that aren't written in higher-level languages so I don't have to wade through Python, JavaScript, etc. slop that is of no use or interest to me. A HandMade network for source hosting, in a way.
The world is turning (has turned?) into a global village:
The multiplication of far-reaching techniques of communication has two important results. In the first place, it increases the sheer radius of communication, so that for certain purposes the whole civilized world is made the psychological equivalent of a primitive tribe.
...where new religions are being invented all the time and secularism is lost---we are not conducting our human affairs based on naturalistic considerations, but involved with our new religions.
I am incredibly sad, similar to when mitchellh was leaving GitHub, to leave Codeberg. Codeberg was gaining widespread attention and adoption, by people who got tired of GitHub shooting itself in the foot and also who wanted to support their digital commons (like me).
Their blog post sounds a lot more reasonable but as @tchebb pointed out, it's quite different from the Terms of Service change, which itself is more reasonably worded than its proposal: "ToU extension to prohibit LLM-extrusions" (ugh).
First LLM-extrusions, then cryptocurrencies. I would be surprised if they stop there, and they might as well adapt The JSON License: "Software shall be used for Good, not Evil."
From Ghostty to Redis, rsync to Linux kernel, more and more FOSS projects are taking advantage of various AI tools in a responsible manner and forcing them to stay on GitHub is just unfortunate and backward in my opinion. We gain nothing from this change and lose a ton of progress in software freedom and digital commons.
True, but sadly none of the other options (based on a cursory glance) do offer the same set of services (code forge, static web hosting, CI) that Codeberg does.
It's unfortunate to see Codeberg position itself against entropy, but it will at least be interesting to see how many of these decisions are walked back over the next decade.
Interesting to see people flagging this comment as 'troll' but I stand by every word. Codeberg has decided it wants nothing to do with a technology that's already transforming software development, and I have a hard time imagining how that transformation could ever stop in the next decade, let alone reverse. The decision to ban LLM usage on Codeberg is something the organization will have to contend with as the usefulness and adoption of such tools continues to grow. I would not be surprised to see these new policies relaxed in a few years.
Meanwhile, on the broader culture war front, the insanity of the anti-LLM vibes is hopefully reaching its zenith. The models are obviously useful, even the ones that can run on your own computer. Adoption is increasing. Continued progression of LLMs into software development, mixed with their du jour villainy, correlates pretty well with the increasing anger and desperation we are seeing in response against LLMs. How much more can they be opposed though? Many large projects already have AI policies and many of those are permissive, not outright bans.
I'm sure in the near-future I'll be impressed/surprised by a few more desperate attempts by various projects to reject LLMs. However, at the risk of trying to predict the future, I expect we will eventually look back on this insanity and cringe at how judgy and controlling the movement to ban LLMs has been.
I didn't flag your post, but this is troll language, imho.
I'm opposed to LLMs generally, but there's also an incredibly practical reason: there's no economic viability there. Every provider is offering their product as a loss leader, and that can't persist for much longer. The benefit of using LLMs at pennies per token is one thing. When it's suddenly dollars per token, usage will drop. See the recent MS copilot pricing change for an example of what's to come.
I was about to transfer my Abject project to Codeberg but can't now. I've been coding for 30 years and have several open source projects. Abject heavily uses LLMs and itself can generate programs because it has a coding agent inside it. The purpose of this project is to research if the Ask protocol works and empower everyday people, not just programmers and engineers, to create deeply personal software that solves their problems instead of being beholden to SaaS and yes, the benevolence of FLOSS software programmers like myself. The point of the Free Software Movement was to keep source code open for USERS so they can modify it for their purposes and not be constrained by the writers of the original software. Abject in this sense follows the spirit of the Free Software Movement because it goes further and empowers users not only to read, but create their own software without having to have a degree in computer science or spent years learning to code.
Software after all should be about empowering people to both think better and solve their problems.
Now I can't do it. I understand they don't want LLMs to train on their data, but to prevent people from storing their projects there if they used LLMs for code means codeberg is dead to me as a free software advocate.
Codeberg seems to be more aligned with the OSS compared to the FL in FLOSS.
conartist6 | 17 hours ago
I'm like as anti-AI as you can get, but I'm not sure how they think this will play out. They basically said "slop is banned, but we can't go after all LLM use that would be absurd"
Without clear rules that can and will be enforced, it will just become a game of who like or dislikes who...
evert | 12 hours ago
I think people are a bit too focused on the enforcement aspect. I could be wrong, but I think the more important point is that it's an expression of their values and the type of community they wish to foster.
If you want to vibe code projects you can probably circumvent the rules and pretend you're in the community, but you might instead prefer to join a community that share your values.
mitsuhiko | 11 hours ago
It's their rules, and I won't tell them what to do, but I think the issue with "what type of community they wish to foster" is that this is ill defined. If the stance is "no LLM" that is very clearly defined. But beyond that, it's arbitrary and will lead to fights about it.
zetashift | 9 hours ago
I think at the scale of Codeberg, it's fine if they figure that out as they go, it's new social stuff for a lot of developers anyway. It's nice to know that there are options.
Potential ForgeFed improvements might also mean that things aren't going to be too crazy.
dvogel | 4 hours ago
Right, so read it as roughly "we can't define where the line should be precisely but we want to build a community that is interested in figuring out how close to maximum human interaction as is still practical".
intarga | 12 hours ago
The same can be said for the site guidelines here on Lobsters. The lines are blurry, a lot of offences slip through, and those the moderators catch are often caught only after most users that would see them already have. The moderators are also regularly accused of bias and favouritism.
All the same, I'm glad the guidelines exist. They set expectations and let the mods ban the worst offenders, such that the site is a much more pleasant place to be. I think that's what Codeberg is going for.
sjamaan | 12 hours ago
I think this is similar in intent to a "fair use policy". They probably just want a sign to tap when discovering projects that are overusing resources.
abnercoimbre | 17 hours ago
Didn’t they offer concrete examples at the end of the blog post? Those are the kinds of projects no longer welcome in the near future.
conartist6 | 16 hours ago
You don't think it will be acrimonious, the process of making accusations and then executing sentences? Even if all the first decisions are totally uncontroversial and target "the worst of the worst", how far do you take it?
alloyed | 16 hours ago
All moderation decisions are acrimonious: people don't like hearing that they've broken a rule, after all. I still prefer this over an attempt to make "clear rules", which i feel are actually more likely be gamed by people trying to find a way to be 1 step behind the line.
Honestly, though? I don't expect this to be too big a deal. Why would a vibe coder willingly choose codeberg, a hosting site actively hostile to them, over github, a site actively courting them?
conartist6 | 16 hours ago
I generally think rules should be designed to be played like a game. The alternative is lawfare -- the idea that everyone is always breaking the rules and you just choose who you want to weaponize them against, probably based on politics
tchebb | 16 hours ago
I'm concerned by how disconnected this blog post is from the text that's been added to Codeberg's terms of service.
The blog post makes several points about the costs that LLMs impose on society and ways that LLMs discourage collaboration. Of those points, I found this one to be the most convincing justification for banning LLM-generated code on Codeberg:
But then I followed the link in the very last sentence of the blog post and read what they've added to the ToS:
Those are entirely different points, and neither one is so much as mentioned by the blog post: the words "copyright" and "harm" do not appear in it. I suppose the copyright point is somewhat self evident, but the claim that LLM-generated code is more likely to be harmful is anything but.
I'm also surprised that the ToS entry seems not to have been proofread. It should say "Such projects have" instead of "Such projects having", "'generative AI'-tools" shouldn't be hyphenated, and it's inconsistent to write "OpenAI Codex" but not "Anthropic Claude". The text is fine as a first draft, but it makes me wary of Codeberg's governance that it went to a vote of over 400 people and got added to their terms of service without being revised.
pgeorgi | 12 hours ago
Maybe they should just phrase their ToS in German, matching their residence, to reduce grammatical complaints. (OTOH there is a defined-and-kinda-binding Standard German, while the same doesn't really exist for English 🤷)
technomancy | 4 hours ago
I've been a Codeberg member for several years; most of the things I've voted on have been in German with an English translation listed secondarily. Not sure why this one is different; maybe because it's the TOS vs the Codeberg eV bylaws?
louwers | 14 hours ago
What the authors of this post don't seem to realize is that it is possible to care about software while maintaining and developing it with the help of LLMs.
Case in point is MapLibre. We ported one project to Codeberg, but the maintainer makes heavy use of LLMs for development. So we have to close that repo again. I am also thinking of a project like Ghostty. Another serious project that wants to leave GitHub, but cannot move to Codeberg now.
This move by Codeberg is an effective ban on all LLM usage, because the rule is very vague and even the guidelines in this blog post are contradictory. No open source community with even a slightly liberal policy on LLM use is going to risk having their community booted off Codeberg.
pgeorgi | 12 hours ago
Since the rationale is "Many LLM-based projects use an extraordinary amount of resources, especially compared to the developers and users involved", a good rule of thumb seems to be that they'll kick you off if the "heavy use of LLMs for development" involves spawning hundreds of CI builds on their gear, or sending gigabytes of artifacts back and forth, in a way that is atypical for human-centric development.
Also, Codeberg is people: Getting the boot there isn't the automatic anti-abuse abuse we got used to with the IT megacorps. I'd expect "Uhm, your project might not be a good fit" as a discussion starter, potentially followed by an offboarding plan to get your project set up elsewhere plus a permanent limited presence on codeberg to forward folks.
louwers | 11 hours ago
I don't see that "rule of thumb" mentioned anywhere in the blog post. This rule targets all projects that use LLMs. Whether they use significant resources or not.
Which effectively means "Your project is not welcome on Codeberg". I understand that it's possible to self-host Forgejo. In understand that there are alternatives. But I had high hopes that Codeberg could be a welcoming community for various kinds of open source projects developed in good faith, with different stances on LLMs. I have not given up hope yet that Codeberg can be that. My member application was actually already pending.
pgeorgi | 10 hours ago
I derived that concept from
and
Given
it seems quite clear that they're looking for projects where people interact to create code (see also the paragraph titled "The development team of none").
"You and your AI monkey" is a bit of a mismatch to that goal of theirs (just like my solo-developer project, really), but they also don't seem to be interested in kicking off accounts as long as their presence doesn't make the system worse for everybody else.
krig | 10 hours ago
These are deliberate choices made as much by MapLibre and Ghostty as by Codeberg. Codeberg changing their policy would just be letting someone else impose their views on the Codeberg community. In the words of Michael Bolton; “Why should I change? He's the one who sucks.”
We’ll see how it plays out, I don’t agree that policies have to be completely closed to interpretation, I don’t think that is even possible. I am maintaining a fork of an LLM-heavy project on Codeberg which has a no-LLM policy in place, but still contains LLM-generated code inherited from the upstream project before the fork. My interpretation leads me to think we are still OK, but if not I’ll respect that and will have to move elsewhere.
viraptor | 12 hours ago
Also it's a weird idea of what a successful project is. Some good software will have a community. Some bad software will too. Some will have a single maintainer, no love, and barely anyone caring about it until it goes missing. We had the corejs drama already. Does tzdata have a "community" as such? I'm the only maintainer of a mostly vibecoded utility which allows some other projects (with communities) to compile - does that quality?
I can't force them to not be selective, but the line they try to draw is weird and feels anti-opensource to me.
martinald | 10 hours ago
Agreed. If it's about resources (CI use or storage) then just cap that for non-popular projects?
If it's about community, that seems to have little to know correlation to quality?
I think it's a real shame a lot of the open source world has came down extremely anti LLM. I have done so much tinkering with open source projects (and finally switched to desktop Linux, as it's trivial to fix annoying bugs and config options with an agent), yet what seems to be like a majority of open source projects have taken an aggressively anti LLM stance.
I do totally understand that a huge influx of AI slop is time consuming and the mostly volunteer community can do what they want - but it seems to me the right thing to do would sit down and work on better tooling (for example, auto rejecting pull requests from people with no open source "track record" on your project, with some way for people to bootstrap it) to help shift through it.
sjamaan | 8 hours ago
For a lot of us, open source is where we go to unwind and tinker and mess around in our spare time. We don't want to deal with any corporate bullshit. Pounding tickets and uninspiring work should be squarely left at
$DAYJOB. We're in it for the fun and intellectual challenge. We like building things with our own hands. Suddenly loads of people turn up and ruin the fun by submitting slop PRs, and hammer our services with scrapers, turning our fun side project into an admin chore. What's the point then?martinald | 7 hours ago
I totally get that. But for many people "new(er)" to open source, they are getting the fun and intellectual challenge from building and fixing stuff with LLMs.
I've fixed so many problems in Linux locally that were fun to have an agent fix (for example, reverse engineering the new Android Quick Share protocol which was broken on Pixel devices for sharing files to my Linux desktop). It was a fun evening testing it and working with the agent to figure it all out. I don't have the time to do a multiday reverse engineering project like that by hand, and now something that was bugging me daily works on perfectly on my desktop.
Is this LLM "slop"? Yes certainly. Is it valuable to me and others? I'd say yes, it took time to test with a real device that someone doesn't need to spend now. How do we square this circle?
But of course, I understand the admin problem, but my view is that the platforms need to solve this for open source project maintainers, like dealing with scrape protection and building in automoderation tools that reject nonsense contributions, while still giving some sort of path .
Obviously everyone's opinions are different on this and I'm not trying to tell people what to do. But it seems to me that open source has went from "we don't have enough contributors" to "we have too many contributions, ban them all" when I think there has to be a middle ground of some sort.
bendmorris | 5 hours ago
Regardless of how much fun you're personally having with LLMs, the average quality of those contributions dropped to negative. Previously, someone who showed up to contribute to a project showed drive and curiosity, and they were likely to try to learn more and stick around. Some of these people would eventually become maintainers - this is how open source projects grow and survive, and was also an on ramp to programming for many of us. Now, the "contributor" is likely to be a bot with little to no human involvement and a very low likelihood of any involvement after.
What a massive loss for all of us this has been, and LLMs are the cause.
Unfortunately some projects are banning contributions entirely - it is unfortunate but I also understand the decision practically. Others are just trying to get back to their previous signal to noise ratio.
conartist6 | 6 hours ago
Gosh it would be bad if there were still just one person maintaining CoreJS. Glad we dodged that bullet! https://github.com/zloirock/core-js/commits/master/
dbushell | 15 hours ago
Makes a lot of sense. GitHub is already the dumping ground for slop, no need to turn Codeberg into a second landfill.
mtlynch | 7 hours ago
Ugh, this is frustrating.
I've been a paying member of Codeberg for 2 years. I didn't even realize this vote was happening. I just checked my email and found this line buried in the middle of an 1,800-word email.
This was the only announcement I received that this vote was even happening. The email says there will be a ballot, but I never received one. I'm not sure if I'm an anomaly or if many other members did not receive notice about this vote.
jmelesky | 6 hours ago
When you became a paying member, you should have had the option of becoming an "active" member or a "supporting" member. If you're an active member, you should have received the invite to the annual assembly where all the proposals were discussed and debated, and then links to vote on each of the changes. If you're a supporting member, though, you don't get an invite or a vote (or a bazillion emails discussing pull requests to the org docs).
You should be able to change your membership to active whenever if you'd like.
If not, you can also keep an eye on the pull requests to the org documents: https://codeberg.org/Codeberg/org to keep up to date on possible changes.
mtlynch | 5 hours ago
I believe I'm an active member. I was invited to the annual assembly, although it's possible I switched to supporting at some point because I got exhausted by all the noisy changes to the org docs.
stephank | 6 hours ago
I believe when you sign up, they ask you if you want to be a supporting member only, or also an active member with voting rights.
pie_flavor | 14 hours ago
They've also declared that they're blocking cryptocurrency related projects. They have explicitly said "this is not a neutral space, Codeberg is political" on the subject.
Not to concern troll, my day job is in crypto. But still. This is your backup service, to save your work in case your computer fails; this is your public identity as an open source community project. Why would you select, for these tasks, a service that advertises that tomorrow it may not like your face, due to the shifting winds du jour, and your backups and identity are gone? You want your infra to be neutral; that is the job of infra.
sjamaan | 12 hours ago
I see your argument, but my gut says no. I actually like that they have the balls to apply some ethics to what they're doing. Ethics seems sorely missing from computing.
Whether you agree with their particular values or not, perhaps you can agree that the modern enshittified ad-infested user-hostile tech hellscape is a direct result of organisations not publicly stating any values and standing by them. Google silently dropping their "don't be evil" motto is a culturally significant symbolic milestone IMO (even if they had already jumped the shark long before).
pie_flavor | 11 hours ago
I don't agree with that, no. Enshittification is a business strategy that usually implies being owned by someone who would prefer you not have values, but I don't think the causality runs in the other direction, where values prevent enshittification (the abstract concept of values, obviously specific engineering values contradict it but those aren't the ones Codeberg is expressing). I also don't think GitHub (the obvious comparison here rather than Google) qualifies: the core service, despite the constant outages, is basically unchanged and just as free. If you're migrating from GitHub to a neoforge, it's probably related either to outages, Microsoft's existence, or the aesthetic of Copilot "intrusion" (which I suppose is a better counterargument to what I said: this is their audience).
sjamaan | 11 hours ago
I mean, obviously specific values will not prevent enshittification (as an extreme example: profit and growth overriding all other concerns). So not having any clearly stated values should be a huge red flag, as you have no idea where they stand. But you could counter that talk is cheap anyway, and companies have been known to completely reverse course.
Personally I've disliked GitHub for a longer time than the Microsoft takeover: their gamification (with badges, stars and whatnot) and focus on "social coding" was a huge turnoff for me. In retrospect it's an obvious (but arguably failed) attempt to make the site more "sticky" and addictive. We can guess what values underlie those changes. Also, their "macho bro culture" was well-known even back in the day when they just got started. I distinctly remember a picture of Scott Chacon wearing sunglasses, walking towards the camera with explosions in the background, whisky glass in hand.
So yes, the service is essentially unchanged but I'd argue many people haven't had a close look at GitHub's apparent values in the first place and just went with "the default".
The values of Codeberg specifically stated here don't prevent enshittification, for sure. But the fact they are a collectively owned nonprofit at least rules out the typical reason for enshittification (profit, duh) and makes it harder for a single "owner" to make unpopular decisions. Of course there are (AFAIK unstated) values underlying the decision to set up the organization like that.
prayerie | 12 hours ago
nah, learning that just made me like codeberg even more
mxey | 10 hours ago
Codeberg is not a backup service.
kryptiskt | 9 hours ago
I don't get this at all. If you're in cryptocurrency you can pay for your own hosting instead of leeching off a non-profit free forge. Like, you're in it to make a buck, taking resources from idealists is really low.
Halkcyon | 5 hours ago
They work in crypto, what can you expect?
mariusor | 12 hours ago
We must all make ethical choices every day. Having people disagreeing with yours is not a case of "tomorrow it may not like your face" (sic). The fact that you say that tells me that you perhaps haven't evaluated enough the ethics of make believe currencies to understand why others might not feel the same way you do.
Al3xFor | 7 hours ago
I select them exactly because they are not afraid to take a stand AND they let me freely use the software which runs that very service ( that they have all the rights to manage as the community sees fit).
I don't expect Codeberg or anyone else to be my "public identity as an open source community project", I thank them for giving me the freedom to do that for myself: https://lobste.rs/s/jhcyuq/deep_dive_into_my_forgejo_setup
louwers | 11 hours ago
Wow this is quite shocking already. This can literally mean anything. If you have an anti-fascist project that alerts communities of ICE raids hosted on Codeberg, it could be seen as "harming the reputation of Codeberg", because the MAGA movement will not like it.
If your physics library starts being used by some war company, boom, your project can now deemed to be harming the reputation of Codeberg.
I really think Codeberg should see the ToS as a sort of constitution, requiring a three-fourths majority to change, because otherwise they will ban more and more content that half of the members don't care about or have prejudices against.
jmelesky | 6 hours ago
The LLM change passed 70/28.
The cryptocurrency change passed 62/31. Also, the change itself didn't add "harm the reputation", it added the cryptocurrency clarification. You can see the pull request here: https://codeberg.org/Codeberg/org/commit/f383b648fd733ab88083dc390b91e77cd7589980
rau | 18 hours ago
The project has a right to do whatever it wishes.
I suspect that forbidding projects which use LLMs to produce code will seem as strange in twenty years as forbidding projects which use high-level languages strikes us all now.
Back when compilers were invented, did folks point out that a compiler uses more energy than does an assembler, and an optimising compiler uses even more? Did they complain that using a high-level language enabled a single programmer to produce what used to take a team? Did they complain about a ‘digital divide’ between those who could afford procuring and running high-level languages and those stuck with machine code?
At a different inflection point for technology, I do recall that a lot of folks spent a lot of time believing that they could manage memory better than a garbage collector (pretty sure I did!), and that automated bounds-checking was too costly for performance (guilty!), and so forth. But even back then I would have been shocked if, say, SourceForge had refused to host projects in Lisp or Haskell.
It’s just so weird that a technology comes out which empowers programmers (which most of us here are) and software users (which almost everyone on Earth is) to such a frankly amazing extent, and to see folks so profoundly negative about it.
hoistbypetard | 18 hours ago
No. And there were also no allegations of human rights violations committed during the development of, nor enabled by compilers or higher level languages.
Nobody complained of a digital divide because compilers didn't drive up the costs of computing equipment due to an investment bubble. Nor did compiler developers abusively spider information services, making it more difficult and expensive for smaller organizations/individuals to keep their services online for everyone.
LLMs pose a qualitatively different set of issues than the analogies you suggest. It would be more constructive to engage those issues on their own terms, without strained analogies. The LLM industrial complex is unlike compilers, and unlike higher level languages, in every line of criticism explored in that blog post.
It's so weird that their proponents keep trying to emphasize what similarities there are rather than addressing the criticisms directly.
simonw | 16 hours ago
Do Amnesty International have a position on why scraping the web to build a search engine isn't incompatible with International human rights law?
alloyed | 15 hours ago
Well, GDPR guarantees the right to data erasure, and google will in limited circumstances let you have a "right to be forgotten", as in, they'll intentionally de-index pages on behalf to maintain your right to privacy. Can't do that with an LLM without totally retraining it (i think? not an expert). This is the only point in the doc (section 6.1) that seems directly related to normal search engines as an abstract concept.
That said, you can find docs on their website about google, predating LLMs, and equally here they view Google (as a broad organization that includes ads and adtech) as not practically respecting the right to privacy, either: https://www.amnesty.org/en/documents/pol30/1404/2019/en/
simonw | 15 hours ago
Yeah, the complete inability to remove data that has been trained into an LLM's weights is a good answer to the difference.
mitsuhiko | 13 hours ago
Unfortunately it’s really an argument against open weight models more than anything. Right to be forgotten is implemented via filtering on the way out on search engines, not by data deletion.
I’m afraid people might be reaching for the GDPR gun and end up potentially in a worse situation where open weight models are restricted.
tmcb | 8 hours ago
If I’m not mistaken, the GDPR requires data holders to delete the data. Search engines instead should not list data that is not hosted by them.
So, if a LLM provider implements it by filtering alone, it is probably in breach of the GDPR (IANAL/TINLA/etc.). It makes no difference if the model is open or not.
Also, if open model maintains show a good track record of complying to the GDPR in that regard, that is even better than fake compliance by obscurity I believe.
mitsuhiko | 6 hours ago
The GDPR does not really regulate data storage but processing. You just need to stop processing that data. In practice a data deletion request is implemented as writing a row somewhere that data was deleted and will no longer be processed.
In particular that is needed because you cannot delete those rows from backups.
In theory the text says erasure but in practice we know what data protection agencies are okay with and that is just to ensure that production systems are seeing that data.
natkr | 5 hours ago
How would output (or input) filtering prevent information from leaking through second-order queries? "What is the former creator of Flask up to these days?" doesn't need to mention "mitsuhiko" or "Armin Ronacher" anywhere.
The core premise of LLMs is completely counter to the idea that they could ever meaningfully comply with basic privacy. So either they are incompliant or must be banned, or they are worthless and we'd lose nothing by banning them anyway.
mitsuhiko | an hour ago
It's unlikely that I could use GDPR to remove my Wikipedia page, so there are some limits there with regards to public information. But to answer the non strawman version of this question: labs are already doing this! Look into how refusal steering works. The "issue" is that on open weights models you have a lot of control over that and you can undo a lot of it. People have been doing this to get Chinese models to be less restrictive when it comes to certain topics that are trained on, but not politically acceptable in China.
sjamaan | 12 hours ago
Perhaps also the difficulty in understanding where information comes from. With a search engine, that's obvious, and you can go after both the search engine to de-list something and the originating web site to get the publication itself pulled. With LLMs, things pop up in the output and it can be impossible to tell where it came from.
Case in point: I've had "AI search results" which claimed something untrue (AFAICT) and pointed to sources which didn't actually contain these "facts". So either the LLM fabricated it out of whole cloth, or it appeared elsewhere in the training data. No way to know.
eldondev | 15 hours ago
I don't know, but I can hazard a guess at one line of reasoning. Suppose on a website, someone has posted a picture of Alice wearing mirrorshades. Alice doesn't like the picture of herself wearing mirrorshades, but perhaps the site/operator in question isn't kind and won't take it down. She can ask the search engines to avoid linking to that site, and most could easily remove it from results. It's then also usually trivial to check if mirrorshade Alice is on the search engine at that point.
It's much harder to know if a picture might end up somewhere embedded in the weights. It's not clear what the training data of many of these systems is, and it's not clear if you are somehow embedded in the weights. If someone asked for a picture of Alice, they might get regular Alice or, if the training data is minimal, they might only have mirrorshade Alice to offer. More subltely, if they're asked for "cyberpunk Alice", that may yield mirrorshades Alice instead of, for example, hoodie Alice.
Asking the AI company to redact mirrorshade Alice from a (possibly open-weight) model sounds pretty complicated. They might remove mirrorshade Alice from the training data on request of Alice, but it might be hard to know she's in there to begin with unless the company makes their training data public. They might try to proactively remove Alice from results yielded by AI companies, but the capabilities for that don't seem mature. (with a search engine, it is often easy to know via
site:style directives if something is included/linkable, and many index-based search engines can be intuited by humans).OTOH, some search engines might be incompatible with human rights laws (depending on what part of the web you're scraping and how).
joshka | 16 hours ago
I guess one way of looking at this is that there's a history of AI proponents seeing arguments they think are overblown or made in bad faith by the anti-AI crowd, and set against benefits that are obvious to them, it's easy to start downplaying even the stronger criticisms.
You could make similar arguments on both sides about organized religion. Abuses by organized religious institutions would not, by themselves, justify excluding every project influenced by a religious person. I'm not trying to argue that AI and religion are the same, but I am pointing out that in both cases a complicated argument about institutions, individual responsibility, and actual harms can drift away from reality and towards identity pretty easily even when both sides claim to be arguing about the "facts".
Even if the alleged harms are real, they do not automatically justify excluding a particular LLM-assisted project. You still have to connect that project to those harms and show that excluding it would actually help, and isn't disproportionate.
(For what it's worth, I'm actually kind of fine with Codeberg saying the direct economic burden is enough reason here. I just doubt that was the main thing driving the votes.)
mro | 13 hours ago
I expect the asbestos vibe in hindsight, but we will see. Asbestos was a magic wonder material before is became cancerous. We use less since.
tmcb | 12 hours ago
Most of the criticism against so-called “generative AI” has nothing to do with the technology itself but rather with the absolute insanity that has taken over the industry.
Just like people like myself have nothing against nuclear fission technology itself but strongly oppose to nuclear WMD.
spc476 | 16 hours ago
No, but in the 50s, they did complain that using a "compiler" was a waste of perfectly good CPU time back when you leased the mainframe from IBM or the seven dwarfs and you paid plenty in electrical costs to keep it running. Computers were cheaper than programmers at that point.
When I got into computers (mid-80s, as a teenager), the only language that was "free" was BASIC that came with the computer. Everything thing else, assemblers, compilers, linkers, applications, all cost money. I had to save money to get my first assembler. I paid for two C compilers (one for MS-DOS, one for the Amiga). And I only got the C compilers in the 90s, when compiler technology had advanced enough to make it generate "slightly less than horrible code".
Only for a certain class of programmer. There were programmers in the 70s that begged compiler writers to keep bounds-checking intact. Sadly, management drank the benchmark Kool-Aid and thus only now are we getting languages with bounds-checking.
I grudgingly agree with "software users"; I hate using PHP, but I don't begrudge anyone who uses it to get their vision implemented. I just don't want to have to use or maintain such a thing. And I'm less sold on "empowering programmers." Sure, some programmers have been empowered. I haven't.
pgeorgi | 10 hours ago
People pointed out that you're better off putting the source code into the code repository rather than the compiled result (and that also included heated debates about scripts generated by autoconf/automake, where the rough consensus these days seems to be that they belong in release tarballs but not in a repo).
So far I haven't seen repos that are useful while consisting only of prompts, so the comparison "compiler <-> LLM" seems to be stretched on that part alone.
dbremner | 7 hours ago
John von Neumann supposedly had similar views of assemblers and compilers.
gonz | 12 hours ago
This is not really matching the spirit of what you wrote, but I would love to have a place where I can find only projects that aren't written in higher-level languages so I don't have to wade through Python, JavaScript, etc. slop that is of no use or interest to me. A HandMade network for source hosting, in a way.
boramalper | 12 hours ago
The world is turning (has turned?) into a global village:
...where new religions are being invented all the time and secularism is lost---we are not conducting our human affairs based on naturalistic considerations, but involved with our new religions.
I am incredibly sad, similar to when mitchellh was leaving GitHub, to leave Codeberg. Codeberg was gaining widespread attention and adoption, by people who got tired of GitHub shooting itself in the foot and also who wanted to support their digital commons (like me).
Their blog post sounds a lot more reasonable but as @tchebb pointed out, it's quite different from the Terms of Service change, which itself is more reasonably worded than its proposal: "ToU extension to prohibit LLM-extrusions" (ugh).
First LLM-extrusions, then cryptocurrencies. I would be surprised if they stop there, and they might as well adapt The JSON License: "Software shall be used for Good, not Evil."
From Ghostty to Redis, rsync to Linux kernel, more and more FOSS projects are taking advantage of various AI tools in a responsible manner and forcing them to stay on GitHub is just unfortunate and backward in my opinion. We gain nothing from this change and lose a ton of progress in software freedom and digital commons.
pgeorgi | 12 hours ago
They actively promote other options, see https://docs.codeberg.org/getting-started/what-is-codeberg/#alternatives-to-codeberg, so I'm not sure if that's what they're up to, anyway.
boramalper | 11 hours ago
True, but sadly none of the other options (based on a cursory glance) do offer the same set of services (code forge, static web hosting, CI) that Codeberg does.
intarga | 10 hours ago
All four of the listed alternatives offer all the services you mention. Unsurprising, as three of them are the same software as Codeberg.
hoistbypetard | 5 hours ago
Interesting response from a user who migrated to codeberg rather recently:
https://マリウス.com/i-regret-migrating-to-codeberg/
(PSA: if you have javascript enabled, that site is rather aggressive about encouraging you not to browse arbitrary sites with it enabled.)
hungariantoast | 14 hours ago
It's unfortunate to see Codeberg position itself against entropy, but it will at least be interesting to see how many of these decisions are walked back over the next decade.
hungariantoast | 4 hours ago
Interesting to see people flagging this comment as 'troll' but I stand by every word. Codeberg has decided it wants nothing to do with a technology that's already transforming software development, and I have a hard time imagining how that transformation could ever stop in the next decade, let alone reverse. The decision to ban LLM usage on Codeberg is something the organization will have to contend with as the usefulness and adoption of such tools continues to grow. I would not be surprised to see these new policies relaxed in a few years.
Meanwhile, on the broader culture war front, the insanity of the anti-LLM vibes is hopefully reaching its zenith. The models are obviously useful, even the ones that can run on your own computer. Adoption is increasing. Continued progression of LLMs into software development, mixed with their du jour villainy, correlates pretty well with the increasing anger and desperation we are seeing in response against LLMs. How much more can they be opposed though? Many large projects already have AI policies and many of those are permissive, not outright bans.
I'm sure in the near-future I'll be impressed/surprised by a few more desperate attempts by various projects to reject LLMs. However, at the risk of trying to predict the future, I expect we will eventually look back on this insanity and cringe at how judgy and controlling the movement to ban LLMs has been.
jmelesky | an hour ago
I didn't flag your post, but this is troll language, imho.
I'm opposed to LLMs generally, but there's also an incredibly practical reason: there's no economic viability there. Every provider is offering their product as a loss leader, and that can't persist for much longer. The benefit of using LLMs at pennies per token is one thing. When it's suddenly dollars per token, usage will drop. See the recent MS copilot pricing change for an example of what's to come.
Of course, maybe I'm just insane with my vibes.
mempko | an hour ago
I was about to transfer my Abject project to Codeberg but can't now. I've been coding for 30 years and have several open source projects. Abject heavily uses LLMs and itself can generate programs because it has a coding agent inside it. The purpose of this project is to research if the Ask protocol works and empower everyday people, not just programmers and engineers, to create deeply personal software that solves their problems instead of being beholden to SaaS and yes, the benevolence of FLOSS software programmers like myself. The point of the Free Software Movement was to keep source code open for USERS so they can modify it for their purposes and not be constrained by the writers of the original software. Abject in this sense follows the spirit of the Free Software Movement because it goes further and empowers users not only to read, but create their own software without having to have a degree in computer science or spent years learning to code.
Software after all should be about empowering people to both think better and solve their problems.
Now I can't do it. I understand they don't want LLMs to train on their data, but to prevent people from storing their projects there if they used LLMs for code means codeberg is dead to me as a free software advocate.
Codeberg seems to be more aligned with the OSS compared to the FL in FLOSS.
oceanhaiyang | 12 hours ago
Can I still upload handmade slop?