- cross-posted to:
- hackernews@lemmy.bestiver.se
- cross-posted to:
- hackernews@lemmy.bestiver.se
I’m perfectly happy with this.
Why?
Because at the end of the day, the users of the code forge can be broken up into three categories:
- Traditional coders with human made projects
- Vibe-coders who slop out 100 throw away projects, each of which is a dead project as soon as that user moves onto the next project
- Vibe-coders who focus work on only a single project.
Category 1:
- Code forges were already valuable before LLMs, so we already have a proof of value/utility/whatever for hosting projects from users in category 1.
Category 2:
- Projects from the users in category 2 have a half life of a couple of weeks.
- They’re so quickly made that most of their contents is going to be rehashes of previous existing content, so there’s a reduced value holding on to them as reference when the original sources exist.
- Since they were generated so quickly, a more up-to-date version can be generated in the future if they’re ever needed again.
- Since the projects aren’t going to be used by others, they don’t actually use the features of a code forge, so it’s not a great idea to allocate so many resources that aren’t going to be utilised to host them on a community code forge. A better place for them would be cloud storage (dropbox, google drive, etc…). Most cloud storage services have a free tier. Anyone with a crazy amount of projects and assets that would exceed free tiers and who just wants to hold onto the projects for the future can use something like S3’s glacial deep archive which is about $1.20USD a year to hold onto 100GB of zipped up projects.
Category 3:
- These projects may actually get used by people
- Since they’re developed with LLMs, there’s so much churn that there’s no actual point of reading the code since it’s quickly out of date. That means there’s no use for using the code forge’s source file features
- Issues that reference code will go out of date as soon as the code goes out of date which is far more likely than with category 1 users.
- These projects tend to have so much code churn that their builds are more inefficient than hand coded projects. This means that they’ll tend to want more powerful hardware to run their CI/CD pipelines than a community code forge can provide. So from their perspective, they’re better off running their own code forge anyway.
Keeping category 1 users on community code forges, saving the code forge’s resources from being used by community 2 users, and letting community 3 users tailor their own code forges for their own needs seems like a win for everyone actually using the community code forge, and not those treating it as a file dump.
Hot take: I don’t like this. It states a rule, but the line is vague.
How is vibe coding detected? How much is too much? How does this apply to my personal vibe coded util™? Does this apply to private repos?The rule feels ill defined while the consequences are quite extreme.
I think in practice they are saying that any shit that happens cant be blamed on anything but the submitter.
You can already make a project out of copypasted bits of code. So long as you understand what they do, vetted and vouch for it…it’s fine, no?
The change feels like a slippery slope where it could be easily abused if there’s enough incentive. Like the DMCA (I think that’s the acronym) process on youtube or github, or etc.
Maybe you’re right, time will tell.
Some good questions in the comments:
What does “written by” mean? Who defines authorship? What does “mostly” mean? Who defines ratio? Is this going to be enforced in practice? Who is going to enforce it? Based on which reporting?
While there are cases, when the garbage nature of genAI code wouldn’t cause that much of an issue, the issue will be that these projects ultimately will drive attention away from human-made code, which can have devastating effects on said projects.
I don’t really understand that line of reasoning. By definition if genai code is garbage, a human made competitor would be better. Why can’t it stand on its own qualities rather than rely on blanket banning the alternative?
I think it’s a matter of volume, it’s extremely easy and quick to publish large amounts of generated code, and by far most people won’t go searching your codebase in depth to see if the code is decent, if it works as advertised they don’t have a reason to look into it further. If you add a third party to vet all code then you’re still giving massive amounts of work to whoever has to review the mountains of slop.
It’s not like we didn’t have an issue like this before, about people copy pasting without understanding anything, but these tools made it more accessible and streamlined, and the worst part is that vibe coders believe sharing their slop is of value.
Yeah, this makes sense. They won’t be able to kick every vibecoded project off, but it gives them a rule to point to when they do find and remove slop.
Some of the commenters on the Codeberg issue are expressing concerns about copyright status, but my main concern is security issues with these projects, which I have expressed before. Somebody vibecoded a script to sort files? I don’t really care.
Someone vibecoded a network facing application, shipping a docker container, and is trying to encourage people to deploy it? If I investigate, and see the bad security practices common to vibecoding, I wouldn’t want to be providing that for people to download. Even though I’m technically not responsible for it, it just doesn’t feel right to host something that could explode later. I would want to be able to take it down, and having a rule to point to prevents complaints when Codeberg does so.
They won’t be able to kick every vibecoded project off, but it gives them a rule to point to when they do find and remove slop.
Exactly. This eliminates discussions over and over again with developers, maintainers and users, for every application, if and why they remove it. And it makes it clear and signalizes to such devs before they upload.
while I agree with the sentiment. what if someone just sloppily hand-coded an insecure docker container, should you be able to kick that off too? I fear security isn’t a good baseline to decide if a project should be hosted because security is a gradient and it becomes very subjective to say what constitutes secure. following that, would you give remediation opportunities for hand-coded insecure projects? It’s a can of worms.
I would argue that they can argue it purely from a capacity perspective.
They can estimate the scale of handwritten code growing over a period of time, but not the scale of AI written code scaling, meaning their platform cannot anticipate the scale and therefore it is unsustainable. And at that point the scale of security awareness as a whole can be considered, e.g. “how many attack surfaces, e.g. repos, are we hosting?” vs “how many of our repos are insecure?”
what if someone just sloppily hand-coded an insecure docker container, should you be able to kick that off too?
I can talk to them and have them fix it without them being weird about it. But if they don’t understand the poor design decisions in the first place then I have no confidence won’t be able to actually fix it, and keep it fixed in the future.
It’s fuzzy, and many generalizations are gonna be made.
Ultimately, it really comes down to human judgement and not your instance not your rules. They are hosting for free. They don’t have any obligation to host your projects. The rules and everything are nice conventions, and a good attempt at transparency, but at the end of the day, there is someone with access to the admin panel who is going to make the decisions, and you’re not really going to be able to do anything about them.
It’s the same thing with lemmy tbh, although it’s less annoying because it’s much easier to self host and migrate to Forgejo, as opposed to hosting Lemmy. That’s why I phrased it as “I would feel uncomfortable hosting this content” rather than trying to address the fuzzier arguments about the necessity of human judgement, the difficulty of detecting LLM generated code, or the possibility of incorrectness.
Did Codeberg even have a slop repo problem?
I can’t imagine a person who chooses to use Codeberg will be bad at utilizing LLM’s in any case, especially since SlopHub is right there for the worst sloppers to publish in and maybe win some fake points.
This just feels like some people couldn’t stop themselves from doing some posturing!
Seems reasonable but good fucking luck enforcing that.
IMO they should instead simply charge cost price for hosting. It can’t be much surely?
I don’t think the cost of hosting the repositories itself is the issue, more the legal problems and the fact that real projects get drowned out by the noise
I applaud this at least in theory. As others have pointed out this may prove thorny to implement. But, I love the sentiment and hope it goes well for them.
How the hell does a user see a voting like this? I’m using Codeberg for quite a while and reading here something like “vote closed” is weird to me.
Now I’m not talking about the content of the core itself here.
Are you a member of Codeberg e.V.? Votings are exclusive to members AFAIK, and get announced via mail.
Correct. Membership to the nonprofit Codeberg e.V. is 24 Euro a year (potentially less if you request a discounted rate and are approved). You will be supporting Codeberg’s mission and can optionally take part in these discussions and votes. There is an organization on Codeberg you’ll be added to where discussions occur and then vote notifications are sent out by email.
See also: what is codeberg e.V.?
deleted by creator
This feels like a dangerous slippery slope.
Really easy to abuse to just ban projects for other reasons.
deleted by creator
Considering AI generated code is public domain, so not exactly open source licensed, this is a good move.
Considering AI generated code is public domain
Is it though? There is no law that says that, so its just an interpretation of someone. I argue Ai generated code CAN be public domain if it does not contain any licensed code. Therefore we cannot assume the code being public domain without checking.
It is.
https://sciactive.com/human-contribution-policy/#More-Information
Edit: since apparently no one actually follows this link, this is a link to a series of quotes that come directly from the United States Copyright Office that specifically apply to the question of whether AI generated code can be copyrighted. This is not an opinion piece, it is a series of direct quotes, which I share below.
In the Office’s view, it is well-established that copyright can protect only material that is the product of human creativity. Most fundamentally, the term “author,” which is used in both the Constitution and the Copyright Act, excludes non-humans.
– II. The Human Authorship Requirement
If a work’s traditional elements of authorship were produced by a machine, the work lacks human authorship and the Office will not register it. For example, when an AI technology receives solely a prompt from a human and produces complex written, visual, or musical works in response, the “traditional elements of authorship” are determined and executed by the technology—not the human user. Based on the Office’s understanding of the generative AI technologies currently available, users do not exercise ultimate creative control over how such systems interpret prompts and generate material.
– III. The Office’s Application of the Human Authorship Requirement
Some technologies allow users to provide iterative “feedback” by providing additional prompts to the machine. For example, the user may instruct the AI to revise the generated text to mention a topic or emphasize a particular point. While such instructions may give a user greater influence over the output, the AI technology is what determines how to implement those additional instructions.
– Footnote 30
(All from this document from the US Copyright Office: https://www.copyright.gov/ai/ai_policy_guidance.pdf)
This is just an interpretation of one organization, not an universal law.
Ai is trained on different licensed code. It does not automatically become public domain just because the Ai processes it. There is also no guarantee that the output is free of licensed code that already exists.
It is the interpretation of the Judicial Branch of the United States government and the United States Copyright Office, the best authority on copyright law in the United States.
But yes, much like every single other law in existence, it’s not universal. (I mean, unless we’re also talking about the laws of physics.)
And yes, if the AI generated code is a copy of copyrighted code, then it would still be copyrighted, but that’s an even worse problem for Codeberg than it being public domain.
It is the interpretation of the Judicial Branch of the United States government and the United States Copyright Office, the best authority on copyright law in the United States.
Codeberg is based in Germany where German copyright law applies, and German copyright law is quite different from USA copyright law.
They still need to consider all copyright laws, and something being considered public domain in the US is a pretty big deal with regard to open source licensing. For example, if code is in the public domain, and is “released” under the AGPL, I don’t actually have to follow the terms of the AGPL when I use it.
The quoted laws do not say that generated code is automatically public domain, that is an interpretation of the law by some organization (here ScieActive). The laws just say, that a person using a prompt cannot take ownership and copyright of the generated code. It does not state it becomes public domain for everyone. Besides that, this is only in the US, not universal. And its not even tested in court yet. Its like saying in Brazil (or the EU in example) exist a law that does not allow Ai, therefore its the law for everyone. This is not universal.
The quotes are not laws, and those quotes are from the Copyright Office’s official statement on AI generated material. Please, just click the link at the bottom of my earlier comment. Here, I’ll even link it again:
https://www.copyright.gov/ai/ai_policy_guidance.pdf
If something cannot be copyrighted, it is in the public domain.
It doesn’t really matter if it’s not international law. If I don’t want to follow your open source license requirements, and your project is public domain in the US, I’ll just copy it in the US, and you can’t sue me.
Ok, I’ve editing my earlier comment to explain what I’m linking there, since it seems that neither you nor anyone else bothered to click that link.
Should users who don’t self-host start recording themselves while coding just as a precaution, in case their host will ask for evidence of work authenticity?
Let’s discourage user transparency and AI disclosure
OMG that’s a great move.
You can’t just make murder illegal! All you’ll do is make people lie about doing it!
THIS.
All my murderer friends have moved their activities to jurisdictions where murder is allowed or tolerated.
I think CodeBerg should push for giving Vibe Coders the death penalty, to make sure deterrence will work.










