studies show a clear trend – output is up (more code, more commits, bigger diffs), but outcomes don’t reflect that trend. If anything, the average team is taking longer to ship worse software
There are some cardinal sins within the annals of software development in the workplace. The relevant two here are:
- Do not build a scoreboard for productivity
- Do not equate keystrokes with effort or value rendered
These create perverse incentives that can really screw everything up, including creating permanent damage to company culture, products, and productivity. Anytime you see these things done, it’s because you have (or are) lazy-ass management.
output is up (more code, more commits, bigger diffs)
We’ve known measuring output by lines of code is counterproductive for a long time.
Can you retell my company that?
Has management known that?
A lot of problems seem to be downstream from “management are idiots and assholes”
I’m sure there were pockets but it legitimately seemed like that dragon had been slain until recently. Trying to assign more meaning to scrum points has been in vogue for a while though.
My team assigns both hours and points to tasks. I’ve never seen the points used for anything but they still spend time on it.
One study found a significant correlation between confidence in AI output and belief in the paranormal.
💀 💀 💀
That’s HILLARIOUS
Lol that’s funny. My belief in having a divinely created soul is exactly why I think humans can’t be replaced by these supercharged drunken parrots. :O
…But yeah this is probably referring to the ones who, under the right conditions, will start to sense “ghosts in the machine” and go down the rabbit hole to chatbot psychosis…
I am 100% genuinely not trying to start a fight, this a legitimate question to which I really would like to know the answer.
What is the difference between a divine creation and the "ghost in the machine "?
I will give context , faith is genuinely confusing to me.
I have no issue with personal faiths unless you are trying to force everyone to have the same faith as you.
I consider anything above small scale organisation of religion to be the worst thing that can happen to faith, because people are people and power corrupts.
I suppose my real question is , for a concept that prohibits proof as part of it’s definition, how do you determine that one faith (or system) is better than another?
This doesn’t seem to cover there is also no LLM that doesn’t plagiarize, or where the training data appears to be compatible with such behavior (e.g. CC0). Now I don’t know what that means legally, but morally it seems to be tossing away other project’s licensing and I think for FOSS as a whole that’s no good.
Also something worth reiterating: https://machinelearning.apple.com/research/illusion-of-thinking LLMs apparently can’t do basic logical reasoning. Even a junior coder can do that. I’m always surprised anybody would let LLMs near their code, at all.
I recommed you reading this
It summarizes really good not only the moral, but also the legal problems of AI, vibecoding and “AI-assisted/AI-boosted” programming/engineering/development.
Some of this does not line up with my lived experience pretty starkly.
Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.
Not perfectly well, but llms are designed specifically NOT to be perfect deterministic executioners. Still though, pretty well.
I have seen that in a jr engineers hands llms get to bad outcomes fast, and unintuitively (to leaders…) usage of llms in coding does not provide a path for a he engineer to upskill into a sr engineer. A sr engineer with llms though is almost always radically augmented regarding their output speed on task completion.
Repo level markdown files with architectural guidance not working for example
I’ve seen it become less and less effective as the size of the file(s) grew and as the codebase grew - they got increasingly more diluted or even lost in context compression. After several months of a 6 man team working on the project the rate at which they got ignored started affecting output a lot.
I agree with your last paragraph. We had about 6 weeks of unlimited AI spend before the costs reached executive leadership, and in that time I saw the least experienced developers spend the most with the least to show for it.
But I will say that another factor is thinking that if you get 10% gains from a little AI, then a lot of AI will get you 100%.
But I find the article is right about repo-wide docs. At least on their own. I find having small markdowns (often in the form of skills/commands), focused on specific tasks reduces spend (especially when your execution agent is a low cost model, leaving the reasoning to dedicated agents) and gives better outcomes. Loading massive docs into every task reduces the attention to the task at hand and often confuses AI as the reasoning part of the model becomes overwhelmed and starts inferring wrong things confidently.
I suppose it heavily depends on the scale of the repo though. A large microservice with multiple upstream services it needs to call spends a lot tokens on API which is unnecessary for most tasks. And then it decides to use the wrong one… I have stories lol.
Great summary.
It’s obvious if you actually use software beyond the average literacy of a talking chimp. I use hundreds of apps across iphone, mac, and linux os’s. Literally none of them have noticeably increased in quality, stability, or feature-set beyond their average between 1-5 years ago.
Mac and iphone appear to have more bugs and shittier quality control than at any other point in the last decade.
Quality software is getting harder and harder to find thanks to all the slop-abandonware being promoted by slop-content and slop-SEO on slop-enshittified search engines.
I notice far more idiocracy-grade errors in digital content, cx, business processes, product listings, etc than ever before.
Weather forecasts have gone to dogshit in the last 2 years. Even same-day forecasts can shift on a dime unpredictably. It’s at the point where I check 3 apps. Until this year, I never had a time where I woke up to 0% chance of rain and sunny, then looked outside to see rain. Not a sun shower. A rainy day hour-long downpour.
Auto-generated subtitles are great for content that was never going to receive human attention, but they’re clearly being used to replace humans. At least once a week I notice a major contextual error that completely alters the perception of the line/scene, and there’s no way to submit corrections.
Art, culture, and knowledge are being actively corrupted, bastardized, and destroyed.
The future simultaneously sucks while being dumb as all fuck. Complete clown show run by the idiot criminals, rapists, pedophiles, scammers, and thieves.
Yes it blows my mind everyone can’t see all the historically perfectly fine software starting to crumble to dust
This aligns with my experience, largely. Of course it’s still my job to maximize LLM effectiveness within my organization. Which is a delicate balancing act to protect my teams from overeager executive leadership looking for huge gains.
My own summary is that AI can be an accelerator, but the harder you lean into it, the worse outcomes will be. No matter how much code is written, you still need actual human minds to understand it and they can only handle so much volume before getting overwhelmed.
Also, if AI gives you 20% productivity gains, but that 20% goes into playing with AI trying to get more, you haven’t really gained anything. Usage needs to be standardized rather than developers constantly negotiating with AI trying to coax out better outcomes.
This is a tale older than AI. Most of the AI productivity pushes I struggle to get adopted fail not because of AI bad or its too hard to do. They fail because of a broken CI/CD pipeline. They fail because some team thinks their process is sacred and unique.
human minds … can only handle so much volume before getting overwhelmed.
One might even consider flourishing employees as opposed to not-burned-out ones.
Lol. Who knew?
Hehehehe
LLMs struggle with negation. Telling them not to do something can often have the same effect as telling them to do it.
The future looks… unreliable.
Models may get more powerful, but not significantly more reliable. This it folks – work with what you’ve got!
We all know that true AGIs becoming smarter than humans seems inevitable, but that could be like a hundred years from now, if ever. What’s unclear is what will happen two years from now, involving matters having little to do with the technology & what it is capable of and instead more to do with the economy and what jobs will be available then.

AGI isn’t possible with current or near tech. Anyone who says otherwise is huffing paint or selling AI crap.
I literally said “if ever”, and also “seems” rather than “is”. I also never so much as implied current tech, with my comment about a hundred years from now.
You are reacting against what I never said.
Though I choose to upvote your comment anyway, since at least you said it rather than simply assumed it and moved on.
Wow, what a solid argument you’ve made. It definitely doesn’t reek of someone trying to hype themselves up in the face of an unknown threat.
We invented a next token predictor. That isn’t intelligence nor is it on the path to it, either. The word rocks aren’t any closer to real intellect than the math rocks were. You just think they are cause words are the things that humans use to communicate across time and space.
The only ones spouting nonsense that it is are those whose business models require AGI to be achievable within the next decade. But, ya know, wish in one hand and all that.
15 - 30 years of pain followed by the end of life as we know it those who adapt may thrive but will continue to be exploited.
Tbf to LLM manufacturers, the end of life as we know it was coming either way.







