I joined Lemmy back in 2020 and have been using it as @qaz@lemmy.ml until somewhere in 2023 when I switched to lemmy.world. I’m interested in systemd/Linux, FOSS, and Selfhosting.

  • 175 Posts
  • 90 Comments
Joined 3 years ago
cake
Cake day: June 10th, 2023

help-circle




  • I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.

    Surely there has to be a better way to do this (probably the point).

    Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?

    EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership

    EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?






  • Too many people gloss over the fact how insane it is that some people genuinely believe non-believers burn forever in a magic post death torture world. Like how can someone think that while talking with someone, and be like, aw sucks for them, glad I submit to this ancient chain email creepypasta.

    It still baffles me when I realize someone thinks like that.









  • Deepseek recently published a paper in which they describe that vision tokens contain more information than text tokens and that this can be used to compress context.

    We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.

    Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10×), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20×, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.

    It reminds me of LLM caveman speak, it used to have another option to use Chinese instead of English. A language like Chinese is seemingly better at encoding information in fewer tokens and I think this is the same mechanism why OCR tokens work so well.

    That said, I also doubt that voice messages are more efficient than text prompts, but it’s best not to waste too much time engaging with these sorts of LinkedIn posts (and LinkedIn in general).