Looks like Anthropic is currently doing a giant rugpull and they seem to be degrading their models over time while at the same time redacting the "thinking" responses to try and hide it.
This is what you get when you are dependent on proprietary saass products.
There's nothing you can do except complain and beg thatthey will revert the changes. Absolutely pathetic.
@theorytoe@shin.mugicha.club I believe this specific model is impossible to run on average hardware and without proprietary software such as CUDA, but at least doing it this way you are not dependent on SaaSS.
getting a decent coding model running on consumer grade hardware is kind of a pipe dream. Being able to buy compute from anybody I want is enough to make me happy, I'm not so much of a fosstard that I'm worrying about proprietary dependencies or whatever.
@theorytoe@shin.mugicha.club I personally don't think it is a pipe dream. Over long enough time hardware will get better and models will also get increasingly more efficient. Being able to buy compute from anybody I want is enough to make me happyYes, but what Anthropic is doing is not renting compute power. They decide how their models and how it behaves, you have no control over it. It's not like renting a VPS where you actually still decide what runs on that server, which would be much better. I'm not so much of a fosstard that I'm worrying about proprietary dependencies or whatever.Well I am a freetard, not a fosstard. And I think these thinks are important.
@nigger@detroitriotcity.com claude code is 100% LLM produced and therefore public domainThere is unfortunately no legal precedence for who owns the copyright on LLM code yet. But this could be a possibility though...
@SuperDicq claude code is 100% LLM produced and therefore public domain, which is open source and you can still modify and share code leaks if you do not respect the law
@lolitechengineer@loli.church@theorytoe@shin.mugicha.club Also with these "flash" models are extremely quaternized and with a very low amount of RAM (i.e. 10GB) you can not do agentic coding. As you will probably be limited to a very small context size.
Still works perfectly fine for a regular chat model though.
@lolitechengineer@loli.church@theorytoe@shin.mugicha.club I do wanna say I'm not an expert on AI crap at all outside of the fucking around I did with ollama on my RX 6700 XT in my gaming PC and some RAG stuff I implemented at my dayjob.
@SuperDicq@minidisc.tokyo@theorytoe@shin.mugicha.club yeah, I wouldn't really know. The most "agentic" thing I have mine do is when I'm researching something I'll have it do a web search, download a few pages it thinks satisfy my requirements, and summarizes them. Helps filter through the shit.
@theorytoe@SuperDicq I'd like to have a good enough model that can consume documentation of programming langs and knows of SO/blog output to help me code the stuff I do, but while explaining me the things, not really doing my work for me.
it's not just context, it's how smart the model is. Try the best Qwen 3.5 you can run locally, then the best one cloud providers offer. it's night and day
@theorytoe@shin.mugicha.club You mean the companies who have scraped the entire internet that are currently leading losses on 10 gorillion GB of VRAM perform better than my gamer GPU? No way!