They are more than a accessory. You wouldn't say someone who funds a contract killing is merely an accessory to murder. VCs should likewise be considered as principal offenders, unless there is proof that the company was not transparent with the VCs.
Yes they are quite good, but are not able to run on a 16GB RX 9070.
Quantized Qwen 3.8 Flash Next could maybe run eventually on that card with a highly optimized inference engine that dynamically caches the hottest layer experts. Even then you run into some hard limits.
It will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM.
If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI.
Have you tried FreeToken yourself? I was hoping to find some benchmarks on their github but took a quick pass at their research paper and it seems they're showing ~2x performance on qwen 3.6 35b when compared to llama.cpp - but llama.cpp is so sprawling and has so many options I find that a difficult comparison.
I haven't yet, though plan to when this model is released. The models that I've been daily driving (Gemma 4 26B, Qwen 3.8 27B) have fit nicely on my 3090. I think FreeToken only offers a perf increase for MoE models that you can't feasibly fit in VRAM.
Yeah I'm sometimes unsure how to get best perf out of llama.cpp, and honestly thought it already did what the FreeToken paper discusses, but from everything I've been able to find since llama.cpp has no dynamic expert cache for GPU. An RFC discusses adding such capability and there's impressive results some are claiming from a fork, but I had to stop reading the thread, reading all the LLM generated comments and summaries from people was making me dizzy.
> It is still slow, a lot slower than what you are used to with claude and co.
That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus.
In general we haven't reached the ceiling for what performance we can get out of consumer hardware. As evidence by FreeToken which hasn't even added MTP/speculative drafting support yet, which will add another boost.
> Then when it runs for 30 minutes for something claude needs 5, your device will get hot.
I doubt the timing differential here, but even still I run my 3090 pretty heavily with inference workloads and it stays cooler than when I use it for gaming.
> And even a used 3090 is apparently now between 1-2k.
Yeah I guess the price went up significantly in the last couple months, used to be hovering around 1k. 3090 isn't the only option though.
> That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus.
Yes, a lot of Qwen3.8 27B setups are actually quite snappy, as long as they fit 100% in VRAM. In my testing, I wouldn't go below 32GB of VRAM, though—you really want a 6-bit quant and 8-bit K/V quants minimum. I've seen too much weirdness out of 4-bit quants since Qwen3.8 shipped. I think it may be damaged more than 3.6 at similar levels of quantization?
If hyperscalers hadn't bought up almost all the fast RAM production for the next several years, 32GB of VRAM would be tolerably cheap—a lot by "home PC" standards, but not terrible by "professional tools" standards. Sadly, the RAM market is amazingly ugly right now.
> I doubt the timing differential here, but even still I run my 3090 pretty heavily with inference workloads and it stays cooler than when I use it for gaming.
Yeah, running inference on a laptop is likely to run quite hot. But in an ATX case with decent cooling, it's generally a lower load than gaming. One handy tip: Many Nvidia GPUs (and some from other manufacturers) support power limits. For example, limit a 5090 to 400W instead of 600W, and it will run much cooler. You might lose 11% off your tokens/sec (depending on the exact card).
I have 2 4090 and last time i played around with it, the context window killed it for me.
The normal LLMStudio stuff works great, but then i tried out anything with subagent things or parallel stuff and it trashed my cache and got super sluggish/slowish.
Qwen 3.8 27B is so far the closest I have run, though it suffers on speed compared to Gemma4 26B MoE model which I still use. Neither are going to match Opus or Sol though, but they can be as fast or faster depending on what you are using them for.
I don't use a harness, all the mainstream ones I have tried have tanked my productivity. I know that's not a common sentiment, but it's been my experience. None seem built for the way I work. I program mostly in my head first away from keyboard, then go type it out (faster than it would take to describe the solution to an LLM). Also a perfectionist who likes to learn, and tends to work on out of distribution problems. All I need is a simple chat interface for light research, quick small scoped prototypes, and generating simple scripts.
Not saying a harness is out of the question for me, just all I have seen and tested so far are not for me. Maybe if someone builds a more deterministic harness that doesn't rely on plain English skill files that bloat context and only sometimes do what you want.
True, but you can link them up over thunderbolt or Ethernet. If your goal is to run local LLMs, not all weights need to live on the same computer. You can segment the workload by layers and pass the activations along the lower bandwidth interconnect with a small perf penalty. Also you get double the CPU/GPU cores allowing for better multi user/agent performance.
I was thinking the same as you as far as price per value, it does make seance to get 2x of these things IMO, but what throughput hit would you see in linking versus one machine? latency does matter, and there must be a trade off no?
No that is not inconceivable. What is inconceivable is believing that group of individuals is not a small minority, and the question will filter everyone else out.
It would be a decent interview question if such a large portion of compensation did not come from equity.
Otherwise it's basically the same as asking how someone would feel if all the sudden they lost 50% or more of their wealth. No one feels okay about that, except maybe the delusional few.
The issue is that Anthropic is just one company and China doesn't care in the slightest - so it's an entirely moot point and almost narcissistic (on Anthropic's part) to believe that they can have any impact whatsoever on this situation.
This is why I don't expect Anthropic to stop, and so I don't expect the value of the company to go to zero.
But the question is conditional on them deciding to, which (to the extent you trust Anthropic leadership which they're probably also filtering for in interviews) is conditional on believing that a decision that sent their stock to 0 would have a real impact.
The irony is of course that China is going so fast because of Anthropic, by diluting their models and learning from them. would make for a fantastic Greek tragedy if we weren’t talking about something that impacts the entire world economy and societies
Have you ever actually seen the inside of most Chinese manufacturing operations? The lack of PPE is astonishing. The lack of safety interlocks, protective features, etc. on their equipment is equally incredible. Safety simply isn't a cultural goal.
Spend 5 minutes with DeepSeek, spend 5 minutes with Fable, see how many rejections you get with DeepSeek for the exact same prompts. Sure, if you engineer some nuclear/biological/WMD prompt, they might block it, but they won't reason to do so.
It seems like at every turn the word private is avoided and replaced with "non-public".
I understand that the system probably is sound, the data just isn't encrypted, which I wouldn't necessarily expect. Maybe it's completely clear for users familiar with AtProto, but for others it needs to be clearer on some things:
- Where exactly is the data is stored
- Who is the "space authority" and why should they be trusted
- Strategies where you can manually encrypt/decrypt the data put in spaces
And avoid sketchy sounding terms like "non-public".
I am not developing with the ATproto, but keep an eye on the specs. There is a lot brewing in that ecosystem so I understand your skepticism.
I think that avoiding the word “private” is absolutely the right move here, because they are not building for the individual privacy. That’s not what the spec is intending to unlock. Their goal is to enable the next set of apps, that need to enable sharing of content with a specific and controlled group of users.
The data is stored in a PDS. You can use the one that Bluesky PBC or any other service provider provides, or you can self-host.
Space Authority is a DID, per the link above. So, in ATProto land, that’s just an account. The authority DID controls the accesss to the the space for other DIDs.
I can’t answer your question about encryption. Hard to answer it without a specific use case.
reply