Hacker Newsnew | past | comments | ask | show | jobs | submit | alexchantavy's commentslogin

I like the concept of deontological ethics where instead of just plain good and bad, we consider what is one’s duty. People make decisions based on who they are responsible for.

First thing I thought of too, I love it


How many tok/s are you getting? What gen mbp?


My M4 Pro 48GB gets about 13tok/s, in both 3.6 and 3.8 27b Qwens. Qwen A3B and Gemma get closer to 100tok/s from memory but the results are pretty poor for coding tasks.

Edited to add: for agentic workflow I’m running omlx which tells me it has about a 90% cache hit rate (tradeoff is some disk and mem space) - that noticeably changes the felt speed.


I get around 20 tok/s, 4 bit quant, MTP, 4 bit KV cache quantisation. On an M4 Pro 48Gb.


i suspect ppl dropping generic "its awesome" comments are not actually using it and prbly just managed to get it running for a prompt or two.


I feel like it's 50/50 between people doing that, and people that have spent a lot of time tuning a system they are pointing at focused and well specified problems.


Seems threads about local LLMs on Apple hardware feature comments listing M3/4/5 at 48GB 64GB and not 128GB.

That is, users with M-series hardware that have less-than-max RAM share results whereas users with max RAM do not.

Speculating (not extrapolating), maybe users with machine that have max RAM are less interested in running local LLMs and are less averse to paying services for compute?

Personally, I’d love to see what output max RAM M-series Apple hardware in these threads.


Qwen3.8-27B runs at 59.5 tok/s on my M4 Max, 40-core GPU, 128 GB

I use it occasionally for classification and other tasks but I wouldn't trust those smaller models with the real work and for larger data processing it's too slow, e.g. a dataset I wanted to classify would've taken 56 days on my laptop vs just paying the cheap Luna prices to openai and getting it done in a few hours.


59.5 t/s is really good. Which engine/quant are you using?


Not sure I would trust Luna with that. Deepseek Pro Max and Code Mode I would be more inclined to trust.


I've been using Qwen3.6-37B-A3B on an M1 Max w/ llama.cpp and for my practical uses I prefer it to qwen3.8. When 3.8 does answer its slower and, qualitatively, marginally better than qwen3.6, but 3.8 often ends up in unresolved thought loops and runs slower. The Moe 3.6 on my setup is much faster, 500t/s peaks, 30t/s typical, vs 3.8 150 peak, 4-9 t/s typical.

While I've spend a little time tuning, I'm assuming there will be deeper tuning for 3.8 that might close the gap.


Yeah, that's my experience. It's a big "wow" factor to get a non-trivial LLM running on my Mac, but it's actually not that useful. Like trying to use Photoshop at 8 FPS.


It’s not because you didn’t find use cases for local LLMs that there are none.

I use local LLMs on my Mac Mini M4 Pro with 48G to review text messages tone, act as a text correction tool, act as a code review tool, to do code agent work, generate code snippets, etc

Gemma 4 26B A4B gives me steady 20 tps.


Regarding this analogy, fps don't matter as much for Photoshop, since it's not an immediate mode GUI. 8 fps would be quite ok for comfortably getting feedback on live image filters and such.


M3 Pro 36GB. I am getting 17 tps with MTPLX.


Prompting and having the model fill out the rest makes it feel like jazz or at least an improvised result of a song


It has to be Old Ebbitt Grill


It is. 1856.


Pressure test


> my parents thought it was much more serious than it was

How'd it end up? I'm wondering why it ended up actually less serious


It basically ended up with me learning that what is mostly pursued is the sharing part of piracy, and not doing that.

I got in a bit of trouble by my parents at the time and probably just stopped pirating stuff for a bit. No more issues from the ISP as we complied with the C&D and I stopped seeding stuff or uploading to shared FTPs (I can't really recall which was the issue at the time)


> PlayStation users are moving to PC en masse

PC is even more digital-only than Playstation. No one buys physical games on PC. The only difference is that Valve has been a very good steward over Steam. Theoretically, PC can get as enshittified as PS.

I guess there are other DRM-based purchasing platforms, and there's also DRM free ones like GOG so PC gamers have choice, but those feel niche mostly.


The difference is that there is no one steward. You don't need steam, there are lots of other ways to get your games. GoG, Itch.io, Windows Store, or just the developers webpage.

On PlayStation, switch, or Xbox you have only one gatekeeper, and they do not respect you


The thing about PC though is that there is no exclusivity. Steam built an extremely well respected brand but if that were to turn, the moat is shallow, install a competing client and buy games from there instead. The only internal competition for consoles is digital or physical retail (Or I guess buying game codes could be pseudo-digital).


Voice really matters in writing. If everyone uses Opus to write without editing, then it all sounds the same regardless of who it came from.


I miss that 2010s startup look


That lobster font we all used for our startup names was legendary


lobster.ly


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: