Hacker Newsnew | past | comments | ask | show | jobs | submit | argee's commentslogin


> the gist of the post is "do something because you feel like doing it, not because you feel like you have to." and "the goal is to just do your own thing.".

So basically, a main character life, the life of a player. A player can even choose not to play the game, and it isn't their whole world.

The reason this post is incendiary is due to their overwrought metaphor which ultimately works quite poorly, if at all.

Plus, it reads like a 15 year old's diary entry.


while it's not how i would have written it, i am struggling to understand many of the interpretations being posted in the comments. it's like people skipped over all the lines that say (or imply) "just do your own thing", and got immediately hung up by the use of "npc".

>Plus, it reads like a 15 year old's diary entry.

okay?


I'm just trying to explain why people are posting what they are posting. E.g. "likely from a stress tolerance so shallow and under-developed that the slightest inconvenience or expectation feels like a crisis" stems partly from the immature writing style.

But given that you are posting passive-aggressive "okay?"s to my explanation while purporting that you don't understand the comments while alluding to the notion that you do understand exactly why...makes me think you are just a troll, or worse.


>makes me think you are just a troll, or worse.

not trolling, just found it to be an odd thing to add on. none of the other comments that i read at the time of my comment, except yours, brought up the age of the author.

i typically try to avoid using age as an insult. i find it rude. you don't, apparently.

(out of curiosity, what is the "worse" thing you think i am?)


> The idea that optimization matters less over time is short-term thinking.

The more you think about things and the deeper you investigate, the closer you get to realizing that most aphorisms, advice, and even "facts" are a combination of best-effort philosophy and selection biases from the lowest common denominator. Add in the Barnum effect and you end up with an ocean of inconsequential platitudes being parroted by every Tom, Dick, and Harry.

It's very difficult, requires active effort, and goes against human nature to act rationally. Most people don't even want to do it.


I have an M4 pro (48 GB ram) and I run Gemma 4 26b a4b at 52 tok/s and Qwen 3.5b a3b at 72 tok/s. Both 4bit quantized. These are enough for my needs and the performance is more than good enough. I'm not running the MLX version of the Gemma model, if I did the inference speed would likely be a bit better. I wouldn't use them for coding features though.


My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?

Wow, that's just about half the perf. I'm not sure what you're doing differently, though our hardware is a bit different: I am on a Macbook Pro M4 Pro, while you're on a Mac Mini.

I would try a different version of the model from HuggingFace while ensuring it's MLX. I'm also using LM Studio, not oMLX, and I've seen some threads like these:

https://www.reddit.com/r/LocalLLaMA/comments/1spuwir/omlx_10...


I have a (now discontinued) 64gb mini pro and I’ve found the same qwen model to be almost unusable unless I kill Thinking on each turn.

What are you using them with/for?


I do turn thinking off most of the time for both models. I made a separate comment detailing my use cases.

> enough for my needs

Which are...?


Some examples (keep in mind this is all indefinitely free for me, no burning quota away):

1. Getting information (such as information about hardware unfamiliar to me) when not connected to the internet, which happens occasionally in my case.

2. Continuing to learn Rust by way of toy examples, puzzles, and comparing aspects of various solutions, for example from LeetCode.

3. Reformatting data, for example from a PDF to a markdown table, or converting receipt images to text.

4. Simple translation/explanation (e.g. I'm teaching my wife one of the languages I speak but sometimes may not know/have the words to explain the full nuance of a translated word).

5. Summarization. One of the webnovels I'm reading has some very boring parts I don't want to slog through, in those cases I simply make the LLM summarize that part and move on.

Etc., you get the idea. It's not unusable for coding, but it would make many mistakes when making a whole feature and the context lengths are limited to around 30k-40k tokens by my RAM. I could give it access to the web but I simply use an online model when I need that sort of thing, again partly due to the context limit.

Edit: The MLX version of Gemma 4 26b a4b does about 62 tok/s.


What kinds of tasks are you using this for?

Isn't that begging the question? For one, people can wear a lion headdress, making it something that does exist in real life. Secondly, it seems to be carved from a single tusk, which might have forced some choices when it comes to anatomy, rather than being "as conceived".


Wouldn't you be wearing a lion headdress for some kind of symbolic reason though? Like it's not a very practical piece of clothing. I would guess you either want to draw some kind of parallel between yourself and lions, or maybe remind people that you're strong enough to have killed a lion. I don't think there are any animals that take hunting trophies like that.


Heh. "What does it signify, this lion with a man's body?" - "ran out of space".


But that's also just ignoring the fact that humans seem inherently drawn to the fantastic, whether in religion or something else. There isn't really any reason to believe that wouldn't have existed 40k years ago, correct?


That's what they mean by begging the question; you don't really get much useful information from trying to cite something as evidence of what you've already assumed.


This is the earliest evidence of such fantasy we have. It's not that this piece at this time was the earliest human fantasy, but there is a cutoff line somewhere between humans and our last common ancestor chimpanzees.


Or as one might say, predictably irrational.


> I kinda feel that this person makes a lot of noise in this website.

This person pretty much lives on HN:

https://news.ycombinator.com/user?id=simonw

Personally I do think that people with huge amounts of karma contribute to a bit of a groupthink atmosphere, but Hacker News is set up to be more in their favor (which I think makes sense, to get a lot of karma they've been around a while and also understand what gets upvoted around here).

> Can we rate limit them a bit?

This submission was not made by simonw themselves, so I'm not sure what the ask is here. I think people ought to be submit to post whatever websites they want? It would be interesting to block simonw's own votes from influencing simonwillison.net submissions, but that would be hard to impose fairly across all webmasters.


> It would be interesting to block simonw's own votes from influencing simonwillison.net submissions, but that would be hard to impose fairly across all webmasters.

I think that would be undesirable. Let people vote on themselves as long as they are using their own account and not a second one. If their own vote is all they get, HN will filter them out anyway.


Part of being smart is being able to diagnose certain ideas as dogmatic, which is far more than can be expected from most.

"Most people would die sooner than think—in fact, they do so" — Bertrand Russell


Read this.

https://openai.com/index/where-the-goblins-came-from/

> We retired the “Nerdy” personality in March after launching GPT‑5.4. In training, we removed the goblin-affine reward signal and filtered training data containing creature-words, making goblins less likely to over-appear or show up in inappropriate contexts. Unfortunately, GPT‑5.5 started training before we found the root cause of the goblins. When we began testing GPT‑5.5 in Codex, OpenAI employees immediately noticed the strange affinity for goblins, and we added a developer-prompt instruction (opens in a new window) to mitigate. Codex is, after all, quite nerdy.

Note that the permanent solution was not just adjusting the prompt, and in fact being perfectly aware of that option they decided on a different course of action. That means either you are wrong or they are wrong.


> making goblins less likely to over-appear or show up in inappropriate contexts

So inappropriate goblins are still likely, just less so…

Hey, remember when tech bugs were things like buffer overflows or cross-thread performance impacts? I miss the days when our war with system goblins was purely metaphorical.


Did they ever discuss what the root cause of the goblins turned out to be?


> Anyway, you might have more luck just writing to it in your native language.

This is potentially expensive advice (at least for many mainstream options). Where an English word like "literature" is one token, a couple of Chinese characters that spell a word can be 4 tokens. You'll pay more for input/output and get less of a context window (per word) too.


Incidentally, according to https://gpt-tokenizer.dev, in gpt-5, "literature" is two tokens ("liter" + "ature"), whereas "文学" is one.


My bad, the English word I had in mind was "technology" (技術) but I misremembered due to feeling like "tech" should be its own token.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: