Hacker Newsnew | past | comments | ask | show | jobs | submit | howdareme's commentslogin

Qwen is trained off of gpt’s outputs. This is both a positive and negative

They are not training a whole model in a matter of days


They where working on the problem for a year using codex.


Models are very obviously continuously updated.

Model editing to remove PII that slipped through, all sorts of things of that sort.


pretraining is months but they can totally fine tune in a few days


They don't need to train a whole model. They can feed it new information and fine tune it.


How can a vision model have taste?


If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.

For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.


> if the interface elements have strong information hierarchy

...but that's an example of a UX/usability matter that can be assessed objectively and non-subjectively.


I disagree.

Is 16 px or 14 px a better font-size value for a subheading, in a hypothetical layout? Immediately that kind of decision, where both options are objectively good for 12 px paragraph text, suddenly becomes an issue of taste that cannot be evaluated crudely by an algorithm.


Replace taste with consistent if that helps you. Can it follow a design system...


As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)

But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.

Of course only if the design is achievable in the design system.


This is not my experience at all.


So, formulaic output…the opposite of taste


Not really. Compliance with the letter of the law doesn't mean the intent is complied with.


Claude’s native harness is pretty bad though relatively. Pretty sure at one point it was scoring last on harness related benchmarks using Opus at the time


TIL harness benching.. https://www.harness-bench.ai/leaderboard.html - is this legit? Never even heard of nanobot.


Google’s pro models are almost certainly bigger than Openai’s lol


Why would that be? I am curious why do you think that.


E.g. because they are behind on research and so must compensate with size to achieve similar level of intelligence. At least this is what I heard.

For intelligence/size only OpenAI and Anthropic are the frontier. Google has more compute so it can compensate for that with size of the models...


I'd argue Qwen is pushing the Pareto frontier considerably further when you take size into account.


Because TPUs are more efficient, and its cheaper for them to field them in higher quantity since they own the chip.


Going to assume you didnt capture the data but could you add time taken to completion for each if you have it?


Does anyone know how they turn html to enable powerpoints so seamlessly?


The benchmarks they released


What do you mean? In most cases, the benchmarks show a larger number for Muse and a smaller number for Opus.


In Multimodal yes, but Opus is definitely edging out in Text/Reasoning and Agentic benchmarks.

I think the general skepticism is because they are late to race, and they are releasing a Opus-4.6-equivalent model now, when Anthropic is teasing Mythos.


LLM comment spotted


I mean looking antigravity, jules & gemini cli, they have have no problem with their developers fighting for resources


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: