I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.
Grandparent comment has zero to do with the article. It's just GP generically bitching about Meta. (Your "quote" of the comment does not appear anywhere in the actual comment.)
Sorry what is the punch-line here supposed to be? These investments obviously increase correlation coeffs., but these are highly correlated stocks to begin with.
valuation gains on investments are one-offs, not signs of sustained improvements in profitability that would warrant higher market caps
when those valuation gains are in turn the result of circular financing schemes (a bakery giving out money so that people buy bread from it), we're getting to a dangerous situation
Circular financing is an issue if the wealth accumulation stays in the chip -> model provider ecosystem. However it seems like the labs are making quite a lot of revenue from the chips (customers).
Demand for compute is far outstripping what was projected in the initial rounds in which we pearl clutched over 'circular financing' - which is hardly distinguishable even from the classical bill of exchange, the core phenomenon of the money market in Bagehot's day, in which I provide you inputs and you give me a share. I keep reading people saying what amounts to: the primitive bill of exchange was circular!!; if the seller of inputs has reason to lend, so does everyone else etc etc. In fact if the projections about final sales are correct, it is plain all these deals will be fine.
The issue is simply that the posted article begins with a review of recent earnings/ revenues, but fails to discuss that a substantial part of those revenues are investment markups.
Whether it matters we don’t know yet, but it’s a fact worth noting. A better article might have tried to argue why it doesn’t matter
I disagree. A chess engine has a very well defined goal and is working its patterns to reach that goal.
In software development, except in a minority of cases, the goal is being set by non-technical people, who have no clue how the resulting system should end up, only what it should look like to a user. I think a well-versed programmer can give the AI direct instructions on how to implement something, and just let the LLM write the code to implement/refactor the infrastructure behind the feature they're developing. For me the plan mode of Claude Code is (anecdote incoming) much faster(TM) and better(TM) if I give it concrete instructions - then it writes the plan, asks questions - and then implements it. Very little complaints after that (I usually don't let Claude do any sort of visual QA)
People take much more time off during the summer in the USA than during other times of year. It’s somewhat regional, but it’s been true in all the place I have lived. Probably July is the most common month though because in some regions schools start sometime in August. It’s not uncommon for people to take 1-2w off during that period. More than 2w often requires extra approvals, so is less common.
Hijacking this comment to ask about Huberman, can he really be lumped in with the other two in that list?
I listen to his podcast for a while and stopped when I realized most of what he is talking about is marginal stuff or ”new research suggests”, which I found less interesting.
But has he really ever been a grifter selling complete bunk? Not talking about the ad-segments, which are of course payed endorsements.
> You can think of Q collar as a baby version of health scams promoted by wealthy and powerful medical-adjacent celebrities such as Dr. Oz, Andrew Huberman, and RFK Jr.
This sentence does not imply that Huberman has endorsed this product.
That is fair, especially on this occasion. On the other hand I have found some of his material very helpful, particularly around sleep and helping me maintain my circadian rhythm.
reply