Hacker Newsnew | past | comments | ask | show | jobs | submit | arctic-true's commentslogin

Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.

The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.

[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/


Recursive self improvement of their upcoming IPO value maybe.

They are fluffy PR pieces otherwise.


I tried to build a procedural 3d asset pipeline for a specific use case.

Before the Opus upgrade in November it was basically no way of doing this. I gave up very quickly.

After November i tried again, and no model could build me anything relevant.

Now it just works. Took me an hour to progress to a point were i'm happy.

Whatever they do, progress is still real, still way faster than I assumed

The list of Ubuntus 2404 LTS CVEs is HUGE. Another indicator that a lot of stuff got a lot better fast.

Feel free to be as dismissive as you want, but if you are not careful, you might be 'suddenly' surprised and you might not be prepared for the conclusion of AGI level agents.


A new account affirming OpenAI’s PR message doesn’t persuade me. It’s actually harder to trust.

What would being prepared look like? A house in the woods with a store of food? Knowing how to pick door locks and set up a militia in the desert?

That's really interesting. Can you go deeper on how you achieved that? A loop comparing procedurally generated model with reference image?

How can you possible say this sort of thing in context of what looks like a millenium prize being solved.

I swear there's nobody blinder than those who won't see.


Because it seems like most of the work may have been done by human mathematicians and cribbed by OpenAI at the last minute

We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

> We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".


Only if you interpret statements as being binary logic.

"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.

Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"


Yes, you were guessing. That's the only thing you could be doing, since, as you said, nobody actually knows.

Other people are allowed to have their priors, too. Even under a non-informative prior, the weight of the evidence (Tristan's account, OpenAI's announcement, and Bubeck's "denial", if you want to call it that, plus multiple other mathematicians coming forward with similar experiences) pushes the probability mass toward some degree of impropriety.

What exact evidence are you incorporating into your prior to come out with this posterior?


> not being a simple case of intellectual property theft

No, it's an aggravated case, since it's the same way they got all of their training data in the first place.


Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).

Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.


We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:

> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].

But it is in no way equivalent to this:

> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft

That is a different sentence.


An opportunity to correct myself? Respectfully, I decline. If you have trouble parsing my sentence, it's not meant for you.

> it doesn't seem to be the case

Based on what? Your crystal ball?


Based on: my experience working on AI for 32 years, including a decade at Google including working on large-scale model training systems that used user data and complied with various user policies around data retention, along with a few decades working in science/tech making decisions around ambiguous data.

In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.


The human mathematicians didn't solve the Navier-Stokes problem, they solved the Euler problem. And they were extensively using LLMs to drive the work, as described in the Buckmaster statement.

Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.


I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.

You have to look at the incentives

I swear to god, people would look at the successes of Xerox palo alto and just shrug and say - "yeah, but I mean, this is all marketing"

Xerox is incidentally a really good example, because precisely nobody ended up using the desktop experience Xerox made. They ended up using the desktop experience that Microsoft and Apple made and shipped while Xerox the actual company faded and memory of those original parc research teams faded into obscurity.

This is what I keep saying, and it feels like I'm taking crazy pills here!

Is nobody else astounded by this?


The thing only nerds know about that wasn’t in the news every day? Not the same thing.

Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.

A working Lean proof doesn't care what the incentives are.

People will cling to views as long as they possibly can, despite evidence slapping them in the face.

Eventually it won't matter. Arguments over whether LLMs are "truly" intelligent are going to be a matter of philosophy, and look a little silly.


It literally reminds me of climate denialism.

“Line go up! That is bad! Planet might become unsafe for human life.”

“Nuh uhh! Malankovich cycles and humans are a drop in the bucket! Krakatoa! See!”

“All of California is on fire!”

“Haha, stupid shrill liberals! Go rake your woke forests! Drill baby drill!”

“Are you kidding?! Look at this graph.”

“That’s just propaganda from the elites of the Build-a-bear group!”

“Do we even live in the same universe?”

“And 5g causes COVID!”

“Oh, I guess we’re don’t.”


How? By being knowledgeable, intelligent, and intellectually honest.


This comment was applicable 2 years ago. It isn't any longer.

I found this post interesting in that reguard: https://www.lesswrong.com/posts/thXohzXrWCA2EhZCH/mateusz-ba...

Compute will always be the bottleneck even if this were true.

As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.

In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).


If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.

We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.

I'd argue we've been in an "everything-has-an-API" world for a long time now — it's just that discoverability of said APIs is still crap.

And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.

Humans do this too.

Yes, but for humans it creates friction. Seeing a captcha makes me think twice and thrice if I really want to visit that site so badly that I'll endure the suckage. LLMs don't care, for them the friction doesn't exist.

For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.

Yes. In other words: the singularity. I'll only believe it when I see it though.

I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.

> coming around to not liking the term singularity

Bit ironic given the model’s alleged finding…

Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.


From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".

It doesn't imply that. The singularity is just the inflection point.

Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.

And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.


Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.

Certainly. The singularity sort of assumes that there is not a fixed limit to intelligence, or at least that if there is, it's quite a ways away. That may not be true.

I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.

But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.

From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"

The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.


> the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.

Such a good description. No sentience here, just raw computational power


How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.

Hello good sir from the 1700s pre-industrial revolution who doesn't think that mines and factories can be automated.

The factories that supply the equipment, maintain the equipment, the energy inputs, the financials of those mines are not automated.

People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.


The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.

https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.



Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.

Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged


I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.

It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.

I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.


I 100% agree with your impression. For a good while to come there's going to be a bunch of Jevons Paradox to all of this, but adoption of architectural changes like that photonics adoption is exactly the type of adaption to circumvent bottlenecks I'm referring to. We're going to hit hundreds of bottlenecks and each one will inevitably breed new approaches and technology directions. And the forcing function won't be talking about them, but implementing them, seeing who wins and taking lessons.

pi-fs will solve all our data compression problems.

Eventually recursive self-improvement includes reducing bottlenecks.

Eventually the bottleneck might be people themselves.

Improbably, the real bottleneck is energy.

Which is to say, scalable and open-ended capability of ramping up physical infrastructure.

I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.


And the goalposts move again

Is it? The first article says they’re on-track to building an automated AI researcher by March 2028. So they haven’t achieved RSI yet, I would think?

> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.

Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?


I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.

It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model

What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question

If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?

Edit:

OpenAI have admitted they were training on prompts at the time they made their breakthrough

https://mastodon.social/@tristanbuckmaster/11723647135247030...


If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.

If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.


The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools

The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem


In that it isn't able to genuinely solve problems

Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.

This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.


I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach

You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.

I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?

If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.

Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?

OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct

Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.

Terms and conditions are almost entirely about the company doing things that would otherwise be illegal.

I cannot parse that sentence.

If you think a business’s terms and conditions change the illegality of an action, you should consider researching that assumption and speaking to some lawyers.


Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?

Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.


You don't need to speculate here, since much of the story is not being disputed.

The recent ground-breaking work on Navier-Stokes "blow-ups" was done over a period of years by mathemtaticians Diego C´ordoba and Luis Martınez-Zoroa.

NYU professor Tristan Buckmaster and Anthropic employee (& mathematician) Levent Alpoge took the above work as a starting point, and over a year with LLM assistance developed a blow-up proof under certain conditions.

Buckmaster: "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments."

Buckmaster says he thinks that Martınez-Zoroa, whose work this all builds on, deserves the Fields Medal for his work.

OpenAI claim that on Sept 1st they heard a rumor the problem has been solved (which happened on August 15th), and then decided to re-solve it themselves using a 2-week old model, then later reached out to Prof. Buckmaster and Levant to come to some agreement to co-publish.

OpenAI: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ". In other words, not only did they deliberately choose to tackle a problem they heard had already been solved (in turns out only partially solved), but they may have done so using a model that was aware of the successful way to attack the problem.

It seems there are three potential scandals here:

1) OpenAI by their own admission chose to try to scoop mathematicians who they had heard had already completed a proof

2) OpenAI may have used a model that had seen "de-identified" messages indicating the direction to take

3) An OpenAI employee essentially threatened to "ruin the career" of the NYU professor who had been working on this if he did not cooperate with them

The direct plagiarism possibility, 2), while it should be a warning to anyone using OpenAI's models, doesn't need to be true for OpenAI to have benefited from the researcher's work. It's enough that they heard Navier-Stokes had been solved and could then go out with their swarm of 10,000 agents and $20M of compute to hunt out the latest research and brute force it.

Magnus Carlson once said that if he wanted to cheat all it would take would be for someone to indicate to him (a wink from someone in the audience perhaps) when a position warranted more time to be spent on it (because there was something important to be found if he did). It seems that, at absolute minimum, this is what OpenAI did here, although in context of math this is not cheating - the "wink" was a rumor, originating from who knows where, that a proof existed (but had not yet been published) and therefore there was potential to rush in and scoop rights to publish or co-publish.


The rumors were that Anthropic had solved Navier Stokes and was sitting in it to maximize IPO hype. It was all over twitter. In the first place the intent was never to scope mathematicians, but their closest rival.

The proof turned out to be different from what Levant and Tristan was doing.


I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?

All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.

That's it. The rest appears to be wild speculation.


Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.

Never ever touch those requests. If you get a side by side comparison just resend the prompt.


do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning

I was thinking that the training is more curated, so that the methods are learned from experts, and that measurably successful behavior is reinforced.

Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.


hmm, maybe not sentiments, running commands can produce binary results to reinforce, but that's also speculation

>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.


even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?

Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.

Publication, though? Slimy is right.

But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.


>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.


An Astra sized model takes months to train - they are not saying 2 weeks to train from scratch. The only only interpretation of this "2 weeks" claim that is consistent with reality is that they mean 2 weeks of additional training on top of whatever their starting point was, so it's more like this:

|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model


The things you don’t care about are highly relevant to that claim

Elaborate?

If the model started from work that Tristan Buckmaster had already done, and was aided by an entire team of mathematicians as he alleged, then it's not as capable as you seem to be precluding


With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.

It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).

So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.

Second biggest fumble after Google.


Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".

Don't they own a large portion of OpenAI? Things could be worse

Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.

With 10,000 agents and $20M of compute this is just brute force search.

It's a bit like telling 10,000 kids there's an easter egg hidden over there, pointing to one corner of your yard (or having "heard a rumor" it was hidden in that corner).

If you have $20M to spend on your problem, then yes, AI brute force search is an option, but unless you know a solution is possible (as OpenAI did here), you may still be wasting your money.


You jest and that is OK. Brute force search is not something you can do over math problems of that difficulty or anything with combinatorial complexity.

To me it feels closer to taking the top 10k human mathematicians on a large retreat for a year and having them self organize to collectively solve this problem—not kids and easter eggs.


I'm not joking. Compare to a super-human MCTS system like AlphaGo or Stockfish - once you condense the expertise of your top 10K world experts into a board evaluation or policy function, then the rest is brute force.

Whether this type of agentic swarm approach can be considered closer to MCTS (search), or closer to a less structured GOFAI blackboard type approach (perhaps more like your mathematician retreat) I'm not sure - I don't think they've released any details of the prompt(s) and how these agents were collaborating and building on each others work.

The other part of my easter egg analogy is the direction to "look over there", corresponding to OpenAI specifically asking their hoard of mathematicians to work on Navier-Stokes since they knew it was solvable/determinable, and they certainly had the public work that Buckmaster/Levant were building on as further direction, as well as perhaps their prompts. Unlike Buckmaster/Levant, this wasn't just a couple of humans with a university research grant budget, this was apparently a not-so-small team at OpenAI (says Buckmaster, per a group call he had with OpenAI), with an unlimited budget, so it's hardly surprising (or in the least bit impressive) that they were able to duplicate and surpass their work.


If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.

As far as I understand the 10k agents worked on the proof. The lean formalization came later and was easier/faster than getting the proof.

What makes you think they were coherent?

They managed to solve a problem that was beyond current human ability.

That was the net effect (assuming what they solved was the actual problem and not a loophole in the problem statement or a lean bug). My point is they need not all work coherently to do that -- for example, for all we know 3/4 of them went off the rails, their results were pruned, and the relevant results came from a random subset that happened to produce something useful.

If you work with distributed systems, you still call that scenario a success. On the other hand, if the 3/4 of agents going off the rails bring down the whole mission, that is a failure. The latter would have been my guess with current models scaling to 10k agents.

I am not an expert in lean4, but I could follow parts of the high level lean definitions of the problem statement in the repo. A lean bug would be a fun scenario; I am certain this proof will receive the deserved scrutiny, and if it uncovers a bug, it will make the story even more exciting. It is extremely unlikely to be the case, however, because the 10k agents working on the proof didnt use lean, so it would have to be a math logic error that translates to a lean bug—perhaps something the agents picked up during training?


My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.

Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.

When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.

We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].

Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.

[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...


The singularity happening under trump? We could have had star trek, instead we're getting the combine.

pick up that can

"I love Singularities. I am the best at Singularities. Everybody knows it ..."

Beautiful Singularities.

That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.

And then it will be shut down, proper guard rails put in place, and the new version will accelerate the cleptocracy.

You think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi

The scifi pov has a good track record as this point, you people gotta be more open minded

it proved navier-stokes taking over the us government is easier imo, any idiot gets to be president

Many present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?

If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra.

"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable", or "twice as" for that matter), then put those 2 weeks of training in to focus on those gaps.

At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.


>If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra

OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up.

>They are not shy - if there was a more impressive claim they could make, they would have made it.

What claim would that be ?


> What claim would that be ?

Huh? I'm saying there isn't one.


Okay. I was just confused.

Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?

OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.

it's possible to do a RLHF or RLVR pass pretty quickly. I'm almost certain a full pretraining run isn't possible within that time frame.

Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.

Astra was trained more than two weeks ago.

Astra was in use by OpenAI employees for more than 3 months internally from rumors I heard

The internal model they mention is different from Astra.

They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.

I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works

the model is a set of weights, you can take a snapshot and test it. Reinforcement learning itself is largely testing and tuning.

And... they found this "solution" in 88 hours or so.

It's all gas no brakes now boys and girls. Hold on to your hats!


Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.

Yeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch

The chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.

That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.

Isn't the y axis just what portion of the open problems it could solve? The axis is unlabelled though, I'll give you that

The x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?

Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.

Yeah it's Bel

it's buried because due to the drama the evidence is scarce

Pre-IPO marketing?

Even if it is, Anthropic better have a few things up their sleeve

If pre-ipo marketing pushes them to train a model capable of resolving a millennium problem in mathematics in a weekend, then, to quote XKCD:

  "Mission. Fucking. Acccomplished."

https://xkcd.com/810/

I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?

They are highly capable, no doubt about that, but:

1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.

2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.


"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."

Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.

And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.


1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars. 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.


> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.

Did I say otherwise?

> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

I know, but I don't know how that relates to my point, which is about the way they are doing it.


Sorry, I misinterpreted point 1), on X they said they didn't have people specialized in that specific field for prompting and steering the agents, just a group of mathematicians and physicists.

The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.


Of the seven Millenium problems, Navier-Stokes was the one most thought to be in reach.

I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)


the goalposts are on Pluto at this point.

I'd put good money on the fact that we will have a lot of distilled intelligence and yet the world won't look much different.

i mean that is already true

Is there any reference to a current goalpost position? Since you are claiming they had been moved, it's interesting from where exactly.

I'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample).

I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.


There were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathemati...

I brought this up here at HN, and in the ensuing discussion Buzzard himself replied saying he was somewhat joking (https://news.ycombinator.com/item?id=49011950).

Certainly, but the key word there is "somewhat". Progress is now happening so incredibly fast that I no longer know what to consider implausible.

Highly capable of writing math proofs, no doubt.

It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).


Lmao my friend, the whole "drama" is that there are allegations of plagiarism.

Or they trained a LoRA on the victims chats in order to launder their plagiarism.

The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.

Brain has loops and parallel connections.

Loops and parallel connections make transformer go brrr


I suspect some amount of long form writing will go the way of code - long form writing for the purpose of consumption by other AIs. Writing as a means of exchanging qualitative information, with no regard for how the reader will feel about it (beyond understanding what the words mean). Not everything can be distilled into data, but this doesn’t mean it is beyond the reach of LLMs.

On the other hand, long form writing for human consumption seems like it may evade LLMs for much, much longer.


What makes you think it will take a long time? AI seems capable of imitating any writing style if prompted to do so already, and I think it will get better on this quickly since the AI writing style is a main focus of AI labs right now. I can see no reason at all to believe this is a matter of years still, more like a few months.

The style you refer to is the "container" of the writing. The medium. Like the specific encoding of the message. What @arctic-true was talking about was the "content" of the writing, which is bounded from above[1] by the information content of the prompt.

So, I'm not sure if it's a question of time at all: if a LLM text contains some piece of information beyond the information that went into the prompt, where does this "extra" information come from? [Note, I'm not thinking about facts which could trivially come from the training corpus, I'm thinking specifically as information in the sense of intended message from sender (author) to receiver (reader)]

[1] cf. this comment where I explain this analogy between LLMs and noise channel in communication theory: https://news.ycombinator.com/item?id=49510244


You can have a great “writing style” and still put together really crappy long-form work. The problem is that AI writing, particularly creative writing, is too repetitive, too predictable, too trope-laden.

All of the things you say are very true in the near term for short form writing - a page or two of Claudeslop will probably be much easier to swallow in a year or two than it is now. But I don’t see a path to fully AI-generated novels or long-form investigative journalism becoming mainstream in the next couple of years.


Presumably, a truly superintelligent AI in a utopia scenario could figure out a way to enable people to continue working, even if the work doesn’t actually contribute anything, if it is really the case that work is essential to human wellbeing.

Now you might say “but those would be fake jobs.” If so, I have bad news about how many present-day jobs are fake jobs.


I agree with this. I thought the AI 2027 paper put it well:

> There are even bioengineered human-like creatures (to humans what corgis are to wolves) sitting in office-like environments all day viewing readouts of what’s going on and excitedly approving of everything, since that satisfies some of Agent-4’s drives.

I'll leave it up to you to decide if that's an enviable/desirable future.


I don’t even see why they have to be bioengineered. We have all-natural human beings doing this right now, except they are rubber-stamping human rather than artificial directives.

They don't, it's just at that point in the AI 2027 story the AI has already wiped out the humans due to resource contention (similar to how humans largely wiped out the wolves).

I think you are forgetting who owns these AI companies… their interests aren't your interests.

This is exactly correct. You see some discussion of a superintelligent AI becoming “superhumanly persuasive,” such that it can talk anyone into doing anything, and using this as a means to take power. But even with AI that was orders of magnitude weaker than superintelligent, people were getting talked into doing all kinds of stupid nonsense.

And that was without any ability to provide them with incentives. Imagine if an agent swarm got its hands on a huge pile of cash?


Otherwise-rational and moral people do bad things at the behest of powerful superintelligences with vast resources all the time. We call them “corporations,” and they can convince bright-eyed college graduates to do just about anything. Now, the corporations aren’t deliberately trying to destroy humanity, but they do lots of things that point in that direction. An artificial superintelligence would be able to do the same thing. We just have no way of knowing whether it would deliberately try to destroy humanity.

One thing that the AI doom discourse reveals is just how comfortable people allowed themselves to feel in the pre-AI world. There’s this belief that AI creates a risk of human extinction in the near-term which did not exist before. I think it’s telling that many of the leading lights of this movement are in their 20s or early 30s - too young to remember the Cold War. The truth is, we were never safe, and if all of the GPUs on earth were zapped out of existence right now, we still wouldn’t be safe. Life is random and chaotic and violent for most people most of the time, and we all just find ways to get through it. I suspect it will be much the same if an artificial superintelligence arises. Maybe it will kill a bunch of us, maybe it will kill all of us, but anyone who remains will eventually convince themselves that everything is okay.


Sure, the world was never safe for individuals, and even entire civilisations have gone extinct. For climate change and even nuclear weapons, most of humanity might perish, but there's the hope that some will survive. So, in the past humanity was not in a position to extinguish humanity as a whole.

Civilizations do not go extinct - species go extinct. And even that isn't as bad as it sounds: many genetic traits remain from extinct species and still contribute to the well-being of other species.

I guess I don’t see why “everyone will die” should scare me more than “almost everyone, including me, will die but someone somewhere might survive in a barely-habitable world.”

And this is where I think some of the problem comes from. A lot of AI doomers swam in the same waters as transhumanists, life extension enthusiasts, etc. and got very good at pretending that they were never going to die, or that they would live for a massively superhuman lifespan such that they did not need to think about death. AI doom is much scarier - and thus much more worth posting about - if it represents the first time you’re grappling with your own mortality, as I suspect it is for a lot of younger people in this space.


Autonomy is not the same as general intelligence. We already have all kinds of fully autonomous technologies that are nowhere close to generally intelligent. Plus, at a certain level of abstraction, human beings also need to be “prompted” to some extent by stimuli. And this is the funny thing about general intelligence as a concept: most of the definitions that come close to internal coherence rely on references to human intelligence, a concept we feel like we understand because we all live it all the time, but whose actual nature and structure is extremely slippery.

Yes everything needs an initial condition.

You could have the smartest human political operator, but if he has no context, no motivation, not much is going to happen.


I like this much better than chucking random numbers and letters at it like OAI was doing a year or so ago. I’d much rather have civilization destroyed by something called Astra or Fable than GPT-6.8s-latest.

I prefer "Smith", "Oracle", "Merovingian", "Architect"...

The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)

I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.

Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.


At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.

I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)

Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.

I’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.

I haven't tried asking it in Hangul but these particular artists (and the photos I'm using actually) are linked to their romanized english names on e.g. Fandom so it's not unfindable on the internet

This is exactly right. Now, that does mean something about Zitron: he’s a pundit, not an expert or a forecasting genius. You could point to any number of analogous booster types. Twitter/X somehow loves pushing these people onto my recommended feed. I remember reading breathless threads about how o3 was going to single-handedly end white collar work. There are just more of that sort of person, so none of them in particular gets the same amount of attention as Zitron, who seems to be the only person willing to go on the record against AI. Maybe his overall worldview is sound, and maybe it isn’t. But he’s not really making confident, specific predictions about the future.

Which leads me to a criticism of the piece: several assertions are described as “Wrong,” with no explanation or citation, which are not obviously wrong to my mind. For example, the assertion that “DeepSeek has commoditized the [LLM]” is at least debatable. It is a prospect that the major labs seem to have at least considered.


In my experience, people who make fun of AI's ability to eliminate white-collar work are exactly the same people who are white-knuckling it, hoping they can make another $300k next year to buy that Rivian.

People who don't need a salary look at the situation more objectively and, in general, can see that a lot of white-collar work is in peril.


I think that is tied to the general trend that people see themselves as far less replaceable than others. And I’m not sure that people who don’t need a salary are exactly objective about it - unless you mean people who live off the state or something, those are people who need their capital reserves to continue to appreciate in order to sustain their lifestyles, and are thus relying on the destruction of white collar work to justify the valuations of the technology companies which are underpinning the strength of the markets. It’s like saying in the 19th century that only the propertied classes should be allowed to make decisions about labor policy, because the laborers themselves have their judgment clouded by their position.

But I take your point, and I only brought up the o3 example to emphasize that people on both sides of the booster/doubter debate can be prisoners of the moment. There are boosters eternally convinced that the utopia (or armageddon, for the doomer-inclined) is either already here or imminent, and there are doubters eternally convinced that we have reached the peak. One of them will eventually be right, but neither has been yet. You can’t fault one of them for calling their shot if you’re not willing to see it happening on the other side.


> I think that is tied to the general trend that people see themselves as far less replaceable than others.

What if (almost) everyone is right, that their work really is more difficult to automate than others realize?

This could be the lesser minds problem. Based on it I am doubtful most knowledge workers are as immediately replaceable as everyone suspects.


The flip side of that is that at least half of all white-collar work was already useless before AI. If it hasn't already been eliminated, there's no reason to assume a-priori that AI can eliminate it.

If half of all white-collar work was useless before AI, and AI can replace those workers, then one can logically conclude that AI is useless.

Not completely. It just means you can have an AI do the useless busywork instead of a human. Something can be economically useful but still have proponents in a firm.

A lot of the "useless work" is busywork you do being a part of an executive's empire so they can claim they managed X number of people.

How can it do that?


Call your project Gas Town, write a blog about it, and investors will be swarming all over it thinking managing agents is even more impressive than managing people.

>In my experience, people who make fun of AI's ability to eliminate white-collar work are exactly the same people who are white-knuckling it, hoping they can make another $300k next year to buy that Rivian.

In my experience, people who believe that some white collar industry can easily be automated with AI usually don't work that same role or even that industry, and suffer from Dunning-Kruger bias. In my line of work, I'm desperate for Claude to actually do a better job at not producing a big mess, and it's failing terribly.

I don't work in law, but I feel like I barely need a lawyer if I can ask an LLM to interpret a contract for me. I don't work as a physician, but why can't I just cut and paste an MRI report and ask Claude what to do. I'm no plumber, but it I take a picture of some fixtures I am thinking about replacing, Claude will give me some items to add to the shopping cart from a plumbing supply store.

In all cases Claude was super confident and it wrote what felt really rational.

But in my experience it gets things right, in some amazing ways, but it gets things wrong, in shocking ways.

Everyone thinks it's someone else's job that it'll automate.


IME many of the people who truly believe it's replacing white collar workers are being replaced are invested in the stock market bubble.

is this comment written by ai?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: