>Maybe I'd waste a day trying to get an API to do something it turned out it couldn't do.
I was working on a side project recently. I had spent months designing the data model in my spare time, thinking through how to make it as elegant and durable to change as possible in the long term, since (if I launched it) the repercussions for getting it wrong would be significant.
Once I had a working design, it probably would have been several more months to build a working prototype and start testing it.
Instead, Claude knocked out the prototype for me in an afternoon. And it immediately became clear that it didn't work: not because the data model didn't solve all the problems I wanted it to solve, but because it didn't fit the shape of how I quickly learned a normal person would need/want to interact with the product. I was so focused on the long term, that I never thought about what the first five minutes of a user with hands on the thing would need. And the changes needed would be significant.
Maybe there's some variable reward mechanism. But I sure was glad to be able to pull that particular slot machine handle and learn that than waste even more of my time on what was a dead end.
>No. People have an inner monologue, partial results and ideas and they remember that.
Yes, but the vast, vast majority of decisions you make either don't take place via an inner monologue, or include details that were not actively/consciously "thought" and reasoned with in your inner monologue.
And yet, when asked why you did something, you're not likely to respond "sorry, that decision was made subconsciously". Instead, you use your inner monologue to try to backfill in a reason why. That reason may be correct, or it may not be. You don't actually know, since you have new data that may be updating your own internal state as you try to rationalize it after the fact.
But is it really what we want, machines with the same defects as humans? I don't want a pocket calculator that make mistakes "sometimes" so I have to double-check the results, I want a pocket calculator that works (to those who want to argue that pocket calculators don't give the correct result for (1/3)*3: STFU).
>But is it really what we want, machines with the same defects as humans?
Sort of, actually. I think we humans actually have some intuition that we'd be more effective if our cognition were augmented more directly by machine strengths: the ability to run precise calculations, more memory, ability to look facts in some sort of knowledge graph.
I think we're on the right track, but instead of augmenting humans with machine strengths, we're building intelligence in hardware in a way where it can access that augmentation. Plus, then we can quickly distribute updates, run parallel instances, etc.
If intelligence is compression, and hallucinations are essentially loss, then as the models grow in size performance (at least as far as hallucinations) should reduce. Or we'll get things fast enough that we can afford to stop relying on model weights for memory and check an increasingly larger set of discrete facts as part of reasoning.
Right now, the models are making trade-offs. As compute grows, and inference gets faster, we can make fewer of those trade-offs and start to use the unique strengths of machines to fill the gaps we're seeing, I suspect.
One of the biggest strengths of a computer is reproducibility. The worst software bugs are inconsistent or non reproducible. The least useful calculators apply rules inconsistently, to your example.
the inconsistency of LLMs is by far one of the biggest gripes I have with them. Closely related to their apparently deep desire to avoid following instructions.
I know these are both a byproduct of noise (which is somewhat tunable) and noise is inherent to these systems in a lode bearing way.
I still hate it. it’s holding the technology back. I don’t honestly see how we can safely or even successfully approach the idealized realm of AI without bypassing this problem, which to my understanding, probably means not using language models at all and trying a totally different approach. But I really don’t know much about machine learning, I’m a super novice compared to a lot on this website.
When we are really thinking about something we do it forwards, backwards and middle out, and regenerate and distill many times.
When we do meta thinking about that process after the fact, two things happen. 1, we change our total “thought” by adding that meta thinking to it. And 2: it’s a very lossy process, because we don’t have very good data about what our brain or mind was actually doing during that first think and emotional factors are nearly always at play and even more complex.
Now for the more complex AI, the fragmented process of multiple agents and loops and reruns are pretty similar to that first think we do. At least structurally. But the meta think is where they differ. They have no emotion, but they also have even worse data about its own function. They constantly degenerate so I would argue their “changing the thought by thinking about it” factor is also generally way higher than ours.
Getting better at consistent/reproducible thinking, with many ‘steps’, that leaves good documentation of that thinking behind for future analysis, has to be one of the more important areas for the big flagships going forward. I’m certain that “what is this fucker doing and why” is the biggest pain point for AI researchers. Or the math, it’s usually the math.
But you’re correct in the general structure; they generally do the same post hoc analysis we do, just noticeably worse because of their opaque nature(even to themselves) and general degenerative instability.
Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.
Certainly. The singularity sort of assumes that there is not a fixed limit to intelligence, or at least that if there is, it's quite a ways away. That may not be true.
Why? If there's no evidence of fraud, why are you imagining some new category of fraud that these companies have independently conjured up?
Suggesting that an entire industry is in cahoots to invent and participate in an entirely new fraud-like scheme, including companies that have lots of wealth and growth and much to lose from such an exposure...seems awfully conspiratorial, don't you think?
>Ed's rant about how impossibly weird and inhuman the framing was revealed how little he knows, and cares to know, about the mechanics of the businesses he claims to profile. Yes, something like "on plan" is jargon, but a shred of reasonable journalistic curiosity (a Google search) would explain what it means and how it's a normal statement for a career-sales CEO.
This is precisely the sort of feedback an LLM could provide him before he hits the publish button, funny enough.
It wouldn't surprise me if we start to see minimal performance gains from incremental changes to base models. It seems like the gains from the Opus 4.5+ incremental updates were a result of Anthropic learning a lot about post-training, the gains from RLVR, etc.
If new post-training techniques are seeing diminishing returns, we could just be back to waiting for new large pretraining runs at larger sizes for gains (even if those ultimately end up getting distilled down into smaller models because the economics for serving anything larger than Fable isn't practical).
it seems to me that OpenAI is the only actual lab that truly understands reasoning. they have the best reasoning efficiency, they get pretty uniform improvements with more reasoning compared to other labs. (theres been plenty of graphs where models do worse with more reasoning), and i suspect their models are a lot smaller than we think.
i think the next gen of openAI models are going to be quite insane tbh.
I never thought I'd be the sort of person to go to Burning Man, and this is my first year off since I started going nearly a decade ago.
I started by going to a regional burn (or what some people call a "ten principles event") because a girlfriend had a bunch of friends that wanted to go. I had always been a bit in awe of the art and such from BM itself, but weary of the drugs/party culture (I was a nerd that made the annual trip to San Diego Comic-Con, PAX, etc) and didn't think it'd be my scene.
Without going down an entire rabbit hole: it changed the entire trajectory of my life. I've made the best friends I've ever had and now spend a good chunk of my time working on large art projects, helping run a theme camp, and being involved in this group of weirdos year round.
A few years of being involved locally made the desert seem like an inevitability, and when I finally went, it wasn't the life changing experience everyone told me it'd be: because it already was, even without me being there. But it did quickly become one of my favorite places in world.
My honest advice? If you're anywhere near a Burning Man regional event, find a way to get involved and check that out first. I think that's a great way to see if you vibe with the culture around the event, and it'll give you a sense of whether trekking out to the desert is a good fit.
Also, as an aside on the wealthy tech/finance front: one of the first camps I shacked up with was mostly a bunch of blue collar types who liked to work hard, with two retired grandmother-type figures who helped run camp logistics. Tech money (including mine) certainly helps fund things (because the city is expensive to build), but at least in the scene I run in, the tech types are not the majority, and sure don't act like your typical tech bros.
I was working on a side project recently. I had spent months designing the data model in my spare time, thinking through how to make it as elegant and durable to change as possible in the long term, since (if I launched it) the repercussions for getting it wrong would be significant.
Once I had a working design, it probably would have been several more months to build a working prototype and start testing it.
Instead, Claude knocked out the prototype for me in an afternoon. And it immediately became clear that it didn't work: not because the data model didn't solve all the problems I wanted it to solve, but because it didn't fit the shape of how I quickly learned a normal person would need/want to interact with the product. I was so focused on the long term, that I never thought about what the first five minutes of a user with hands on the thing would need. And the changes needed would be significant.
Maybe there's some variable reward mechanism. But I sure was glad to be able to pull that particular slot machine handle and learn that than waste even more of my time on what was a dead end.
reply