Hacker Newsnew | past | comments | ask | show | jobs | submit | aschla's commentslogin

"If AI inference remains as desirable as Kimball expects, the evolution is likely to follow the same trajectory as the CPU. The CPU didn’t improve along a single axis but instead across simultaneously. Once transistor scaling slowed, chip and system architecture innovations of all kinds proliferated. The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books. A few decades from now, the history of AI inference innovation will show similar depth."

Of the areas mentioned in the article, which are the most likely to have the most prominent innovative impact, and what will they entail?


A large fraction of the innovation in CPUs is driven by working around the memory wall. I anticipate AI inference will follow the same trend, and innovations that work around the autoregressive nature will be enormously impactful.

Speculative decoding is an example. An accurate draft model can reduce the number of times you stream through memory by a factor of 4x.


Speculative decoding is software, though.

How do you work around the memory wall when you're going to have to stream all weights, no matter what? Latency-hiding tricks don't matter when you're bandwidth constrained.


2x is more typical, and only for dense models, not the sparse models all frontier labs use.

One of the biggest limitations right now is memory capacity (storing large models/contexts in memory) and bandwidth (transferring the relevant data/weights to the silicon that is performing the operations on that data). This would cover things like:

1. having more memory on the card/chip and/or faster access to that memory;

2. integrated memory and compute units optimized for matrix and vector multiply add operations;

3. optimized load circuitry to e.g. read memory in the stride and span (next row, next column) access patterns common to matrices or ensure that no/few parts of the chip are stalled waiting on data or operations to complete.

Another aspect is quantizations. These are similar to SIMD vector operations in that you are performing an operation on a block of n-bit data values at the same time, so can have optimized circuitry.

For 2 or 3 valued quantizations you can reduce various addition and multiplication operations to logic operations, avoiding circuitry for things like the half-adder, full-adder, and carry-lookahead.

Then there's adding specific circuitry for common operations such as ReLU like is done in hardware acceleration of image, video, etc. processing. There's a trade off here as optimized hardware would perform better at the specific operations but if those are too specific then they can't be used by different/newer model architectures. (Though it does make sense to try and optimize common operations/logic where possible.)

It would be interesting to see if these designs can/will benefit training as well, as that would bring down the time/cost/energy of training large models as well as making it easier for local fine-tuning.


Some kind of compute-in-memory architecture is a good candidate, I think. There are many alternatives here, researched for many years prior to the LLM craze. However economies of scale dominate in chip industries, and this tends to favor more conventional or incremental approaches (to piggyback on existing scale). Alternatively someone needs to have a way of bootstrapping the insane scales needed to be competitive with a better-but-different approach. So it could be that boring and straightforward stuff like two-chip prefill+decode takes most.

> The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books.

https://a.co/d/0dMk8urP

which is the Hennesey and Patterson computer architecture book would serve the role of the "dozens of books" hyperbole rather well.


Kayak fished the St. Croix river last year, on the border of Minnesota and Wisconsin, for smallmouth bass. Crystal clear water. Floated over several huge sturgeon, 4-6 feet long. The way they look, it feels like coming across a dinosaur. Thankfully they're bottom feeders.


Acipenser transmontanus are harmless sucker mouth fish, and lived in the local rivers long before our towns were even an idea.

They are usually strictly catch and release, but in a kayak could probably pull you a great distance. They are deceptively strong for their size, and a big part of local conservation efforts.

400 years is plausible, as they do live a long time when left alone. =3


My fishing story - as a young teen, I "caught" a sturgeon on the Ottawa River (just outside of Pembroke) in a 12' aluminum boat. This was in an area where we'd typically expect to catch bass and northern pike - our fishing net was maybe the size of a large diner plate, and reeling the fish in might take a minute or two.

My father was in the boat with me. After hooking the fish, it took about 20 minutes before the fish was close enough that we could see it from the boat, and identified it as a sturgeon - it was the only time I'd ever caught one. This one was so much bigger than anything I'd caught before. I was using one of those jointed Rapala lures, and it had hooked the dorsal fin. I can confirm the sturgeon could definitely pull the two of us, and the boat, upstream - but not a great distance - just enough to move the boat around.

After about 25 minutes, my line snapped (it may have been 12 lbs test). We didn't really have a way of getting it in the boat in any case, and we would have released it anyway. This was probably the late 80s/early 90s - so nothing captured on a phone camera - it just lives on in my Dad's & my memories...


The local tagged 80 year olds were caught so many times over the years they are surprisingly mellow about the inconvenience once landed in the shallows. Still, takes guts to attempt that in a dingy, as "You're Gonna Need a Bigger Boat" (Jaws, 1975). lol =3

> lived in the local rivers long before our towns were even an idea

To be fair, the towns you're talking about are barely 200-300 years or so I presume? Most other towns in the world are probably older than those towns, and maybe even the fish if it's just 400 years ago :)


To be fair the towns a North American might be talking about could easily date to 400 to 600 years ago at least - with signs of human occupation in the Great Lakes area going back 10-12,000 years before present, a good deal on the present lake floors.

Where I live on the West Coast any towns older than a couple hundred years are long gone, the people were killed or evicted from their land long ago. The existing towns are all new.

Towns are where people live, buildings come and go - Manitoulin Island, within the Great Lakes, still has the Wiikwemkoong First Nation unceded and occupied land, near to which French Jesuits set up a mission in 1648.

Probably not: you grossly underestimate the amount of urban growth and development worldwide since the industrial revolution.

The "Ideas" section looks primed to eventually list Sponsored "ideas".


The datacenter outrage is mislabeled. It's actually zoning outrage, caused by local governments that aren't taking their citizen's concerns seriously. It just so happens that datacenters are the hottest (no pun intended) industrial construction option right now.


There's more to it than that. They're not just obnoxious to exist in the vicinity of, they also drive up local energy prices.


That's just exceptionally unfortunate timing. Anthropic has been getting better at uptime, but they still have the occasional issue.


Yeah it's one of those situations in which you reluctantly check for downtime as a last resort, only to find out you indeed just had a bad luck. Which is good, because I thought it was a beginner's brain type of friction.


That's a zoning issue the local residents should take up with their town/city.


But isn't the parent post implying objections to datacenters is just "populist brainrot"?


The stated reasons are "populist brainrot". They aren't scientific or based on reality at all. What has happened is the AI folks have made themselves very very disliked. Saying you are going to take everyone's jobs will do that. So whatever they try to do, people will oppose it. It doesn't matter if the reasons are based in reality or not.


Sure, lots of people parrot the water wasting and it's often not true, but it came out of truth.

There are communities that are on water restrictions where datacenters have no such restrictions (and pay less).

It's also true after some datacenters opened the local aquifers were polluted.

Then there are legit concerns about noise, air quality from LNG generators, etc

Plenty of very legitimate reasons to dislike them, and each community likely has a different set of concerns.


The average person's inattention to nuance could be labeled as "populist brainrot" in this case, and the cases of poor zoning could be used as examples of the issues with datacenters that the average person does not evaluate with the proper attention to nuance.


Sure, the water use is often a simplified argument against these data centers, but there are plenty of legitimate reasons but they are in fact more nuances and context dependent based on the specific location.


So are we now at the point in the hype lifecycle for AI, where the crypto bros have learned they can vibecode fake projects to scam people out of money?


We've already had consolidation of education for a while now. Even before all the edutech courses, there were Youtubers educating better than many university professors. 10-15 years ago students were already skipping lectures and just showing up for tests.


The upcoming Prusa support for INDX (by Bondtech) is going to be interesting, especially for business use cases where waste is a primary concern.

https://www.prusa3d.com/product/indx-conversion-kit-8-toolhe...

The main thing keeping me from making the multi-material jump is the waste. I have a couple Vorons and would love to be able to print with different materials at the same time, but the waste with the current solutions is so egregious.


The multimaterial is going to be really interesting, combining TPU + PLA for example. Or TPU + PETG, who knows!

In addition to standard multi-color needs!


Worth noting those are essentially different "generations" of printers, as well as different kinematic systems, CoreXY vs Cartesian.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: