Hacker Newsnew | past | comments | ask | show | jobs | submit | _air's commentslogin

It's funny to pick a birth year from our era and think of how history will represent us:

Martha was born in the countryside near Short Pump, in 1993 CE. Her village holds about 300 people. She makes her home in a timber-framed detached house. Martha has neither brothers nor sisters. She eats mostly wheat bread and maize, pork and dairy, and potatoes and garden vegetables, with apples in season, and coffee, soft drinks and snack foods too. She knows her letters. Her height is about 171 cm (5 ft 7 in).


An “Eastern Woodlands” person born 1989 is recorded to have experienced a “volcanic winter” and “dimmed sky” at age 2 as a result of the eruption of Mount Pinatubo.

This is an… interesting contextualization.

I am wondering how many historical situations are sexed up to be something significant, yet were nothing to write home about in the moment. All it took was a historian with an active imagination to bring it to the forefront.


Switzerland is ranked 67th in country population density. For reference, the United Kingdom is ranked 48th and the United States is ranked 183rd.

https://en.wikipedia.org/wiki/List_of_countries_and_dependen...


I wonder if that number can be adjusted based on the amount of arable land, or based on the ease of construction (quite the nebulous term here admittedly). The number of mountains presumably makes this hard to compare.


India is 31, Netherlands is more dense than India. Would not have expected that, but then I remember that India has a massive desert, and the Himalayas. So I guess it makes more sense now.


America is soooo big and soooo sparsely populated


That is an utterly meaningless statistic. Canada, with four times the population, ranks #233, because most of the country is uninhabited / uninhabitable.


Population weighted density is a better metric for this use case. It's more stable than population density when adding large areas of sparsely populated land, because the denser, more highly populated areas are more heavily weighted. It shows, roughly, the density experienced locally by the average person in some region.

https://en.wikipedia.org/wiki/Population_weighted_density

The problem is it's difficult to compare across polities because nobody will agree on the right granularity of parcel size to use (and indeed, it is not really obvious what the right granularity is, and choice of parcel size can drastically change the number).

It's similar to the metrics of "average class size" vs. "student-weighted class size". https://allenschwenk.wordpress.com/wp-content/uploads/2023/0...


Whichever of the reasonable parcel sizes you choose, it's still miles better than the population density based solely on the territory of a country.


> That is an utterly meaningless statistic

It's very meaningful, when the main argument is population overcrowding.


The entire human population can fit within Los Angelas. It’s not a good metric in general. Pressure on public services, resources and housing is far more useful


The entire human population could stand there: each person would have roughly 1.5 to 5.4 square feet of space—less than the size of a single chair. If you actually did get all of humanity in LA it would instantly break out in war.


It's completely useless, because you're comparing country area level.


Read the rest of the post.


Neat! Here’s orange guy In “camouflage”: https://imgur.com/a/9xpEG2a


Yeah, seems like a standard supply level to me.

"California’s inventory has averaged just over 20 days of supply over the last five years (2019–23), compared with the U.S. average of 21.6 days."

https://www.eia.gov/todayinenergy/detail.php?id=63944


I think they mean that it's 4-6 weeks until they hit zero, accounting both for stored products and the current rate of production/imports.


Usually when something like this is reported, it is because of some other milestone.

Like, they have 6 weeks, on hand, in tanks already delivered.

But, all of the ships in-bound are now done.

After the war started, there was a record number of ships, already filled, already in-transit. But now they have all reached their destinations. So there is no more incoming.


Grab a metal detector to the left of the volleyball court. Then hold it near the bottom left edge of the court


Opus 4.5 was released on Nov 24 last year. It’s only been 5 months!


Wow you're right, okay not so bad then.

That brief two week period when Opus could eat entire tickets was simultaneously fantastic and a bit alarming


Black has no legal moves because of the knight but they aren't in check


"The amazing champions inspire boldly like brilliant genius and incredible legends admire splendid talent."

Hard to forget a sentence like that!


Do we have metrics for the uptime of other major services? Would be interesting to see if this is just a GitHub problem or industry-wide.


Bitbucket Cloud incident history: https://bitbucket.status.atlassian.com/history

Though I will be the first to say I don't fully trust it based on the flakey git clone errors we see in CI.


This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains


Only way to have hardware reach this sort of efficiency is to embed the model in hardware.

This exists[0], but the chip in question is physically large and won't fit on a phone.

[0] https://www.anuragk.com/blog/posts/Taalas.html


I think you're ignoring the inevitable march of progress. Phones will get big enough to hold it soon.


Instead of slapping on an extra battery pack, it will be an onboard llm model. Could have lifecycles just like phones.

Getting bigger (foldable) phones, without losing battery life, and running useable models in the same form-factor is a pretty big ask.


I think the future is the model becoming lighter not the hardware becoming heavier


The hardware will become heavier regardless I'm afraid.


Good. It's ridiculously tiny and lightweight these days.

Especially with phones; the first thing everyone does after buying their new uber thin iPhone is buying a case for it, which doubles its thickness.


I think for many reasons this will become the dominant paradigm for end user devices.

Moore's law will shrink it to 8mm soon. I think it'll be like a microSD card you plug in.

Or we develop a new silicon process that can mimic synaptic weights in biology. Synapses have plasticity.


One big bottleneck is SRAM cost. Even an 8b model would probably end up being hundreds of dollars to run locally on that kind of hardware. Especially unpalatable if the model quality keeps advancing year-by-year.

> Or we develop a new silicon process that can mimic synaptic weights in biology. Synapses have plasticity.

It's amazing to me that people consider this to be more realistic than FAANG collaborating on a CUDA-killer. I guess Nvidia really does deserve their valuation.


> bottleneck is SRAM cost

Not for this approach


That's actually pretty cool, but I'd hate to freeze a models weights into silicon without having an incredibly specific and broad usecase.


Depends on cost IMO - if I could buy a Kimi K2.5 chip for a couple of hundred dollars today I would probably do it.


I mean if it was small enough to fit in an iPhone why not? Every year you would fabricate the new chip with the best model. They do it already with the camera pipeline chips.


Sounds like just the sort of thing FGPA's were made for.

The $$$ would probably make my eyes bleed tho.


Current FPGAs would have terrible performance. We need some new architecture combining ASIC LLM perf and sparse reconfiguration support maybe.


Wouldn't it be the opposite of freezing weights?


On smartphones? It’s not worth it to run a model this size on a device like this. A smaller fine-tuned model for specific use cases is not only faster, but possibly more accurate when tuned to specific use cases. All those gigs of unnecessary knowledge are useless to perform tasks usually done on smartphones.


Probably 15 to 20 years, if ever. This phone is only running this model in the technical sense of running, but not in a practical sense. Ignore the 0.4tk/s, that's nothing. What's really makes this example bullshit is the fact that there is no way the phone has a enough ram to hold any reasonable amount of context for that model. Context requirements are not insignificant, and as the context grows, the speed of the output will be even slower.

Realistically you need +300GB/s fast access memory to the accelerator, with enough memory to fully hold at least greater than 4bit quants. That's at least 380GB of memory. You can gimmick a demo like this with an ssd, but the ssd is just not fast enough to meet the minim specs for anything more than showing off a neat trick on twitter.

The only hope for a handheld execution of a practical, and capable AI model is both an algorithmic breakthrough that does way more with less, and custom silicon designed for running that type of model. The transformer architecture is neat, but it's just not up for that task, and I doubt anyone's really going to want to build silicon for it.


> Realistically you need +300GB/s fast access memory to the accelerator, with enough memory to fully hold at least greater than 4bit quants.

The latest M5 MacBook Pro's start at 307 GB/s memory bandwidth, the 32-core GPU M5 Max gets 460 GB/s, and the 40-core M5 Max gets 614 GB/s. The CPU, GPU, and Neural Engine all share the memory.

The A19/A19 Pro in the current iPhone 17 line is essentially the same processor (minus the laptop and desktop features that aren’t needed for a phone), so it would seem we're not that far off from being able to run sophisticated AI models on a phone.


Agree with the first part - but I can run GPT OSS 20b, a highly capable model on my laptop with 32GB of RAM at speeds that for all practical intents is as fast as GPT-5.4 and good enough for 90%+ of non-technical use cases.

As such I can't agree with "The only hope for a handheld execution of a practical, and capable AI model is both an algorithmic breakthrough" - we are much closer than 15/20 years to get these on a phone


With this work you can run a medium-sized model like GPT OSS 20b at native speed even while keeping those 32GB RAM almost fully available for other uses - the model seamlessly starts to slow down as RAM requirements increase elsewhere in the system and the fs cache has to evict more expert layers, and reaches full speed again as the RAM is freed up. It adds a key measure of flexibility to the existing AI local inference picture.


KV-cache is still quite small compared to the weights. It can stay in memory for reasonable context length, or be streamed to storage as a last resort. This actually doesn't impact performance too much, since we were already limited by having to stream in the much larger weights.


This should be the top comment


A long time. But check out Apollo from Liquid AI, the LFM2 models run pretty fast on a phone and are surprisingly capable. Not as a knowledge database but to help process search results, solve math problems, stuff like that.


It will never be possible on a smart phone. I know that sounds cynical, but there's basically no path to making this possible from an engineering perspective.


No one needs more than 640K!


Quantum computing is right around the corner!


This comment will age well.


Is 100 t/s the stadard for models?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: