I suspect the only reason those alternative providers have better up-time and more generous quotas, is because they don't have nearly the same amount of demand. Notice that Deepseek recently had to increase their pricing, once it gained it popularity.
Demand > supply. It's impressive that customers have not migrated en masse to other providers, given the frequency of these outages. Perhaps switching costs are greater than some would believe. Or, qualitative differences between models continue to exist, despite matching on public benchmarks.
That's quite a reach. More likely someone merged and deployed their vibe-coded PR and is now figuring out how to bring the service back up.
If there's a lot of demand for rollercoaster rides, the rollercoaster will not stop operating; instead the queue of people in front of it will increase.
It's not like a bridge or elevator where we have a certain number of people that can use it, and if one more person joins, the whole structure breaks apart and everybody perishes.
Those guys are running a website that provides an interface for some specific hardware. Just like file hosting providers back in the day selling their terabyte-sized hard disks in 100MB-increments.
> instead the queue of people in front of it will increase.
This can ultimately result in the system breaking. I don't think Anthropic engineers are that much worse than their peers, such that they are 10x more prone to causing outages due to bad deployments.
>> Perhaps switching costs are greater than some would believe.
I haven't switched because there's nothing to switch to that is anywhere as good. I've been making dedicated attempts at using Sol but it falls short, despite what some people claim.
I did migrate 80% of my tokens. But for some tasks claude models are still the best.
Easy workaround is to work outside US peak hours (europe morning). I love this outages, I am hardly affected, and weekly reset usually promptly follows!
Do you really think it’s a coincidence claude outages always happen when US and EU workdays overlap?
Anthropic is a trillion dollar company and employs way more skilled high-paid engineers than you are btw, you really think it’s a systems issue not capacity
Not just each failed attempt, but every transaction. Previously, Radar only checked the card on a subscription at first use. It will now check every transaction unless you opt-out. So not only do you have to opt-out of Radar Standard, you also have to opt-out of having every subscription transaction run through Radar.
They are doing both. Distilling Mythos down to affordable models, so they can continue to fund the business. And training Mythos level models at the high-end, to expand the frontier.
He didn't. That's why the article exists. You have to do the science to see if it does.
He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?
Either you've memorized the outline (or detailed component shapes) of, say, a horse, or you haven't. Memorizing the outline of a pelican isn't going to help you with the horse.
You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.
reply