Yes, I read the 5-7 paragraph article :). I'm not sure what part of it you could possibly consider a rebuttal...? It's just talking about one devs decision to keep working on their purist software despite their anxiety; basically no broader points are made at all, and certainly none are made with any sort of rigor. It didn't seem like that was his goal.
They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
It’s not hypothetical. Magic strings are a known and implemented feature for standard model interaction. Nearly impossible to detect unless you know where to look with current technology.
Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
They can just favour some specific versions of some library that's been compromised. Unlike introducing bugs / flaws directly in the source code, they can claim plausible deniability, and it's much easier to implement without compromising the general coding capabilities of the models.
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.
Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].
(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)
Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
The model is conditioned / pretrained to use a particular version of a library. The model doesn’t know why — it was just taught to use that version. To the model, it is just some insignificant detail in the grand scheme of things. It’s just like how a model would intuitively favour using English without explicit instructions — it’s not trying to sabotage other cultures, it’s just what it does.
But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.
Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
> Also, in that case, there would likely be activations indicating that it is favoring a specific version.
> Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
I doubt it. Can you definitively prove that you can reliably detect the kind of threat I described when model weights are released? Can you be sure that your detector won't miss *any* such sleeper attacks? If not, then that's a threat that will be used to justify the ban of models (open or not) that is not sanctioned by the US government. A model being open doesn't make a difference here.
So if the detector isn't perfect, it isn't useful? Not sure I buy that.
Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.
Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.
Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.
Obviously. But that is missing the point of the conversation. It was suggested that open models offer a security advantage which it seems is not the case.
Right, the opinion shifting to the people being unworthy of uncorrupted data unless they’re wealthy enough to pay the artificial extortion price for integrity.
What you're actually saying is that I shouldn't be allowed to determine what kind of RAM my application needs. Instead, you are saying that someone with a gun should dictate this choice to me.
That's what we mean when we say that people with your point of view are promoting nanny-statism. It amounts to infantilization of adult consumers and subsequent disempowerment.
Occasionally such legislation can be justified, as when genuinely hazardous products are involved, or products whose use involves wide-ranging externalities, but usually not. People can and should be educated -- and expected -- to make their own decisions about things like whether they need ECC RAM.
How do YOU propose doing this when most computer "experts" have never used ECC RAM and when it's only available in high end specialist systems that aren't on sale in regular computer stores?
I think my knowledge of the topic is pretty good but I learned from this discussion that almost all AMD CPUs support ECC RAM and you apparently just need a motherboard to match. When accompanying people who were buying "gamer" PCs (after they rejected my recommendation for a second hand workstation grade computer with ECC RAM) I never witnessed store staff saying "hey you could get a motherboard with ECC support".
It's simply none of my business whether or not someone does the homework to understand what kind of RAM they should buy. I bothered to inform myself; they can, too.
Next up. CPUs that will spit out the wrong calculations 1/1M times. You’ll have to pay more for the enterprise version to get uncorrupted processing. The spare traces cost 5 cents, but we can charge 20% more for the privilege instead.
Just because you can allow it, doesn’t mean you should. I don’t want to live in a world where you need a PhD to avoid being scammed by increasingly sophisticated ploys. Data isn’t free anymore and the right questions are becoming harder to ask.
It's none of their business whether or not others choose to use ECC. It's absolutely his business whether or not authorities force them to use ECC RAM or else.
> People can and should be educated -- and expected -- to make their own decisions about things like whether they need ECC RAM.
So you think its realistic to expect the average consumer to assess the pros and cons of using ECC RAM?
We could solve this by better information which addresses your concerns about it being nanny statism. We would need it to be something that consumers understand - e.g. We could require all computers without ECC RAM to have a warning label saying something like "this computer uses cheap memory so is prone to crashing and losing your data". I strongly suspect the impact of that would be very similar to that of just banning non-ECC RAM.
I would also argue that having reliable RAM is a reasonable extension of existing legal requirements for consumer products: e.g. that they are of satisfactory quality and fit for purpose (to use the current UK terminology).
There are two issues with nanny statism, the first is poorly informed kneejerk overreactions, the second is immediately jumping to heavyhanded policies. In your example even the absurd warning label is significantly less of an issue than flat out banning non ECC RAM.
reply