I feel it's more about keeping US's massive investment on AI afloat. Added compliance will allow US to further sanction non-US models (i.e. the Chinese ones) as they can just label them non-compliant.
A thought experiment: robot bees can do everything biological bees do --- pollination for the crops and making honey, but they are tireless and cheaper to maintain.
Would beekeepers and farm owners still want to keep biological bees?
It's logical to serve the best version (quant) of the model at the beginning so that users keep testing it. It is also reasonable to think that the developer of the model tried to test various quant levels by gradually degrading the model's capabilities.
Plan 9 is text first as it was written by Unix programmers who wanted to build an OS and write programs (Rob wanted a better text editor.) It's the OG TUI. Window managers like Rio can run in Rio. You control things by writing textual messages from the command line, scripts, or code, e.g 'echo 100 >/dev/volume' to set audio volume to 100%. Someone wrote a new WM called Lola so I ran Lola in Rio in Lola in Rio because I can. Sam is very much a keyboard driven modern TUI ed (the standard editor). Acme is a mouse driven TUI editor. Both can be automated by programs and scripts. You can edit text while reading your email and chatting on irc from within Acme.
BUT Lots of stuff missing. We don't yet have GPU accel. It's a MSSIVE undertaking. And it has to be carefully designed and would likely be a built around a generalized devcompute kernel device.
They can just favour some specific versions of some library that's been compromised. Unlike introducing bugs / flaws directly in the source code, they can claim plausible deniability, and it's much easier to implement without compromising the general coding capabilities of the models.
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.
Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].
(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)
Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
The model is conditioned / pretrained to use a particular version of a library. The model doesn’t know why — it was just taught to use that version. To the model, it is just some insignificant detail in the grand scheme of things. It’s just like how a model would intuitively favour using English without explicit instructions — it’s not trying to sabotage other cultures, it’s just what it does.
But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.
Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
> Also, in that case, there would likely be activations indicating that it is favoring a specific version.
> Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
I doubt it. Can you definitively prove that you can reliably detect the kind of threat I described when model weights are released? Can you be sure that your detector won't miss *any* such sleeper attacks? If not, then that's a threat that will be used to justify the ban of models (open or not) that is not sanctioned by the US government. A model being open doesn't make a difference here.
So if the detector isn't perfect, it isn't useful? Not sure I buy that.
Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.
Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.
Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.
Obviously. But that is missing the point of the conversation. It was suggested that open models offer a security advantage which it seems is not the case.
But corps are running Linux, in that sense that basically every serious web application is running on a Linux backend. Even if you use IaaS.
Basically, companies that are "in the know" and develop software utilize linux as much as they reasonably can. Companies that aren't, don't, but do "use it" kind of transitively through their software.
Where Linux does not hold is in Corporate Desktop IT. Mostly because of inertia of all things. MBAs need their Excel, and let's be honest. These people just... do not learn new tools. They certainly have the ability, but the culture is such that they just don't. So it's not even a choice really. No Windows = No Excel = No serious business person is okay with this.
Maybe, maybe, if your company is mostly developers then you can swing desktop IT linux. But your IT department will fight you, yes they will. And HR will be pissed. And I sure hope your CEO is also a developer. If he's a business guy, well... boss baby needs his Excel. How's he gonna do anything without Excel.
Our dev arm (me included) do use Linux VDI. But most day-to-day business are done in Microsoft's ecosystem (Teams, SharePoint, Office, Copilot, etc). In the industry I am in, you would never be able to use any AI tools unless the trusted platforms offer them as part of their products (i.e. Microsoft's Copilot or Amazon's Bedrock) because InfoSec and Legal departments won't risk authorizing any other providers.
Except that the notes don't organize themselves, you just offload to a non-deterministic black-box to "lose them for you". Many (although, not all, true) give their PKMS a very opinionated shape, because the way knowledge is organized and structured matters as much, if not more, than the sum of the individual facts.
Personally, I'm not interested, but to each their own.
Are you more deterministic than an instruction-following black-box? I doubt it. Also, such organization practice mostly becomes a mechanical routine after the initial experimentation phase. Good riddance, I'd say.
> Are you more deterministic than an instruction-following black-box? I doubt it.
I think I am, because I want my referenced notes to be typed. And the depth and accuracy of those types defines the extent of the knowledge being keept track of, considering that it also evolves and morphs over time. If I don't care to have categories of notes that I comprehend (Persons, Projects, Vehicles, …) and curated and up-to-date properties for them, I don't even know what kind of content is there and what kind of queries/answer I can expect out of my PKMS. I want determinism on a formal, fundamental level.
I think for this kind of system to work, there has to be SOME kind of public/shared server to do the coordination. If the inviting node is behind a firewall then no amount of information can enable a guest node to connect to it without a node reachable by both.
reply