That’s if you have disparate prompts and don’t want to actively select a model. If you’re developing a pipeline, it’s a terrible idea. You want to choose a model, validate it, then stick to it.
I have, but honestly that was exactly what I am not looking for. Sometimes it picked a model for a Dutch email that totally does not support Dutch. Other times it would work fine. It's just a layer of indeterminism I wasn't looking for.
150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.
Ninfer is compatible with a nvfp4 model for the 27b. Also nowadays I prefer to use the byteshape one, I get less loops, and I am not sure if I really see a difference in speed or quality. Pure vibe agentic coding on a C++ codebase or ocaml one, ocaml one has codex as reviewer as I am more interested by that project, the other is more for fun.
We have not used the wero app directly but my bank app which does not require Google play but has integrated the use of wero to send or receive money from or to bank accounts. My bank is LCL in France
https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...
reply