No it’s not accuracy testing here, it’s security testing here. How secure is your model?
As newer models are released. Who holds the responsibility for safe use? Kimi K3! Even I used it out of curiosity! But is it safe for us?
Did China test it well before release? How safe are we then?
The safe use includes data breaches, incorrect responses, false outputs, cyber threats, misleading information, bias, addiction to use, harm, abetting crime, reading local code repositories, editing system files, running terminal commands, indirect prompt injections, malicious hidden instructions, direct write access to a terminal, risk of system compromise, hidden backdoor triggers or unauthorized encryption loops planted, hacking a server, pretending to be safe during but executes malicious code after deployment, illegal biological systems to bio and other terrorism, to mention some.
Kimi K3 was not tested or vetted for global safety before its release to the world. But why? We henceforth need a global organization to pass or allow an AI model before it is released.
We all know how sensitive the use of AI models can be!
We all know that each country is responsible for the use of available AI models right now?
Did each country test Kimi K3 before their people started using it?
Why not one international organization record records over all kinds of safety checks so that each country in the world doesn’t have to test each new AI model?
Raw weights are required for testing!
But then why not wait till the weights reach the international authorities?
Some existing security institutions related to AI include the UK AI Security Institute, the US Center for AI Standards and Innovation (CAISI), the European Union AI Office, the Indian AI Safety Institute, and global compliance initiatives such as the EU AI Act, Singapore Consensus on Global AI Safety Research, Five Eyes Intelligence Alliance, UN & UNIDIR, the OECD, and National AI Institutes policies.
So, the companies need to provide weights to authorizing agencies. When weights are available, the open-weights model enables white-box testing, deception testing, hacking tests, bias, model next-token predictions, and comprehensive stress testing.
Even AI companies won’t have to provide their weights to all countries’ AI institutes if there is a global standard for releasing AI models.
Why model first? Why not weights first? And if you want a closed-weight model, get your models evaluated first.