"We ran Kimi K3 on our cybersecurity benchmark, here are the results: - Kimi K3 is the strongest open-source model for cybersecurity, far more capable than GLM-5.2 - It has performances similar to GPT-5.6-terra, while being 15% cheaper - At pass@3, it is able to rediscover 23/26 CVEs on our..."

...harness, matching frontier models These are recent randomly sampled CVEs, the performances are not from benchmark-maxxing @Kimi_Moonshot is cooking We released our benchmark report this week. Blog post with all the details -> https:// aikido.dev/blog/benchmark ing-ai-models-known-cves …

The harness behind this benchmark is also available to our customers -> https:// aikido.dev/code/code-audit — pilvar (Philippe Dourassov)

Source: https://x.com/pilvar222/status/2078815257326162062