"We ran Kimi K3 on our cybersecurity benchmark, here are the results: - Kimi K3 is the strongest open-source model for cybersecurity, far more capable than GLM-5.2 - It has performances similar to GPT-5.6-terra, while being 15% cheaper - At pass@3, it is able to rediscover 23/26 CVEs on our..." Reddit
...harness, matching frontier models These are recent randomly sampled CVEs, the performances are not from benchmark-maxxing @Kimi_Moonshot is cooking We released our benchmark report this week. Blog post with all the details -> https:// aikido.dev/blog/benchmark ing-ai-models-known-cves …
The harness behind this benchmark is also available to our customers -> https:// aikido.dev/code/code-audit — pilvar (Philippe Dourassov)
