I wanted a real production test for Kimi K3 rather than another benchmark comparison, so I used it while building a 75-second deep-sea explainer through eight revision passes.
This was not text-to-video. Licensed footage and credited scientific images formed the photographic layer. The depth gauge, animated density and temperature curves, trench profile, and Burj Khalifa comparison were coded in HTML/SVG/JS and rendered frame by frame in Playwright at 4K.
ffmpeg handled the deterministic production path: normalization, grading, concat, overlays, subtitles, TTS placement, music mixing, exact timing, and final encoding. K3 assisted the code and stayed useful as changes moved between graphics, timing, audio, and compositing.
The strongest result was not autonomy by itself. It was keeping the project editable through repeated revision. Human review still decided pacing and hierarchy, while ffprobe, extracted frames, and level analysis supplied external verification.
What kind of real-world task do you think exposes a frontier model more clearly than its benchmark scores?