In 904 DeepSWE-Durchläufen führt GPT-5.6 Sol bei pass@1, aber Kimi K3 schlägt bei pass@4 mit 2,8x mehr gelösten Aufgaben pro Dollar.
Das Routing zwischen den Modellen erreicht ~85,6 % und ermöglicht kosteneffiziente hybride Coding-Agenten.
In 904 DeepSWE rollouts, GPT-5.6 Sol leads pass@1, but Kimi K3 dominates pass@4 with 2.8x more solves per dollar.
Model routing between the two reaches ~85.6%, opening the door for cost-sensitive hybrid coding agents.