GPT-5.6 Sol Achieves 13.0% Score, New SOTA in Vals AI Space Race Against Kimi K3

Image for GPT-5.6 Sol Achieves 13.0% Score, New SOTA in Vals AI Space Race Against Kimi K3

Vals AI's "AI Space Race" concluded with OpenAI's GPT-5.6 Sol achieving a 13.0% score, establishing a new state-of-the-art (SOTA) in the challenging Kerbal Space Program simulation. The competition pitted the US-developed GPT-5.6 Sol against China's Kimi K3, which scored 5.2%, drawing over 10,000 live viewers to the event. According to Vals AI, the "Final score of the AI Space Race: GPT-5.6 Sol πŸ‡ΊπŸ‡Έ 13.0% (new SOTA) vs. Kimi K3 πŸ‡¨πŸ‡³ 5.2%."

The "AI Space Race" was designed as a rigorous test of long-horizon AI capabilities, tasking models with independently building a space program from scratch within Kerbal Space Program. This involved designing, launching, and landing spacecraft repeatedly, without human intervention or pre-built assets, with each model given up to five days to complete the complex challenge. The objective was to measure an AI's ability to plan, execute, learn from failures, and adapt.

OpenAI's GPT-5.6 Sol, the flagship model of the GPT-5.6 family, demonstrated superior performance in this specific simulation. Recently released for general availability, GPT-5.6 Sol is recognized for setting new standards in intelligence and efficiency across various domains, including coding, knowledge work, and scientific reasoning, according to OpenAI. Its robust computer use and design judgment capabilities contribute to its polished collaborative performance.

Kimi K3, developed by Chinese artificial intelligence startup Moonshot AI, is a 2.8 trillion-parameter open-weight model. While it secured 5.2% in the "AI Space Race," Kimi K3 has been noted for its competitive performance in other benchmarks, often challenging or surpassing leading proprietary models in specific coding and agentic tasks. Moonshot AI positions Kimi K3 as a significant step for open-weight models in the global AI landscape.

This competition underscores the intensifying global rivalry in artificial intelligence, particularly between US and Chinese developers. While GPT-5.6 Sol secured the win in this particular long-horizon test, Kimi K3's overall emergence signifies that open-weight models are increasingly narrowing the capability gap with closed-source counterparts, as observed by various industry analysts. The rapid advancements highlight the dynamic nature of the AI frontier.