SpaceX AI unveiled Grok 4.6 recently, marking a major milestone in the rapid evolution of its large language model ecosystem. Independent evaluation from Artificial Analysis shows Grok 4.6 achieves an intelligence index of 61 points, on par with ChatGPT-5.6 Sol and closely trailing Fable 5 Max at 62 points. The model carries a per-task cost of $0.84, delivering a clear cost advantage compared to Sol ($1.23) and Opus 5 ($2.34). When intelligence, inference speed and operational expenses are considered comprehensively, Grok 4.6 ranks first overall. Musk has expressed optimism that Grok will surpass all competing models in the near future.
1. The Remarkable Turnaround of Grok
It is striking to recap how drastically Grok’s fortunes shifted in just five months. Earlier this year, Grok 4.3 lagged far behind leading models from Silicon Valley competitors. The team faced an unprecedented crisis: apart from Elon Musk himself, every one of xAI’s 11 founding members departed, and Musk publicly acknowledged that xAI’s initial construction attempt had failed and required a complete restart. Grok was then incorporated into SpaceX under the new brand SpaceX AI.
During the same period when rivals Gemini fell into stagnation, two senior product leaders from Cursor, Andrew Milich and Jason Ginsberg, joined xAI and reported directly to Musk. In April, SpaceX and Cursor signed a cooperation and joint training agreement. Four days after completing its largest IPO in history, SpaceX announced the acquisition of Cursor. The core logic behind this strategic move was clear: software engineering represents a vital application scenario for AI. High-quality structured data generated by developers, short feedback cycles and high-frequency usage create a robust positive feedback loop.
In this workflow, developers use Grok for programming tasks, while Cursor records operational trajectories, failure points and areas requiring model improvement. Engineers feed these real-world action traces back into model training and fine-tuning datasets. Reinforcement learning and continuous iterative training steadily lift model capability. The results arrived faster than expected. In July, Grok 4.5 achieved near-parity with GPT-5 in coding performance, and one month later, Grok 4.6 was officially released.
2. Analogous Industry Comebacks: The Template Validated by Anthropic
This growth path is not unique to xAI. Anthropic proved a similar strategy can deliver explosive progress. Two years ago, when DeepSeek made waves across the industry, Claude was still viewed as a secondary player. The launch of Sonnet 3.7 alongside Claude Code, an AI coding assistant, kickstarted sustained rapid advancement, and within less than a year, Claude established itself as a direct competitor to ChatGPT.
Coinciding with Grok’s turnaround, Tencent replicated a comparable growth trajectory with Workbuddy and Hunyuan. These parallel developments demonstrate that foundation model advancement is no longer driven solely by raw parameter scaling. Closed-loop learning built on real user scenarios has become a viable, repeatable route to catch up with industry frontrunners.
3. Tencent’s Co-design Loop: Workbuddy and Hunyuan Synergy
Hours before the launch of Grok 4.6, Tencent highlighted Workbuddy and the Co-design Loop mechanism during its earnings call. Similar to Musk’s heavy investment in Cursor, Workbuddy originated from a small internal team within Tencent. Building upon Codebuddy, the previous AI programming agent, the product grew rapidly following the industry-wide AI agent boom in March. By June, Workbuddy reached 2.097 million daily active users, securing the top position among domestic office AI agents, surpassing the combined traffic of its second and third-ranked competitors.
Unlike general-purpose chatbots, office AI agents operate within complex, long-duration workflows. Users generate massive volumes of authentic work records: drafting presentations, organizing documents, processing spreadsheets, researching materials, writing and refactoring code, and invoking auxiliary tools. In these scenarios, failed attempts often deliver more valuable signals than benchmark errors. Each user interaction, whether successful or unsuccessful, can be extracted as reusable agent trajectories. These authentic task records, operation sequences and outcome cases train the model to acquire practical capabilities that synthetic benchmarks cannot replicate.
Tencent’s chief scientist Zhang Yaqin has long advocated this direction. More than one year ago, he noted in AI Half-Life that reinforcement learning with human feedback (RLHF) reaches a performance ceiling. To accurately evaluate real-world model performance, developers must design assessments based on authentic mission scenarios. After rebuilding Codebuddy and launching Workbuddy, Zhang formally proposed the Co-design method to achieve coordinated optimization between models and product scenarios.
During the earnings briefing, Zhang elaborated on the mechanism: “By continuously deploying models into production environments and feeding real-world user feedback into training, the model-product Co-design loop helps validate model accuracy, identify edge cases and enable faster iteration.” This mechanism constitutes the core driver of Hunyuan’s rapid progress.
At the end of April, Tencent unveiled Hunyuan Preview, followed by the official launch of Hunyuan 4. The model leads domestic alternatives in agent execution and reasoning benchmarks. With only 295B parameters, it achieves performance comparable to models with substantially larger scales. Zhang Yaqin explained that capability gains are most visible in practical experience, covering reasoning, long context, programming, office automation, financial modeling and front-end development. These use cases align precisely with Workbuddy’s application scope.
A mutual reinforcement cycle has formed between Workbuddy and Hunyuan: stronger models enhance user experience, attracting more active users and generating higher-quality data. The richer dataset further elevates model capacity, which in turn drives user growth. Zhang confirmed that Hunyuan 4 will be followed by larger-scale iterations, targeting state-of-the-art (SOTA) performance.
Shortly after Tencent shared its roadmap, Musk released Grok 4.6 and publicly set a SOTA target. He described SpaceX AI’s training corpus as uniquely powerful, and stated he would welcome any model that outperforms Grok on practical engineering tasks. Musk has repeatedly emphasized turning models into ubiquitous consumer applications, a goal shared with Tencent as both companies compete in the consumer AI space.
4. The Value and Far-Reaching Influence of the Co-design Loop
The emergence of the Co-design loop offers a viable path for latecomer foundation model teams to narrow the capability gap. Meta recently launched Muse Codem, a new programming agent built on Muse Spark 1.2. Meta adopted a differentiated strategy by releasing a standard edition and a contributor edition. The contributor version costs 95% less than the standard release, and allows developers to contribute programming data to support ongoing model training.
Compared with the unrestrained capital competition among Silicon Valley giants, domestic AI companies focus more on balancing input and commercial output. The Co-design loop carries greater strategic weight here. Authentic user data continuously improves model performance, and better models attract more users, creating a self-reinforcing network effect. Higher user retention lifts monetization efficiency, especially in vertical productivity scenarios such as coding and office automation.
After Workbuddy launched, Tencent rolled out iterative upgrades and integrated hundreds of auxiliary tools, enabling document parsing, meeting summarization and other capabilities inside the agent framework. The platform aims to collect high-quality real-world task samples, running a continuous execution-validation-training cycle to boost model performance and capture sustainable commercial returns from productivity scenarios. Positive financial signs have already emerged.
Tencent’s financial report shows Workbuddy-generated token consumption has maintained sustained growth. Revenue from this business segment saw a 176% year-on-year increase in Q2, achieving positive cash flow. The management chose to reinvest earnings rather than extract immediate profits, which could have delivered a 30% direct profit boost. Zhang Yaqin explained that Tencent is executing a long-term strategy, prioritizing resource allocation to Hunyuan and Workbuddy and related AI applications.
The company’s leadership believes sustained advantages do not stem from one-off model upgrades, but from replicable iterative systems verified by real-world products. After years of intensive competition, large model teams are gradually reaching a consensus: only organizations capable of building a functional Co-design loop can unlock sustainable growth.
Teams operating multiple model services can simplify unified access control and routing management via 4sapi, creating consistent traffic governance for internal and external model endpoints. This type of layered infrastructure complements the co-design workflow by stabilizing data collection pipelines for model iteration.
5. Industry Outlook: The End of Pure Benchmark Competition
The breakthroughs achieved by Grok and Tencent’s Hunyuan ecosystem mark a clear shift in the large model industry. For a long time, model competition centered on benchmark scores and parameter scale. Now the industry focus is shifting toward closed-loop systems that connect models, end products and real user feedback.
Traditional offline training relies heavily on static, pre-collected datasets. The Co-design loop establishes a continuous flow: end users generate task data, the system extracts high-quality trajectories, the model absorbs new knowledge through retraining, and upgraded models return to products to serve users. This creates a perpetual growth engine.
This paradigm lowers the threshold for catching up. New entrants no longer need to compete head-on with leaders on initial dataset size and training computing resources. As long as they can launch viable application products to obtain user interaction data, they can gradually narrow the capability gap. However, the model also raises new challenges: companies must simultaneously master foundation model training, agent engineering and consumer product operation, a comprehensive capability set that many teams lack.
For enterprise clients, the trend brings tangible benefits. Models optimized through real scenario iteration deliver more reliable performance in practical work, rather than only excelling in standardized testing. In programming, office automation, financial analysis and enterprise service workflows, scenario-tuned models produce more consistent, actionable outputs.
Conclusion
Grok 4.6’s parity with ChatGPT-5.6 Sol, alongside the rapid progress of Tencent Workbuddy and Hunyuan, highlights a transformative industry trend. The Co-design Loop, which tightly integrates foundation model training with end-user product scenarios, has become a critical accelerator for latecomers to challenge established leaders.
The era where large model competitiveness is judged solely by benchmark metrics is fading. Sustainable advantages will belong to teams that can build complete closed-loop systems spanning products, users, data and model iteration. In the coming years, more large model developers will pivot resources toward vertical application scenarios, seeking to establish their own Co-design workflows. The competition between Grok and Hunyuan ecosystems will continue to unfold, and the quality of their real-world iteration mechanisms will determine long-term market positioning.
Learn more:https://4sapi.com




