INSIGHTS / Foundation Models & Learning
What Do We Do in 2026 Before AGI Arrives?
A view of model progress, agents and the practical work ahead while the path to AGI remains open.
The 2026 new year opened with a surge — an AI IPO wave swept through the capital markets as Birentech, Zhipu AI, and MiniMax all accelerated their listing processes, leading media to dub 2026 an 'IPO mega-year.' This also confirms that China's domestic LLMs have officially transitioned from a 'technology exploration phase' to an 'industry deep-cultivation phase'.
Yet faster industrialization does not mean technical paths have converged. With no clear AGI inflection point in sight, what is the main line of technological evolution? How should entrepreneurs seize interim opportunities? This is the core question shared by AI practitioners and investors alike.
At the current stage where technical capabilities are still in linear growth, on one hand, continuously exploring the upper bound of model intelligence remains the most important task; on the other hand, how to effectively activate and release the capabilities models already possess in real, complex scenarios has become an unavoidable practical challenge. These two points constitute the core of current industry focus.
In this context, understanding 'what to do when the inflection point hasn't arrived' becomes the key prerequisite for judging the next phase's technology and industry trends.
01 Before the Inflection Point: Exploring the Intelligence Ceiling Remains the Priority — Scaling Remains the Most Efficient Method
Pre-training has given large models world common sense knowledge and basic reasoning capabilities. More data, larger parameters, and more saturated compute remain the most efficient approach to scaling foundation models. Over the past two years, the industry has made extensive attempts in multi-modality and end-to-end voice, but practice has proven that what truly determines a model's intelligence ceiling is still text. Current models' knowledge and basic reasoning abilities both depend on 'reading extensively' — that is, the text data that carries human common sense knowledge. From a first-principles perspective, the first core element for exploring the intelligence ceiling is larger-scale data, especially common sense knowledge data (real data, synthetic data, and distilled data). Meanwhile, although more efficient methods are being explored at the model layer and infra layer, within the foreseeable future, scale remains the most effective path to improving model capabilities — beyond larger-scale data, larger model parameter sizes and more compute investment are still necessary.
02 Key to Unlocking Existing Model Capabilities: Activation Alignment and Enhanced Reasoning
Activation alignment and enhanced reasoning capabilities — especially activating more comprehensive long-tail capabilities — are another key to ensuring model effectiveness. Even if a model possesses common sense knowledge, if it cannot be effectively activated, its capabilities are difficult to release. In the past, alignment mainly relied on SFT, CoT, and similar methods, but currently reinforcement learning is gradually becoming the mainstream path: by constructing scenarios and data, models self-optimize in reinforcement learning to activate their intelligence ceiling. But real-world challenges exist: the emergence of general benchmarks can on one hand evaluate models' general effectiveness, but may also lead many models to overfit. Meanwhile, the number of evaluation benchmarks is limited, while real-world application scenarios are virtually infinite. With the proliferation of Agents, alignment complexity has significantly increased — if past model alignment was more like taking exams on language and math knowledge points, Agents are closer to real work environments, with diverse job types, complex situations, and even requiring a certain level of 'emotional intelligence.' For example, booking flights and hotels essentially involves introducing Agent data to align with human behavior, thereby activating model capabilities. Therefore, the breakthrough for the next phase is making models align long-tail real-world scenarios faster and better. Mid-training and post-training will make rapid alignment and strong reasoning capabilities across more scenarios possible.
03 Agent Evolution: Cross-Environment Generalization and Migration Remain Difficult — Simplest Solution: Construct Environment Data + Reinforcement Learning
Agent is a milestone in model capability expansion and the key to AI models entering humans' real (virtual/physical) world. Without Agent capabilities, large models will remain at the (theoretical learning) stage, unable to convert knowledge accumulation into real productivity. Over the next year, more Agents will continue to emerge — this judgment is no longer fresh. But to drive Agent large-scale application and deployment, the core first lies in understanding that AI's essence is not creating entirely new capabilities, but systematically replacing part of humans' work or capabilities. This is also the key factor determining Agent's inevitable explosion. Second, although early Agents primarily implemented through model applications, current models can already integrate Agent data directly into the training process, enhancing model generality. However, the difficulty lies in generalization and migration between different Agent environments. Therefore, the simplest path at the application layer, as mentioned earlier, is to continuously introduce data from more different Agent environments and conduct reinforcement learning for different environments.
04 Focus for Model Capability Breakthroughs: Memory, Online Learning, and Self-Evaluation
At the application layer: build applications around different job types, and keep them sufficiently 'thin' to adapt to model capability evolution.
Regarding the application layer, there are two main judgments: First, as large models develop increasingly end-to-end, model R&D and model applications inevitably need to be combined. The application layer must be kept sufficiently 'thin,' able to quickly adapt to foundation model capability changes, and rapidly launch, iterate, and verify value. Second, building model applications should start from the perspective of helping people and creating new value. If you build an AI software nobody uses, it cannot create value, and thus is destined to lack vitality.
Specifically, what should AI software actually do? Returning to first principles, AI does not need to create new applications — its essence is to simulate people, replace people, or help people accomplish what humans must do (certain job types). From this perspective, future applications mainly divide into two categories: AI-enabling past software, changing parts that originally required human involvement to AI; creating AI software that aligns with some human population, replacing human work.
05 Embodied Intelligence: Unlocking Model Capabilities in More Real Scenarios — Data Scale and Hardware Technology Issues Need to Be Resolved
Beyond the virtual world, corresponding to the Agent explosion is the development of embodied intelligence. Robot brains now have fairly strong intelligence — the next key is to properly align and activate large model capabilities with robot behaviors. As hardware technology gradually matures, embodied intelligence is expected to release model capabilities in more real-world scenarios. But like the generalization difficulties Agents face, activating universal embodied capabilities through few samples is essentially impossible. This requires breaking through the problem of large-scale data acquisition.
06 Multi-Modality: A Long-Term Important Direction — Efficient Path: Parallel Exploration of Text, Multi-Modal Understanding, and Multi-Modal Generation
Although multi-modality hasn't yet provided substantial help in exploring the AGI intelligence ceiling, it is undoubtedly an important direction. The current effective path remains parallel development of text, multi-modal understanding, and multi-modal generation, with moderate exploration of their combination to discover new capability emergence. Of course, exploration of their combination requires not only courage in technical exploration but also long-term, sustained capital investment.
07 Domain Models: Rationally Viewing phased Opportunities
Under the AGI endgame, domain models will likely be covered by foundation models — domain data, processes, and Agent data will gradually enter the main model. But in the phase before AGI is achieved, with sustained commitment, and when possessing domain data and maintaining close coupling with foundation models, opportunities still exist to establish leading advantages in specific competitive windows.