From 1G to 5G, wireless communication systems evolved from connecting people to connecting everything. To connect and serve ubiquitous intelligent agents, 6G and future wireless communication systems will shift from ubiquitous connectivity toward pervasive intelligent connectivity. They will evolve into integrated infrastructure that goes beyond connectivity by deeply integrating AI and sensing, while exhibiting strong heterogeneity across multiple dimensions. This heterogeneity challenges both conventional model-driven physical-layer algorithms and today's task-specific artificial intelligence (AI). A single base station can require hundreds of separately engineered models, and many must be retrained with new labeled data whenever the channel environment changes.
To address this fragmentation, a research team led by Professor Xiang Cheng at Peking University, working with collaborators at the Hong Kong University of Science and Technology (Guangzhou), developed WiFo-2, a generalist wireless foundation model for unified communications and sensing system design. Instead of building a new network for every setting, WiFo-2 learns reusable representations of channel state information (CSI) and transfers them across scenarios, system configurations and tasks. The study, "WiFo-2: a generalist foundation model unifies heterogeneous wireless system design," was published in National Science Review. Peking University doctoral student Boxun Liu is the first author, and Xiang Cheng is the corresponding author.
The foundation of the model is LH-CSI, a heterogeneous three-dimensional space-time-frequency dataset assembled for pretraining and evaluation. It contains 11.6 billion CSI points from eight data sources and 78 coarse- and fine-grained subsets. The data cover three acquisition modalities - statistical channel modeling, ray tracing and real-world measurement - as well as frequencies from sub-6 GHz to terahertz bands, varied antenna configurations and a broad range of mobility conditions. This diversity is designed to prevent the model from relying on shortcuts tied to one dataset or system.
WiFo-2 combines a masked denoising autoencoder with a CSI sparse mixture-of-experts mechanism. Its task-aware routing activates only selected experts for each input, maintaining model capacity while reducing computation. A two-stage training strategy first exposes the model to mixed masking and denoising tasks and then adds confidence-enhanced pretraining. The model consequently learns both how to reconstruct missing or corrupted CSI and how to estimate the reliability of its own reconstruction.
The researchers evaluated zero-shot reconstruction on three tasks: time-domain channel prediction, frequency-domain channel prediction and channel estimation. On unseen scenarios and system configurations, WiFo-2 required no task-specific fine-tuning yet outperformed fully supervised models trained with the complete task datasets. It also estimated reconstruction confidence with a mean absolute error of about 2.3 dB. Accurate confidence estimates can guide AI-model updates in practical systems and help maintain reliable communication links. Further tests showed scaling-law behavior shaped jointly by data volume, data heterogeneity and model size, providing practical guidance for future wireless foundation models.
For downstream adaptation, WiFo-2 was tested on nine communications and sensing tasks: scenario classification, angle-of-arrival estimation, Doppler estimation, CSI feedback, wireless localization, cross-band channel prediction, sub-6 GHz-to-millimeter-wave beam prediction, signal detection and vision-aided channel prediction. Using only 1% of the labeled training samples used for the full-shot task-specific baselines, the model achieved state-of-the-art performance across these tasks. The result reduces the data collection and training burden of adapting wireless AI to new uses.
The team also developed a wireless foundation model-powered hardware prototype for over-the-air testing. The transmitter and receiver use Ettus USRP X410 software-defined radios, while an NVIDIA Jetson AGX Orin at the receiver runs WiFo-2. The system follows the 5G New Radio frame structure and places demodulation reference signal pilots on selected time-frequency resources. The overall campaign covered diverse real-world environments, including indoor and outdoor scenarios, and multiple frequency bands spanning sub-6 GHz and millimeter-wave. Within this broad coverage, the team evaluated three representative tasks - frequency-domain channel prediction, channel estimation and scenario classification. Across the tested settings, WiFo-2 showed consistent performance advantages over task-specific and parametric baselines, demonstrating strong generalization across scenarios and frequency bands.
The prototype remains an early-stage demonstrator rather than a commercial-ready next-generation wireless system, and its latency and power consumption still require further optimization for practical mobile deployment. Even with these limitations, the simulations and over-the-air experiments show how one pretrained model can replace fragmented wireless AI pipelines and extend the foundation-model paradigm to wireless channels as a physical modality. The project has been open-sourced on GitHub. Repository: https://github.com/PKU-PCNI/WiFo-2 .
National Science Review
Computational simulation/modeling