Autonomous vehicles and robots need to understand a world that is constantly moving. Vehicles change lanes, pedestrians cross streets, and the roadscape can shift from one moment to the next.
To train and test perception and motion-planning systems, researchers need large amounts of high-quality 4D LiDAR data. However, collecting real LiDAR sequences is expensive, time-consuming and labour-intensive. Existing methods for generating 4D LiDAR data can also struggle to represent large outdoor environments where both scene structure and object movement remain realistic over time.
Researchers from Singapore University of Technology and Design (SUTD) have developed HieraScaffold , an AI framework that generates large-scale 4D LiDAR scenes more efficiently and coherently. LiDAR, short for Light Detection and Ranging, uses laser pulses to measure distances and build a 3D picture of the surrounding world. In 4D LiDAR, time is added to this 3D information, allowing systems to model how a scene changes across a sequence.
Led by SUTD Assistant Professor Zhao Na, the research paper, “ HieraScaffold: Learning Compact Hierarchical Representations for Scalable 4D LiDAR Generation ,” presents a new way to generate large outdoor 4D LiDAR environments that capture both static structures, such as roads and buildings, and moving objects, such as vehicles and pedestrians.
HieraScaffold builds a compact “scaffold” around surface-adjacent regions sampled by LiDAR, rather than representing the entire 3D/4D space densely. It represents the static background and dynamic foreground in separate scaffold components, compresses them into compact directional representations, and then generates the static scaffold first before generating the dynamic scaffold conditioned on it.
This matters because autonomous systems need to understand not only where objects are, but how they move from one moment to the next.
“The real world is dynamic, not a collection of disconnected snapshots,” said Assistant Prof Zhao. “Vehicles and pedestrians move, so a system must understand not only where objects are, but also how they change over time. By adding time to 3D space, 4D LiDAR provides the smooth and consistent sequences needed for reliable perception and planning.”
HieraScaffold learns compact hierarchical representations from raw LiDAR sequences. It reconstructs continuous scene geometry from sparse scans using unsigned distance fields and spatial gradients, then compresses the resulting scaffolds with a neural contourlet representation that preserves important directional geometric and motion cues. Building on this compact representation, HieraScaffold progressively recouples the decomposed components to generate coherent 4D LiDAR scenes.
In practical terms, HieraScaffold generates large scenes as connected blocks instead of trying to produce an entire outdoor environment as one dense volume. This block-based approach makes the process more scalable while helping neighbouring parts of the scene remain consistent.
In tests using the KITTI-360 and Waymo Open autonomous-driving datasets, HieraScaffold generated 4D LiDAR scenes that more closely matched real-world data and remained more consistent over time than several existing methods including LidarDM. The generated scenes showed better overall realism, more accurate spatial structure, and a closer match to the distribution of real LiDAR data. It also produced smoother transitions between generated LiDAR frames, an important feature for modelling moving traffic scenes.
For autonomous-driving and robotics teams, the potential value lies not only in generating more data, but in generating more useful variation. In downstream experiments, the researchers found that 10,000 synthetic LiDAR samples combined with only 1,000 real samples outperformed training with 10,000 real samples alone for both segmentation and vehicle detection.
This suggests that high-quality synthetic data could become a powerful complement to real-world data. It may help perception models generalise better by exposing them to more varied scenes and movement patterns. However, the researchers caution that synthetic data should not yet be seen as a complete replacement for real data.
“High-quality synthetic data can do more than simply increase the size of a dataset,” said Assistant Prof Zhao. “It can introduce useful diversity and help models generalise better. The goal is not to replace real data entirely, but to make training and testing more scalable, flexible and efficient.”
The technology could eventually reduce the cost of collecting and labelling large LiDAR datasets. It could also make it easier to build dynamic simulation environments, allowing more perception and planning tests to be carried out virtually before systems are tested on public roads. Beyond autonomous driving, the approach may support robotics, virtual reality and augmented reality applications that require large amounts of dynamic 3D data.
The team also highlights current limitations. Existing LiDAR datasets still have limited scene diversity and may contain areas that were not fully scanned. As a result, models trained on these datasets may reproduce missing areas as holes or abrupt boundaries. More complete and varied dynamic datasets will be needed to improve robustness.
Next, the researchers aim to train HieraScaffold on richer 4D LiDAR datasets, improve its ability to recover missing scene structures, and evaluate its robustness more broadly before any real-world deployment.
HieraScaffold: Learning Compact Hierarchical Representations for Scalable 4D LiDAR Generation