A new artificial intelligence model that estimates tree heights from standard RGB satellite imagery with near-lidar accuracy could transform forest monitoring and carbon accounting. Developed by a joint research team from Beijing Forestry University, Manchester Metropolitan University, and Tsinghua University, the framework integrates large vision foundation models (LVFMs) with self-supervised enhancement to produce high-resolution canopy height maps at sub-meter precision. The study, published in the Journal of Remote Sensing on October 20, 2025 (DOI: 10.34133/remotesensing.0880), addresses the longstanding challenge of balancing cost, accuracy, and scalability in forest monitoring.
Traditional lidar systems provide accurate height data but are expensive and technically complex, limiting their use for large-scale or frequent monitoring. Optical remote sensing methods often lack the structural precision needed for small-scale plantations, and deep learning approaches typically require massive labeled datasets and lose fine spatial details. The new model overcomes these limitations by combining the DINOv2 LVFM as a feature extractor with a self-supervised feature enhancement unit that retains fine spatial details, followed by a lightweight convolutional height estimator. When tested against airborne lidar measurements in Beijing's Fangshan District—a region of fragmented plantations dominated by Populus tomentosa, Pinus tabulaeformis, and Ginkgo biloba—the model achieved a mean absolute error of 0.09 meters and an R² of 0.78, outperforming traditional CNN and transformer-based methods.
Beyond height estimation, the model enabled single-tree detection with over 90% accuracy and produced strong correlations with measured above-ground biomass. It also demonstrated robust cross-regional adaptability when applied to a geographically distinct forest in Saihanba, maintaining accuracy and confirming its potential for national-scale carbon accounting. The ability to reconstruct annual growth trends from archived satellite imagery provides a scalable solution for long-term carbon sink monitoring under initiatives such as China's Certified Emission Reduction program.
"Our model demonstrates that large vision foundation models (LVFMs) can fundamentally transform forestry monitoring," said Dr. Xin Zhang, corresponding author at Manchester Metropolitan University. "By combining global image pretraining with local self-supervised enhancement, we achieved lidar-level precision using ordinary RGB imagery. This approach drastically reduces costs and expands access to accurate forest data for carbon accounting and environmental management."
The framework was implemented using PyTorch and the fastai library on an NVIDIA RTX A6000 GPU, with one-meter-resolution Google Earth imagery from 2013 to 2020 as input and UAV-based lidar data for validation. Comparative experiments with conventional networks like U-Net and DPT confirmed superior accuracy and efficiency, validating the model's potential for scalable canopy height mapping and biomass estimation. The research was supported by the National Science Foundation of China, the Natural Science Foundation of Beijing, BBSRC, EPSRC, and the Key Research and Development Program of Shaanxi Province.
This innovation bridges the gap between expensive lidar surveys and low-resolution optical methods, enabling detailed forest assessment with minimal data requirements. Future research will extend the method to natural and mixed forests, integrate automated species classification, and support real-time carbon monitoring platforms. As global efforts toward net-zero goals intensify, such intelligent mapping tools could play a central role in sustainable forestry and climate-change mitigation.


