Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings

📄 arXiv: 2609.03480v2 📥 PDF

作者: Alkiviadis Koukos, Spyros Kondylatos, Thomas Nord-Larsen, Lotte Nyborg, Christian Tøttrup, Kenneth Grogan

分类: cs.CV, cs.AI, cs.LG

发布日期: 2026-09-03 (更新: 2026-09-04)

备注: This preprint presents a national-scale tree species mapping framework for Denmark using Sentinel-1/2 time series, National Forest Inventory data, and EO foundation model embeddings. The resulted national map can be found here: https://zenodo.org/uploads/22108850


💡 一句话要点

通过比较光谱时序特征与基础模型嵌入实现丹麦树种映射

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 树种映射 光谱时序特征 基础模型 遥感数据 森林监测 生态研究 机器学习

📋 核心要点

  1. 现有方法在树种分类中面临数据稀缺和特征提取不足的挑战,影响了森林资源的有效管理。
  2. 论文提出通过比较手工设计的光谱时序特征与基础模型嵌入,优化树种分类的输入表示。
  3. 实验结果表明,基于STF的MLP模型在纯林分类中表现最佳,而TESSERA嵌入在数据稀缺情况下具有显著优势。

📝 摘要(中文)

本研究利用国家森林清查样地和遥感数据对丹麦的树种进行映射,同时评估基础模型在大规模森林特征化中的潜力。我们比较了两种树种分类输入表示:手工设计的光谱时序特征(STF)和由EO基础模型TESSERA及AlphaEarth生成的嵌入。随机森林、XGBoost和多层感知器(MLP)分类器被用于评估所有输入表示。结果显示,基于STF的MLP在纯林和混交林中分别取得了0.843和0.653的宏F1分数,表现最佳。最终生成的10米分辨率树种地图为丹麦首个高分辨率国家树种地图,具有重要的生态监测和土地管理价值。

🔬 方法详解

问题定义:本研究旨在解决丹麦树种映射中的分类精度问题,尤其是在训练数据稀缺的情况下,现有方法往往依赖于手工特征提取,导致效果不佳。

核心思路:通过比较手工设计的光谱时序特征与基础模型生成的嵌入,探索不同输入表示对树种分类性能的影响,旨在提高分类精度和适应性。

技术框架:研究采用随机森林、XGBoost和多层感知器(MLP)作为分类器,结合光谱时序特征和基础模型嵌入,评估不同输入表示的分类效果,并在全国范围内生成树种地图。

关键创新:本研究的创新在于引入基础模型嵌入(如TESSERA和AlphaEarth)作为树种分类的输入,尤其在训练样本不足时表现出明显优势,超越了传统手工特征提取方法。

关键设计:在模型训练中,采用宏F1分数作为性能评估指标,针对纯林和混交林分别进行评估,设计了多层感知器的网络结构,并结合了树冠高度信息以提升分类准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,基于STF的MLP在纯林和混交林中分别取得了0.843和0.653的宏F1分数,表现最佳。而TESSERA嵌入在训练样本不足时的表现优于STF模型,显示出在数据稀缺情况下的显著优势。最终生成的树种地图整体准确率达到79.9%。

🎯 应用场景

该研究的成果可广泛应用于森林监测、生态研究和土地管理等领域。高分辨率的树种地图为政策制定者和研究人员提供了重要的基础数据,促进了对森林资源的可持续管理和保护。同时,研究方法的创新也为其他地区的树种映射提供了借鉴。

📄 摘要(原文)

We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.