Deep Reinforcement Learning Based Cross-Layer Design in Terahertz Mesh Backhaul Networks

📄 arXiv: 2310.05034v1 📥 PDF

作者: Zhifeng Hu, Chong Han, Xudong Wang

分类: cs.LG

发布日期: 2023-10-08


💡 一句话要点

提出基于深度强化学习的跨层设计以解决太赫兹网回程网络问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 深度强化学习 太赫兹网络 跨层设计 资源分配 动态流量管理 链路恢复 无线通信 网络优化

📋 核心要点

  1. 太赫兹网的动态流量需求和链路故障导致现有跨层路由和资源分配方法面临挑战,尤其是高指向性和非视距路径损失问题。
  2. 本文提出DEFLECT方法,通过深度强化学习实现动态流量和链路故障的高效管理,设计启发式路由指标和资源分配算法。
  3. 实验结果显示,DEFLECT在资源消耗上显著优于传统方法,且实现了无数据包丢失和快速链路恢复,提升了系统的整体性能。

📝 摘要(中文)

太赫兹(THz)网因其支持超高数据速率和灵活重构能力,成为下一代无线回程系统的理想选择。然而,动态流量需求和可能的链路故障使得高效的跨层路由和长期资源分配成为一个开放问题。本文提出了一种基于深度强化学习(DRL)的跨层设计方法DEFLECT,旨在应对动态流量和突发链路故障。该方法首先设计了一种启发式路由指标,以提高资源效率,并开发了DRL资源分配算法,实现长期资源效率最大化和快速链路恢复。仿真结果表明,DEFLECT在资源消耗上优于传统的最小跳数指标,且在无数据包丢失和毫秒级延迟的情况下,能够在1秒内恢复资源高效的回程链路。

🔬 方法详解

问题定义:本文旨在解决太赫兹网中动态流量和链路故障导致的跨层路由和资源分配效率低下的问题。现有方法在面对高指向性和非视距路径损失时,难以有效应对流量波动和链路中断。

核心思路:DEFLECT方法通过深度强化学习(DRL)实现动态流量管理和链路故障恢复,设计启发式路由指标以提高资源效率,并通过多任务结构优化功率和子阵列分配。

技术框架:DEFLECT的整体架构包括启发式路由模块和DRL资源分配模块。启发式路由模块用于评估资源效率,DRL模块则负责长期资源分配和链路恢复。

关键创新:DEFLECT的主要创新在于结合了启发式路由和深度强化学习,能够在无数据包丢失的情况下实现长期资源效率最大化,并快速恢复链路。与传统DRL方法相比,DEFLECT在链路故障恢复速度和资源消耗上表现更优。

关键设计:在设计中,采用了多任务学习结构以促进功率和子阵列的联合分配,同时引入了分层架构以实现针对每个基站的定制资源分配。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,DEFLECT在资源消耗上显著低于传统的最小跳数路由方法,且在无数据包丢失的情况下实现了毫秒级的延迟。链路故障恢复时间缩短至1秒,极大提升了系统的可靠性和效率。

🎯 应用场景

该研究具有广泛的应用潜力,尤其是在下一代无线通信系统中,如5G及未来的6G网络。通过提高太赫兹网的资源管理效率,DEFLECT能够支持更高的数据传输速率和更灵活的网络配置,推动智能城市、物联网等领域的发展。

📄 摘要(原文)

Supporting ultra-high data rates and flexible reconfigurability, Terahertz (THz) mesh networks are attractive for next-generation wireless backhaul systems that empower the integrated access and backhaul (IAB). In THz mesh backhaul networks, the efficient cross-layer routing and long-term resource allocation is yet an open problem due to dynamic traffic demands as well as possible link failures caused by the high directivity and high non-line-of-sight (NLoS) path loss of THz spectrum. In addition, unpredictable data traffic and the mixed integer programming property with the NP-hard nature further challenge the effective routing and long-term resource allocation design. In this paper, a deep reinforcement learning (DRL) based cross-layer design in THz mesh backhaul networks (DEFLECT) is proposed, by considering dynamic traffic demands and possible sudden link failures. In DEFLECT, a heuristic routing metric is first devised to facilitate resource efficiency (RE) enhancement regarding energy and sub-array usages. Furthermore, a DRL based resource allocation algorithm is developed to realize long-term RE maximization and fast recovery from broken links. Specifically in the DRL method, the exploited multi-task structure cooperatively benefits joint power and sub-array allocation. Additionally, the leveraged hierarchical architecture realizes tailored resource allocation for each base station and learned knowledge transfer for fast recovery. Simulation results show that DEFLECT routing consumes less resource, compared to the minimal hop-count metric. Moreover, unlike conventional DRL methods causing packet loss and second-level latency, DEFLECT DRL realizes the long-term RE maximization with no packet loss and millisecond-level latency, and recovers resource-efficient backhaul from broken links within 1s.