Opponent Modeling-Based Dynamic Resource Trading for Multi-UAV Assisted Edge Computing
Meng-Ze Liu, Zhe Wang, Jin-Xiang Bai, Long Shi, Jun Li, Kang Wei, Heng-Tao He, Jie Zhang, Shi Jin
Abstract
In uncrewed aerial vehicle (UAV)-assisted mobile edge computing (MEC) networks, effectively incentivizing the self-interested UAV servers to participate in cost-efficient edge computing tasks through the dynamic pricing mechanism remains a critical challenge. This article proposes a decentralized resource trading scheme where the profit-driven UAV servers strategically sell the computation offloading services to the mobile users (MUs) with time-varying demands. We formulate the sequential interactions between the UAVs and the MUs as a partially observable stochastic multileader multifollower (POS-MLMF) Stackelberg game, where each UAV acts as a leader to maximize its long-term profits by optimizing the flight trajectories and service prices, and each MU serves as a follower to minimize the weighted sum of its cumulative delay and payment by optimizing its offloading strategy. Due to the coevolution of the multiagent policies, the environment faced by each agent is nonstationary. To optimize the best response policy for each trading agent in the nonstationary environment, we propose two opponent modeling-based reinforcement learning algorithms, namely deep reinforcement opponent network with a dueling double deep Q-network (DRON-D3QN) and the neural fictitious self-play with a dueling double deep Q-network (NFSP-D3QN). These algorithms enable each agent to either explicitly or implicitly model the opponents’ trading strategies, thereby facilitating effective policy optimization in a decentralized learning paradigm. Simulation results demonstrate that our proposed opponent modeling-based algorithms outperform the baselines by simultaneously improving the UAVs’ profits and reducing the MUs’ costs, achieving a win–win outcome.
Source: semanticscholar
Read the AI summary and key takeaways for traders on WOBR Quant Research.