The coexistence of multiple Radio Access Technologies (RATs) and wired access networks (e.g., Ethernet and fiberoptic broadband) provides flexible connectivity but also makes access network selection increasingly complex. Traditional optimization, game-theoretic, and heuristic methods require global information or incur high signaling overhead, limiting scalability in dynamic environments. While Reinforcement Learning (RL) offers adaptability, existing RL and Deep RL (DRL) approaches often rely on centralized coordination or network-assisted data, reducing their effectiveness in large-scale, mobile, and heterogeneous scenarios. We propose ON-policy multi-agent PPO for adaptive access NETwork selection (On-Net): a decentralized, user-centric DRL framework based on centralized training with decentralized execution (CTDE). In this design, users independently make network access selection decisions using only local observations, while benefiting from collaboration during training. We validate our approach in both simulation and a real-world testbed, demonstrating its effectiveness in selecting between various access technologies, whether wireless or wired. On-Net balances high throughput, fairness among users, and reduced handovers with controlled decision times; notably, real-world experiments demonstrated a 90% reduction in handovers.
On-Net: On-Policy Multi-Agent Reinforcement Learning for Adaptive Access Network Selection / Solana, Á., Monaco, D., Sacco, A., Marchetto, G.. - (2026), pp. 1-6. (IEEE/IFIP Network Operations and Management Symposium Roma (Italia) 18-22 Maggio 2026) [10.1109/NOMS69089.2026.11668165].
On-Net: On-Policy Multi-Agent Reinforcement Learning for Adaptive Access Network Selection
Doriana Monaco;Alessio Sacco;Guido Marchetto
2026
Abstract
The coexistence of multiple Radio Access Technologies (RATs) and wired access networks (e.g., Ethernet and fiberoptic broadband) provides flexible connectivity but also makes access network selection increasingly complex. Traditional optimization, game-theoretic, and heuristic methods require global information or incur high signaling overhead, limiting scalability in dynamic environments. While Reinforcement Learning (RL) offers adaptability, existing RL and Deep RL (DRL) approaches often rely on centralized coordination or network-assisted data, reducing their effectiveness in large-scale, mobile, and heterogeneous scenarios. We propose ON-policy multi-agent PPO for adaptive access NETwork selection (On-Net): a decentralized, user-centric DRL framework based on centralized training with decentralized execution (CTDE). In this design, users independently make network access selection decisions using only local observations, while benefiting from collaboration during training. We validate our approach in both simulation and a real-world testbed, demonstrating its effectiveness in selecting between various access technologies, whether wireless or wired. On-Net balances high throughput, fairness among users, and reduced handovers with controlled decision times; notably, real-world experiments demonstrated a 90% reduction in handovers.| File | Dimensione | Formato | |
|---|---|---|---|
|
On-Net_On-Policy_Multi-Agent_Reinforcement_Learning_for_Adaptive_Access_Network_Selection.pdf
accesso riservato
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
366.52 kB
Formato
Adobe PDF
|
366.52 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
|
On_Net___IPSN_2026.pdf
accesso aperto
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Pubblico - Tutti i diritti riservati
Dimensione
296.94 kB
Formato
Adobe PDF
|
296.94 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015431
