The coexistence of multiple Radio Access Technologies (RATs) and wired access networks (e.g., Ethernet and fiberoptic broadband) provides flexible connectivity but also makes access network selection increasingly complex. Traditional optimization, game-theoretic, and heuristic methods require global information or incur high signaling overhead, limiting scalability in dynamic environments. While Reinforcement Learning (RL) offers adaptability, existing RL and Deep RL (DRL) approaches often rely on centralized coordination or network-assisted data, reducing their effectiveness in large-scale, mobile, and heterogeneous scenarios. We propose ON-policy multi-agent PPO for adaptive access NETwork selection (On-Net): a decentralized, user-centric DRL framework based on centralized training with decentralized execution (CTDE). In this design, users independently make network access selection decisions using only local observations, while benefiting from collaboration during training. We validate our approach in both simulation and a real-world testbed, demonstrating its effectiveness in selecting between various access technologies, whether wireless or wired. On-Net balances high throughput, fairness among users, and reduced handovers with controlled decision times; notably, real-world experiments demonstrated a 90% reduction in handovers.

On-Net: On-Policy Multi-Agent Reinforcement Learning for Adaptive Access Network Selection / Solana, Á., Monaco, D., Sacco, A., Marchetto, G.. - (2026), pp. 1-6. (IEEE/IFIP Network Operations and Management Symposium Roma (Italia) 18-22 Maggio 2026) [10.1109/NOMS69089.2026.11668165].

On-Net: On-Policy Multi-Agent Reinforcement Learning for Adaptive Access Network Selection

Doriana Monaco;Alessio Sacco;Guido Marchetto
2026

Abstract

The coexistence of multiple Radio Access Technologies (RATs) and wired access networks (e.g., Ethernet and fiberoptic broadband) provides flexible connectivity but also makes access network selection increasingly complex. Traditional optimization, game-theoretic, and heuristic methods require global information or incur high signaling overhead, limiting scalability in dynamic environments. While Reinforcement Learning (RL) offers adaptability, existing RL and Deep RL (DRL) approaches often rely on centralized coordination or network-assisted data, reducing their effectiveness in large-scale, mobile, and heterogeneous scenarios. We propose ON-policy multi-agent PPO for adaptive access NETwork selection (On-Net): a decentralized, user-centric DRL framework based on centralized training with decentralized execution (CTDE). In this design, users independently make network access selection decisions using only local observations, while benefiting from collaboration during training. We validate our approach in both simulation and a real-world testbed, demonstrating its effectiveness in selecting between various access technologies, whether wireless or wired. On-Net balances high throughput, fairness among users, and reduced handovers with controlled decision times; notably, real-world experiments demonstrated a 90% reduction in handovers.
2026
979-8-3315-9268-4
File in questo prodotto:
File Dimensione Formato  
On-Net_On-Policy_Multi-Agent_Reinforcement_Learning_for_Adaptive_Access_Network_Selection.pdf

accesso riservato

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Non Pubblico - Accesso privato/ristretto
Dimensione 366.52 kB
Formato Adobe PDF
366.52 kB Adobe PDF   Visualizza/Apri   Richiedi una copia
On_Net___IPSN_2026.pdf

accesso aperto

Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 296.94 kB
Formato Adobe PDF
296.94 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015431