The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and deployment increasingly costly. These models are often used as computational monoliths with fixed cost, hindering adaptive deployment across different cost budgets. We argue that nested components, ordered by importance, can be extracted from pretrained models and selectively activated within the available computational budget. To this end, our proposed FlexRank method leverages low-rank weight decomposition with nested, importance-based consolidation to extract submodels of increasing capabilities. Our approach enables a ``train-once, deploy-everywhere'' paradigm offering a graceful trade-off between cost and performance without training from scratch for each budget - advancing practical deployment of large models.

FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment / Zaccone, R., Laskaridis, S., Ciccone, M., Horváth, S.. - ELETTRONICO. - (In corso di stampa). (43rd International Conference on Machine Learning (ICML) Seoul, South Korea July 6 - July 11, 2026).

FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment

Riccardo Zaccone;Marco Ciccone;
In corso di stampa

Abstract

The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and deployment increasingly costly. These models are often used as computational monoliths with fixed cost, hindering adaptive deployment across different cost budgets. We argue that nested components, ordered by importance, can be extracted from pretrained models and selectively activated within the available computational budget. To this end, our proposed FlexRank method leverages low-rank weight decomposition with nested, importance-based consolidation to extract submodels of increasing capabilities. Our approach enables a ``train-once, deploy-everywhere'' paradigm offering a graceful trade-off between cost and performance without training from scratch for each budget - advancing practical deployment of large models.
In corso di stampa
File in questo prodotto:
File Dimensione Formato  
4236_FlexRank_Nested_Low_Rank_.pdf

accesso aperto

Descrizione: Camera ready manuscript
Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 1.23 MB
Formato Adobe PDF
1.23 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013135