Music generation via multimodal machine learning generative models has recently gained significant traction in both the research community and the general public. Among such models, an important category is represented by Text-To-Music models, which take only a textual description or caption as input to specify the type of composition the user wishes to generate. While generative artificial intelligence models achieve remarkable performance, they also require substantial computational power during both training and inference. The latter, in particular, becomes increasingly critical as these models are deployed more widely in real-world music-making applications. In this paper, we address this issue by analyzing the energy consumption of seven diffusion-based and two autoregressive Text-To-Music models. Through a series of experiments, we evaluate how model-dependent generation parameters influence energy usage during inference. Since audio generation quality, prompt adherence, and energy efficiency are all important, we employ Pareto-optimal analysis to explore the trade-offs among these competing objectives. Our results highlight the balance between model performance and energy usage, offering guidance for designing more environmentally conscious generative audio systems.
The following figures are interactive visualizations created with Plotly.
Hover over markers to inspect the exact values, use the mouse wheel to zoom, and click and drag with the left mouse button to move across the plot.
Double-click anywhere inside a figure to reset the view.
Click legend entries to hide or show individual models for easier comparisons.
If some markers are covered by the legend, simply drag the plot while holding the left mouse button.
This experiment was performed on randomly selected audio and caption pairs from the MusicCaps and SongDescriber datasets. Pareto-efficient configurations are highlighted in red.
CLAP Score results have been obtained with the stable-audio-metrics toolkit (higher is better).
The following metrics were computed with the audiobox-aesthetics library (higher is better):
git clone https://github.com/rickgiantsteps/musicdiff-consumption
cd musicdiff-consumption
conda env create -f requirements.yml
@ARTICLE{ronchini2026,
author={Ronchini, Francesca and Passoni, Riccardo and Comanducci, Luca and Serizel, Romain and Antonacci, Fabio},
journal={},
title={Analyzing the Energy Consumption of Generative Text-to-Music Models},
year={2026},
volume={},
number={},
pages={},
keywords={},
doi={}
}