Analyzing the Energy Consumption of Generative Text-to-Music Models

1Dipartimento di Elettronica, Informazione e Bioingegneria - Politecnico di Milano, Milan, Italy 2Institute of Sound Recording (IoSR), University of Surrey, Guildford, UK 3Université de Lorraine, CNRS, Inria, Loria, Nancy, France
*These authors contributed equally

Abstract

Music generation via multimodal machine learning generative models has recently gained significant traction in both the research community and the general public. Among such models, an important category is represented by Text-To-Music models, which take only a textual description or caption as input to specify the type of composition the user wishes to generate. While generative artificial intelligence models achieve remarkable performance, they also require substantial computational power during both training and inference. The latter, in particular, becomes increasingly critical as these models are deployed more widely in real-world music-making applications. In this paper, we address this issue by analyzing the energy consumption of seven diffusion-based and two autoregressive Text-To-Music models. Through a series of experiments, we evaluate how model-dependent generation parameters influence energy usage during inference. Since audio generation quality, prompt adherence, and energy efficiency are all important, we employ Pareto-optimal analysis to explore the trade-offs among these competing objectives. Our results highlight the balance between model performance and energy usage, offering guidance for designing more environmentally conscious generative audio systems.

Experiments

The following figures are interactive visualizations created with Plotly. Hover over markers to inspect the exact values, use the mouse wheel to zoom, and click and drag with the left mouse button to move across the plot. Double-click anywhere inside a figure to reset the view. Click legend entries to hide or show individual models for easier comparisons.

If some markers are covered by the legend, simply drag the plot while holding the left mouse button.

Varying Inference Step Number


GPU energy consumption during audio generation
Efficiency comparison between AudioLDM and AudioLDM2
Equivalent CO2 emissions generated by the inference setups
CO2 equivalents emitted per second by the models

Varying Inference Batch Size


GPU energy consumption for different batch sizes
Emissions released for different batch sizes
GPU energy consumption (lighter models)
GPU energy consumption rate for different batch sizes

Quality Metrics Analysis



This experiment was performed on randomly selected audio and caption pairs from the MusicCaps and SongDescriber datasets. Pareto-efficient configurations are highlighted in red.

Fréchet Audio Distance (FAD)

The FAD has been computed with the fadtk library, using the LAION and EnCodec encoders (lower is better).

MusicCaps Dataset

SongDescriber Dataset

Contrastive Language-Audio Pretraining (CLAP) Score

CLAP Score results have been obtained with the stable-audio-metrics toolkit (higher is better).

Audiobox Metrics

The following metrics were computed with the audiobox-aesthetics library (higher is better):

CE — Content Enjoyment (subjective quality of audio)
CU — Content Usefulness (suitability for content production)
PC — Production Complexity (number of audio elements)
PQ — Production Quality (evaluates technical properties)

MusicCaps Dataset

SongDescriber Dataset

Try it for yourself


git clone https://github.com/rickgiantsteps/musicdiff-consumption
cd musicdiff-consumption




You can create the virtual environment and install the needed packages using conda with the following command

conda env create -f requirements.yml

BibTeX

@ARTICLE{ronchini2026,
  author={Ronchini, Francesca and Passoni, Riccardo and Comanducci, Luca and Serizel, Romain and Antonacci, Fabio},
  journal={},
  title={Analyzing the Energy Consumption of Generative Text-to-Music Models},
  year={2026},
  volume={},
  number={},
  pages={},
  keywords={},
  doi={}
}