Add support for ACE-Step 1.5 and ACE-Step 1.5 XL. Also added dataset captioning through the UI. (#785)

* Base ace step 1.5 xl added. Generating, still wip on training and ui

* Base training code done

* Fix some issues with caching text embeddings. Update sample cards to show audio

* Fix issue with quantizing ace step

* Add album artwork to samples with waveform.

* Cleanup logs

* Add album art endpoint to speed up album art loading

* Made an make video with artwork script

* Make ui handle basic audio models. Make multi line adjustments to the editor and better syntax hilighting.

* Add prompt tagging system for special tagged models.

* prompt tagging processing for ui working.

* Moved default samples to a special file so we can add more when needed and they can be adjusted for a specific model

* Add a captioner job with music captioner that is prepped for use with the ui

* Add basit ui setup for captioning modal and handeling captioning jobs

* Starting captioning job from ui working. Still better management for it.

* Better filtering of job options in the job view for captioning jobs

* Added qwen3 vl as a captioner for images

* Have an indicator when a dataset is being captioned.

* Adjust the way caption jobs look in the queue

* Fix a few issues. Adjust defaults.

* Version bump

* Added ace step to the readme.
This commit is contained in:
Jaret Burkett
2026-04-09 15:02:03 -06:00
committed by GitHub
parent 9ca58e9aa2
commit 78cf049c29
54 changed files with 5589 additions and 536 deletions

View File

@@ -6,6 +6,7 @@ import random
import torch
import torchaudio
from toolkit.audio.album_artwork import add_album_artwork
from toolkit.prompt_utils import PromptEmbeds
from torchao.quantization.quant_primitives import _DTYPE_TO_BIT_WIDTH
@@ -1201,13 +1202,16 @@ class GenerateImageConfig:
raise ValueError(f"Unsupported video format {self.output_ext}")
elif self.output_ext in ['wav', 'mp3']:
# save audio file
audio_path = self.get_image_path(count, max_count)
torchaudio.save(
self.get_image_path(count, max_count),
audio_path,
image[0].to('cpu'),
sample_rate=48000,
format=None,
backend=None
)
if self.output_ext == 'mp3':
add_album_artwork(audio_path)
else:
# TODO save image gen header info for A1111 and us, our seeds probably wont match
image.save(self.get_image_path(count, max_count))