Jaret Burkett
764b5064fb
Migrate to a new DTO for latents to carry more information that a normal tensor such as audio.
2026-08-30 10:30:21 -06:00
Jaret Burkett
da79ebce99
Add D-OPSD as a distillation handeling option for MiniMax H3 ref2va
2026-08-26 11:06:33 -06:00
Jaret Burkett
afd1d92722
Fix issue with double sampling qwen vl video frames for hidream h3 for reference videos
2026-08-17 15:39:07 -06:00
Jaret Burkett
127d6f626d
Rework img/video reference in Minimax H3 to more closely match the comfy ui implementation.
2026-08-15 07:13:54 -06:00
Jaret Burkett
97bf49edad
Add support for video references in MiniMax H3 ref2va
2026-08-15 06:18:09 -06:00
Jaret Burkett
21dc65972d
Only add a blank control if the model has to have it, for unconditionals. Previously it always encoded a blank control on unconditional if the model could take one. Affects DOP and blank prompt preservations.
2026-08-11 08:25:23 -06:00
Jaret Burkett
f421542df4
When dropping out, caching, and doing DOP, make sure we select a cahced trigger word when dropped out so DOP matches the drop out embeddings.
2026-08-11 06:59:20 -06:00
Jaret Burkett
ab5fef8970
Fix caption dropout so it now works identically when caching text embeddings.
2026-08-10 18:53:33 -06:00
Jaret Burkett
257da9b586
Rework DOP so it works with caching text embeddings
2026-08-09 22:13:49 -06:00
Jaret Burkett
72623ed3d6
When doing auto frame count. Ensure the time is not squeezed or expanded to fit tempooral spacing. trime the few extra frames. Also fixed frame counts of buckets.
2026-08-09 12:31:02 -06:00
Jaret Burkett
8c1a4082fd
Allow images to work with auto frame count, and include images in video datasets if they exist.
2026-08-08 20:26:28 -06:00
Jaret Burkett
00a93e3830
Add dataset flag to cache the raw tensors
2026-08-04 15:32:48 -06:00
Jaret Burkett
a9a04547e9
Dont move encoders on and off device when caching until the first instance of needing to process a cache item.
2026-08-03 15:54:30 -06:00
Jaret Burkett
41676bb258
Queue up videos with multiple threads when caching latents so the VAE is not waiting on videos to process
2026-08-03 15:41:36 -06:00
Jaret Burkett
8502a845b1
Add support for MiniMax H3 T2V and I2V training
2026-08-03 10:17:39 -06:00
Jaret Burkett
7b2386c096
Handle video codecs that fail in opencv
2026-07-29 21:01:57 -06:00
Jaret Burkett
3f8afcac7e
Allow dataloader to encode first frame with the text embeddings if the model needs it.
2026-07-29 17:50:45 -06:00
Jaret Burkett
8f2d001eae
Improvements to video frame loading. Added ability to cache as uint8 pixelspace for video
2026-07-24 12:36:47 -06:00
Jaret Burkett
dd08579eda
Add ability to pull control images from same folder group
2026-07-08 05:21:31 -06:00
Jaret Burkett
b1e1a834d4
Added support to train Krea2 as an edit model
2026-07-04 09:12:39 -06:00
Jaret Burkett
ba0b3dbb65
Force batch size when bucket is too small by duplicating items in the batch
2026-06-21 20:01:29 -06:00
Jaret Burkett
c730d64478
Added a flag to keep loading the image when latents are cached. Useful for DFE and other methods that target pixelspace losses.
2026-06-15 05:31:48 -06:00
Jaret Burkett
41157b460c
Added ability to set the caption extention in dataset viewer, captioner, and trainer so one dataset can have multiple caption styles in different files with different extensions. Added dataset caption template for a blank ideogram 4 formatted template.
2026-06-06 08:32:24 -06:00
Jaret Burkett
aecd554128
Add sapiens2 matting as a mask generator. Begin transition to model paths and model folders.
2026-05-20 08:56:16 -06:00
Jaret Burkett
af6458d1b5
Enable caching of ACE step latents.
2026-04-28 13:39:20 -06:00
Jaret Burkett
78cf049c29
Add support for ACE-Step 1.5 and ACE-Step 1.5 XL. Also added dataset captioning through the UI. ( #785 )
...
* Base ace step 1.5 xl added. Generating, still wip on training and ui
* Base training code done
* Fix some issues with caching text embeddings. Update sample cards to show audio
* Fix issue with quantizing ace step
* Add album artwork to samples with waveform.
* Cleanup logs
* Add album art endpoint to speed up album art loading
* Made an make video with artwork script
* Make ui handle basic audio models. Make multi line adjustments to the editor and better syntax hilighting.
* Add prompt tagging system for special tagged models.
* prompt tagging processing for ui working.
* Moved default samples to a special file so we can add more when needed and they can be adjusted for a specific model
* Add a captioner job with music captioner that is prepped for use with the ui
* Add basit ui setup for captioning modal and handeling captioning jobs
* Starting captioning job from ui working. Still better management for it.
* Better filtering of job options in the job view for captioning jobs
* Added qwen3 vl as a captioner for images
* Have an indicator when a dataset is being captioned.
* Adjust the way caption jobs look in the queue
* Fix a few issues. Adjust defaults.
* Version bump
* Added ace step to the readme.
2026-04-09 15:02:03 -06:00
Jaret Burkett
7f3309b291
Add support for audo frame count so datasets can have varrying length videos. Varous ltx 2.3 VAE optimizations such as removing tiling articacts, and doing frame split encoding to reduce vram on encoding/decoding.
2026-03-24 12:20:09 -06:00
Jaret Burkett
5642b656b9
Fix audio issues with ltx2 models. Silent codec fails now raised. Auto convert surround sound audio to stereo. Invalidate old caches just to be safe so they recache now.
2026-03-23 20:08:33 +00:00
Jaret Burkett
1ce2428722
Shrink text embeds to max token length for LTX-2. Drastically reduces cached text embedding sizes
2026-01-28 12:54:49 -07:00
Jaret Burkett
73dedbf662
Do caching of latents, first frame and audio when caching latents for LTX2
2026-01-14 11:05:23 -07:00
Jaret Burkett
5b5aadadb8
Add LTX-2 Support ( #644 )
...
* WIP, adding support for LTX2
* Training on images working
* Fix loading comfy models
* Handle converting and deconverting lora so it matches original format
* Reworked ui to habdle ltx and propert dataset default overwriting.
* Update the way lokr saves to it is more compatable with comfy
* Audio loading and synchronization/resampling is working
* Add audio to training. Does it work? Maybe, still testing.
* Fixed fps default issue for sound
* Have ui set fps for accurate audio mapping on ltx
* Added audio procession options to the ui for ltx
* Clean up requirements
2026-01-13 04:55:30 -07:00
Jaret Burkett
ff14cd6343
Fix check for making sure vae is on the right device.
2025-10-21 14:49:20 -06:00
Jaret Burkett
be990630b9
Remove dropout from cached text embeddings even if used specifies it so blank prompts are not cached.
2025-09-26 11:50:53 -06:00
Jaret Burkett
454be0958a
Initial support for qwen image edit plus
2025-09-24 11:39:10 -06:00
Jaret Burkett
390e21bec6
Integrate dataset level trigger words and allow them to be cached. Default to global trigger if it is set.
2025-09-18 03:29:18 -06:00
Jaret Burkett
f699f4be5f
Add ability to set transparent color for control images
2025-09-02 11:08:44 -06:00
Jaret Burkett
5c27f89af5
Add example config for qwen image edit
2025-08-23 18:20:36 -06:00
Jaret Burkett
bf2700f7be
Initial support for finetuning qwen image. Will only work with caching for now, need to add controls everywhere.
2025-08-21 16:41:17 -06:00
Jaret Burkett
bb6db3d635
Added support for caching text embeddings. This is just initial support and will probably fail for some models. Still needs to be ompimized
2025-08-07 10:27:55 -06:00
Jaret Burkett
77dc38a574
Some work on caching text embeddings
2025-07-26 09:22:04 -06:00
Hameer Abbasi
5e86139e0a
Fix NameError.
2025-06-11 15:07:20 +02:00
Hameer Abbasi
c5d6b74fea
Fix caption loading.
2025-06-11 15:05:31 +02:00
Jaret Burkett
22cdfadab6
Added new timestep weighing strategy
2025-06-04 01:16:02 -06:00
Jaret Burkett
1210050ead
Reworked control generator. It is now significantly faster. Also uses better pose model with better license.
2025-05-08 14:35:55 -06:00
Jaret Burkett
c12036df95
Added ability to use short captions from json caption file
2025-04-16 08:32:28 -06:00
Jaret Burkett
d8bdc03256
Allow full control of caption extensions
2025-04-10 07:42:04 -06:00
Jaret Burkett
96ba2fd129
Added methods to the dataloader to automatically generate controls for line, mask, inpainting, depth, and pose.
2025-04-09 13:35:04 -06:00
Jaret Burkett
1d5f387f54
Fix docker command to work better with runpod
2025-03-27 17:44:46 -06:00
Jaret Burkett
45be82d5d6
Handle inpainting training for control_lora adapter
2025-03-24 13:17:47 -06:00
Jaret Burkett
f10937e6da
Handle multi control inputs for control lora training
2025-03-23 07:37:08 -06:00