Skip to content

feat: add LLaDA-Image support - #1968

Merged
leejet merged 8 commits into
leejet:masterfrom
fszontagh:feat/llada-image
Sep 20, 2026
Merged

leejet merged 8 commits into
leejet:masterfrom
fszontagh:feat/llada-image

Conversation

@fszontagh

@fszontagh fszontagh commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds text-to-image and instruction-guided editing for LLaDA-Image (Apache-2.0): a 6B Lumina2/Z-Image-style NextDiT conditioned by a LLaDA2-MoE diffusion-LLM text encoder, reusing the Flux.2 VAE.

Most of it reuses what is already here. The DiT instantiates ZImage::JointTransformerBlock and FinalLayer with identical hyperparameters, and the text encoder is a new LLMArch::LLADA2_MOE on the existing MoE path. New are the QueryFormer, the text projection, and SigVQ for editing, plus the llada_image reference-image preset and sigma schedule. Editing runs the clean reference and the noisy target in one sequence; adaLN is linear in the timestep embedding, so feeding a per-token embedding selects the right modulation without new modulation machinery.

The LLaDA2 vocabulary is not embedded, so --tokenizer with the model's tokenizer.json is required, per the embedded-data allowlist in CONTRIBUTING.md.

One change outside the model: tensor_should_be_converted now skips *_pad_token. LLaDA-Image stores those around 7.7e24, which overflows both f16 storage and q8_0's f16 block scale, so every converted GGUF rendered blank white without it.

Docs in docs/llada_image.md; rows added to README.md and docs/edit.md.

Related Issue / Discussion

Closes #1937

Additional Information

Verified on an RTX 3060 12GB, CUDA. Both checkpoints work: LLaDA-Image-Turbo does text-to-image and editing at 512x512 and 1024x1024, and the 50-step base model at --steps 50 --cfg-scale 5 does both at 1024x1024. At 512x512 the base model's editing returns the reference almost unchanged, which the docs note.

--max-vram 6, --max-vram 4 and --max-vram 3 all produce byte-identical output to unconstrained execution.

Not implemented: generation_mode="vq", where the text encoder block-diffusion-decodes VQ tokens before diffusion. Text-to-image and editing do not use that path.

Checklist

Comment thread src/tokenizers/vocab/llada2_merges.hpp Outdated
@fszontagh
fszontagh force-pushed the feat/llada-image branch 2 times, most recently from c1e7233 to f642157 Compare September 13, 2026 17:19
@fszontagh
fszontagh marked this pull request as draft September 13, 2026 18:05
@fszontagh
fszontagh force-pushed the feat/llada-image branch 2 times, most recently from 146b53f to 7a19970 Compare September 14, 2026 16:28
@fszontagh

Copy link
Copy Markdown
Contributor Author

Both done: vocab and merges are on a single line now, and the helper scripts are removed.

Since the merge script is gone, I converted the weights for both checkpoints and uploaded them, so the docs link ready-made files instead of asking people to build their own. Links are in docs/llada_image.md.

@fszontagh
fszontagh marked this pull request as ready for review September 14, 2026 18:41
@fszontagh

Copy link
Copy Markdown
Contributor Author

Rebased on master and dropped the embedded vocab, following #1974. LLaDA2 now needs --tokenizer with the model's tokenizer.json, and fails with a clear message if it's missing. Also fixed the encode() call for the new signature.

Output is pixel-identical to the old embedded-vocab build.

Updated the PR description too.

@fszontagh
fszontagh requested a review from leejet September 16, 2026 19:17
@leejet
leejet merged commit 15f335d into leejet:master Sep 20, 2026
9 checks passed
danielhanchen added a commit to unslothai/stable-diffusion.cpp that referenced this pull request Sep 21, 2026
* fix: preserve "token_refiner" token for MiniMax H3 LoRAs (leejet#1864)

* fix: fail with a message when MiniMax-H3 is run in img_gen mode (leejet#1863)

* feat: support INT8 ConvRot safetensors (leejet#1857)

* fix: replace free_compute_buffer with runner_done in vae (leejet#1872)

* sync: update ggml (leejet#1873)

* fix(ci): trigger builds for ggml updates

* feat: add taeh3 support (leejet#1874)

* fix: prevent gallocr hash overflow in tiny graph-cut segments (leejet#1880)

* fix: re-clamp streaming VRAM budget to currently free memory (leejet#1878)

* fix: mark graph cuts with both a prefix and a suffix (leejet#1883)

* fix: make max_order of lms sampler configurable (leejet#1885)

* fix: guard against missing sampler/scheduler names (leejet#1887)

* chore: format code

* fix: use sd_get_preview_interval() (leejet#1907)

* feat: configurable image / video compression (leejet#1909)

* feat: support standard Qwen3-VL weights for MiniMax-H3 (leejet#1910)

* feat: load scaled FP8 weights without upfront conversion (leejet#1913)

* fix: match exact weights in LLM config detection (leejet#1923)

* feat: add LTX-2.5 support (leejet#1893)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: correct MiniMax H3 reference audio encoding (leejet#1886)

* fix: correct MiniMax H3 audio Euler steps (leejet#1908)

* feat: use backend-native FP8 matmul when supported (leejet#1916)

* sync: update ggml

* feat: additional `--preview-interval` values (leejet#1915)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support numbering for preview images (leejet#1895)

* fix: use carrier sampling for MiniMax H3 audio (leejet#1924)

* feat: generalize temporal tiling across video VAEs (leejet#1926)

* feat: prefetch streamed layers during compute (leejet#1905)

Co-authored-by: leejet <leejet714@gmail.com>

* refactor: unify runner lifecycles and weight residency (leejet#1940)

* feat: add verbose logging and log-level selection (leejet#1941)

* feat: enable single-GPU auto-fit with tiered parameter placement (leejet#1942)

* fix: reuse graph cut plans across CFG passes (leejet#1943)

* refactor: split ggml extensions and move implementations to cpp files (leejet#1945)

* fix: preserve K-quantized embedding weights (leejet#1936)

* fix: correct SDXL embeddings loading (leejet#1939)

* refactor: unify model source and weight lifecycle management (leejet#1956)

* docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_extend assert) (leejet#1950)

* refactor: split generation pipeline out of stable-diffusion.cpp (leejet#1957)

* fix: enable VAE decode tiling fallback without auto-fit (leejet#1932)

* fix: preserve BF16 embedding weights for get_rows (leejet#1959)

* fix: handle invalid option numbers (leejet#1961)

* feat: expose the loaded model version name through the public API (leejet#1962)

* feat: add SenseNova U1.5 support (leejet#1935)

* fix: reuse graph plans when scale parameters change (leejet#1963)

* feat: add linear and attention scale overrides (leejet#1964)

* feat: preserve explicit backend assignments during auto-fit (leejet#1967)

* fix: guard GPU memory capacity and propagate encoding failures (leejet#1958)

* fix: bound plain-text runs in parse_prompt_attention regex (leejet#1919)

* feat: add Wan2.2 S2V (audio+img-to-video) support (leejet#1925)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: validate vision projector output dim against LLM hidden size (leejet#1918)

* fix: resolve MSVC narrowing conversion warnings (leejet#1969)

* feat: Add generation parameters into video metadata (leejet#1901)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support external Hugging Face tokenizer JSON files (leejet#1973)

* refactor: require external Gemma 2 and GPT-OSS tokenizers (leejet#1974)

* feat: support Brownian tree noise in all noise injection samplers (leejet#1899)

* fix: use tokenizer-specific pre-tokenization rules (leejet#1975)

* fix: remove vision_model. from ununsed tensors (leejet#1983)

* perf: eliminate temporary allocations in Philox rounds (leejet#1982)

* fix: honor flash attention flag in LLM text encoder attention (leejet#1987)

* refactor: remove obsolete unused tensor filtering (leejet#1984)

* perf: pad small attention heads to 64 for MMA Flash Attention (leejet#1992)

* perf: update ggml for faster direct convolutions (leejet#1993)

* perf: accelerate VAE direct 3D convolutions (leejet#1996)

* fix: propagate CUDA driver dependency to shared library consumers

* fix: prevent clip_preprocess center crop from exceeding the resized image (leejet#1995)

* perf: reduce CPU overhead in graph execution and sampling (leejet#1997)

* perf: parallelize host tensor elementwise and broadcast ops (leejet#1998)

* feat: support building with upstream ggml (leejet#1999)

* feat: add Qwen Image 2.1 support (leejet#1994)

* feat: restore legacy fp8 handling when building with upstream ggml (leejet#2001)

* fix: avoid passing ggml logs as format strings (leejet#2002)

* feat: add LLaDA-Image support (leejet#1968)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: add native CUDA SageAttention support (leejet#2005)

* fix: avoid narrowing conversion in SigVQ patch embedding and format code

* docs: update CONTRIBUTING.md

---------

Co-authored-by: stduhpf <stephduh@live.fr>
Co-authored-by: leejet <leejet714@gmail.com>
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
Co-authored-by: fszontagh <51741446+fszontagh@users.noreply.github.com>
Co-authored-by: Wagner Bruna <wbruna@users.noreply.github.com>
Co-authored-by: vmobilis <75476228+vmobilis@users.noreply.github.com>
Co-authored-by: Piotr Wilkin (ilintar) <ilintar@gmail.com>
Co-authored-by: jk212h20 <101200018+jk212h20@users.noreply.github.com>
Co-authored-by: assouan <750048+assouan@users.noreply.github.com>
Co-authored-by: nan <zjn32202153@gmail.com>
Co-authored-by: Hmission <62598659+Hmission@users.noreply.github.com>
Co-authored-by: LED-M <105789115+xledx@users.noreply.github.com>
Co-authored-by: Maphist0 <28743569+Maphist0@users.noreply.github.com>
Co-authored-by: George <35490284+noctrex@users.noreply.github.com>
Co-authored-by: Санька Четвёртый <CAHbKA-IV@mail.ru>
Co-authored-by: Lin Xuhao <linxuhao84@gmail.com>
Co-authored-by: Fabrice Aneche <akhenakh@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] LLaDA-Image 6B

2 participants