Sign In

MageTrail 2.8B (Danbooru/E621 Finetune) Based on MageFlow

Updated: Sep 30, 2026

base modeldanboorue621mage-flow

Download

1 variant available

bf16 SafeTensor

MageTrail-V0.3.safetensors

BF16, good balance • 5.56 GB

Verified:

You need these files to run this model. We'll show the best match for your preferences.

VAE

flux2-vae.safetensors • 320.64 MB

Verified:

4B • qwen3vl_4b_bf16.safetensors

8.27 GB

Downloads your preferred variants

Type
Checkpoint Trained
Stats

21

Reviews
Published

Sep 30, 2026

Base Model

MageFlow

Training
Epochs: 100
Hash
AutoV2
011CF64118
default creator card background decoration
Followers - 179

179

Likes - 722

722

License:

ComfyUI_temp_vzmql_00015_.png

## I. Introduction

MageTrail is a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, now 2.8B custom architecture, using a diversity maximized condensed 41k images dataset as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset.

(which would have cost 50k+ dollars, no thanks).

Starting from V0.25 onward, the model is now a completely new architecture due to being trimmed down to 2.8 billion parameters. This help the model become near SDXL/Anima range of availability for all users, with faster gen time and lower vram usage for finetuning/LoRA training. The trimming was done by Bluvoll/MageTrail-Shared-AdaLN-SVD.

Please scroll down to the generation recommendation section to download the required custom node.

The model is now also supported for easy consumer LoRA training/Finetuning through bluvoll/mage-flow-trainer: A trainer for Mage-Flow with SDNQ support.

My second attempt at fine-tuning an image model on larger scale, this finetune aim to prove to the open source community on MageFlow 4B having good potential as an architecture for further finetuning.

~ V0.1 to V0.3 have spent 593 dollars and has ultimately shown the model potential in quickly learning and adapting booru concept and tags to its knowledge base, but has outgrown its limited 41k dataset, V0.4 onwards will be artist style tuning at a larger data scale.

~ The architecture behind MageFlow 4B shows good promise for further investment:

* Being 15-20% faster than NVIDIA Cosmos2/Anima on inference despite being 2 billion parameters larger

* From V0.3 onwards, new discovery have prove the model is completely compatible with Flux2VAE (the current best open source VAE available) in training with some code changes, meaning the architecture is now COMPLETELY a Flux2VAE arch, without needing any money to do a realignment tune to new vae.

* Having a decent Qwen 3 VL 4B Text Encoder

* 256-2048 pixels resolution native support

* Being fairly quick to learn and adapt to new knowledge without any knowledge forgetting

Future goal for the project:

***BIG GOAL***: Due to Muinez/mage-flow-ft , I now completely believe in MageFlow being the next best small scale arch to replace Anima in the Illustration sphrere, the model has shown to have the ability to learn concepts/characters/styles efficiently while being a superior architecture. I pledge to finetune a Anima successor if given 25k-100k dollars of funding from the open source community/any generous patrons.

SMALL GOAL: Gather funding of 1000-2000~ dollars to finetune 10k-100k artist styles into the model (200k-3 million images finetune, planned V0.4-V1) and help further with convergence/aesthetic improvement if the aforementioned big goal above hasn't been reached.

Any donation will help with achieving this goal, you can do so through:

Crypto (Prefered, cause Kofi/Paypal money transfer time is ass and they take a big cut)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)

12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)

FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)

Please handle your money carefully and make sure the address you're sending to is correct.

Ko-fi

https://ko-fi.com/talanartvn

## II. Model Details

Base model: Microsoft's MageFlow 4B T2I

Modified Base model: MageTrail-Shared-AdaLN-SVD

Method: Full-Finetune

Trainer: https://github.com/RicemanT/diffusion-pipe-mageflow-ft

Hardware: H100 HBM3 80GB, rented with Banodoco sponsor and donation from the community

Total training time: 64 (v0.1) + 92 (v0.2) + 269 (V0.3) of H100 hours

Total samples seen: 8.1~ million (v0.1-v0.3)

Training resolutions : 1024²

## Training run

## Version 0.1 (initial 20 epoch run → extended 10 epoch run)

Budget: 130~ dollars (25-30 lost due to experiments and mistakes)

Full config: Training and Dataset

- Learning rate: 7e-6

- LR scheduler: Warmup -> Constant -> REX to 0e-7

- Precision: Full BF16

- Optimizer: AdamW8bit with Kahan summation (to offset BF16 precision roundoff)

- Weight decay: 0.02

- Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

## Version 0.2 (70 epoch continuation)

Budget: 143~ dollars

Full config: Same as v0.1, just with changed lr due to only using x4 gpu instead of x8

- Learning rate: 5e-6

- LR scheduler: Warmup -> Constant -> REX to 0e-7

- Precision: Full BF16

- Optimizer: AdamW8bit with Kahan summation (to offset BF16 precision roundoff)

- Weight decay: 0.02

- Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

### Version 0.3 (100 epoch final continuation)

Budget: 350~ dollars

Full config: Same as v0.2, but on x1 cheap H100 to save cost

- Learning rate: 5e-6

- LR scheduler: Warmup -> Constant -> REX to 0e-7

- Precision: Full BF16

- Optimizer: AdamW8bit with Kahan summation and fixed weight decay feature (it was actually broken before, really bad oversight from me)

- Weight decay: 0.02

- Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

## Additional training features

- Tag dropout: 10%

- Caption dropout: 5%

- Mixed captions at 25/25/25/25 ratio (tags only, NL only, tags-nl, nl-tags)

- Tag shuffle

- Caption shuffle

- Artist trigger attribution system

## III. Recommended Settings

From V0.25 onwards, the model require downloading the bluvoll/ComfyUI-MageFlow-Compressed custom node. The model is not supported in any other UI currently. These are the settings used for the sample images above (ComfyUI):

- Shift: 5.0 (or ComfyUI default)

- Steps: 50 (or 30 for faster gen speed)

- CFG: 8 (or 5 for faster gen speed with step 30 above)

- Sampler: er_sde or euler

- Scheduler: simple or beta

These are just my usual settings and workflow — feel free to experiment.

Artist Trigger: This model use the Drawn by artistname trigger, if you want to use the model built-in artist tag, please always put one at the start of your prompt. (Currently V0.1 barely support any artist or characters though)

Tagging System: This model support both danbooru and e621 tag prompting, although currently due to the model being severely undertrained alot of tags and concept are not working properly yet. Tags only and NL only prompt both work, but the model perform better with mixed prompting and clearer/slightly longer prompt.

Example prompt:
Drawn by nyatcha, m200 (girls' frontline), 1girl, :t, backpack, bag, binoculars, building, city, crane (machine), crossed bangs, crossed legs, feet out of frame, grey jacket, headset, holding, holding binoculars, jacket, long sleeves, messy hair, pout, short hair, sidelocks, sitting, solo, star (symbol). Set in a panoramic, dense futuristic cityscape during dusk, M200 from Girls' Frontline sits perched near an elevated railing on the right side of the frame with her legs crossed. She has short, messy grey hair with crossed bangs and sidelocks, wearing an olive-grey communication headset over her ears. She is dressed in a bulky, layered dark grey military jacket with long sleeves and carries a tactical backpack with mechanical gear protruding over her shoulder, resting both hands in front of her while holding a pair of compact black binoculars. The bustling urban backdrop features towering, dark multi-story architecture filled with glowing square windows, exposed steel framing, rooftop antennas, and a large red construction crane against a hazy, overexposed bright sky. Cars with gleaming roofs pass below the railing amidst the muted industrial color palette of beige, black, and amber lighting.

## IV. Dataset

Booru-Essence-2026-41k images

Originally created by Lodestone Rock, the dataset was updated to 2026 tag standard and captioned with SOTA API captioners, see dataset repo for details.

Model training, Dataset and Captioning tooling lives in the utils folder](https://github.com/RicemanT/model-training-configs/tree/main/diffusion-pipe/MageFlow/utils) of the training repo.

---

## VI. Notes from the Training Diary

Full training diary: MageTrail-diary

---

## VII. License

This model is a Derivative of Microsoft's MageFlow and is distributed under the same [MIT License](https://choosealicense.com/licenses/mit/) as the base model, with no additional restrictions.

---

## VIII. Acknowledgments

Beeg thanks to:

- [Banodoco](https://www.banodoco.ai/) and [their Discord](https://discord.gg/yzwcNaSEz) — Their 88.77 dollar initial grant and future made this project possible, the biggest thanks to them

- [Lodestone Rock](https://huggingface.co/lodestones) — Creator of the original version of the dataset that this model is trained on

- [Motimalu](https://civitai.red/user/motimalu) — Inspiration behind finetuning practices and configs

- [Bluvoll](https://github.com/bluvoll/diffusion-pipe) — diffusion-pipe fork derived from to use for training, and general training advice

- [Anzhc](https://huggingface.co/Anzhc) — general training advice

- [Nruaif](https://huggingface.co/Shio-Koube) — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice

- [Astromahdi](https://gpu.garden/) — jupyter workspace where I processed and store the dataset

- [Heato-Red](https://www.reddit.com/user/heato-red/) — Designer of model page logo

- [animetimm/DeepGHS](https://huggingface.co/animetimm/convnextv2_huge.dbv4-full) — Danbooru tagging model

- [RedRocket](https://huggingface.co/RedRocket/Hydra) — E621 tagging model