MageTrail - 2.8B Image Model based on Microsoft's MageFlow 4B

ALL PREVIEW IMAGES HAVE COMFYUI METADATA.

I. Introduction

MageTrail is a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, now 2.8B custom architecture, using a diversity maximized condensed 41k images dataset as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset.

My second attempt at fine-tuning an image model on larger scale, this finetune aim to prove to the open source community on MageFlow 4B having good potential as an architecture for further finetuning.

~ V0.1 to V0.3 have spent 593 dollars and has ultimately shown the model potential in quickly learning and adapting booru concept and tags to its knowledge base, but has outgrown its limited 41k dataset, V0.4 onwards will be artist style tuning at a larger data scale.

~ The architecture behind MageFlow 4B shows good promise for further investment:

  • Being 15-20% faster than NVIDIA Cosmos2/Anima on inference despite being 2 billion parameters larger
  • From V0.3 onwards, new discovery have prove the model is completely compatible with Flux2VAE (the current best open source VAE available) in training with some code changes, meaning the architecture is now COMPLETELY a Flux2VAE arch, without needing any money to do a realignment tune to new vae.
  • Having a decent Qwen 3 VL 4B Text Encoder
  • 256-2048 pixels resolution native support
  • Being fairly quick to learn and adapt to new knowledge without any knowledge forgetting

BIG GOAL: Due to https://huggingface.co/Muinez/mage-flow-ft , I now completely believe in MageFlow being the next best small scale arch to replace Anima in the Illustration sphrere, the model has shown to have the ability to learn concepts/characters/styles efficiently while being a superior architecture. I pledge to finetune a Anima successor if given 25k-100k dollars of funding from the open source community/any generous patrons.

SMALL GOAL: Gather funding of 1000-2000~ dollars to finetune 10k-100k artist styles into the model (200k-3 million images finetune, planned V0.4-V1) and help further with convergence/aesthetic improvement if the aforementioned big goal above hasn't been reached.

Any donation will help with achieving this goal, you can do so through:

Crypto (Prefered, cause Kofi/Paypal money transfer time is ass and they take a big cut)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)

12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)

FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)

Please handle your money carefully and make sure the address you're sending to is correct.

Ko-fi

https://ko-fi.com/talanartvn

II. Model Details

Original Base model Microsoft's MageFlow 4B
Modified Base model MageTrail 2.8B Adaln Approximation
Method Full Finetune
Trainer My Diffusion-Pipe fork
Hardware H100 HBM3 80GB, rented with Banodoco sponsor and donation from the community
Total training time 64 (v0.1) + 92 (v0.2) + 269 (V0.3) of H100 hours
Total samples seen 8.1~ million (v0.1-v0.3)
Training resolutions 1024Β²

Training run

Version 0.1 (initial 20 epoch run β†’ extended 10 epoch run)

Budget: 130~ dollars (25-30 lost due to experiments and mistakes)

Full config: Training and Dataset

  • Learning rate: 7e-6
  • LR scheduler: Warmup -> Constant -> REX to 0e-7
  • Precision: Full BF16
  • Optimizer: AdamW8bit with Kahan summation (to offset BF16 precision roundoff)
  • Weight decay: 0.02
  • Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

Version 0.2 (70 epoch continuation)

Budget: 143~ dollars

Full config: Same as v0.1, just with changed lr due to only using x4 gpu instead of x8

  • Learning rate: 5e-6
  • LR scheduler: Warmup -> Constant -> REX to 0e-7
  • Precision: Full BF16
  • Optimizer: AdamW8bit with Kahan summation (to offset BF16 precision roundoff)
  • Weight decay: 0.02
  • Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

Version 0.3 (100 epoch final continuation)

Budget: 350~ dollars

Full config: Same as v0.2, but on x1 cheap H100 to save cost

  • Learning rate: 5e-6
  • LR scheduler: Warmup -> Constant -> REX to 0e-7
  • Precision: Full BF16
  • Optimizer: AdamW8bit with Kahan summation and fixed weight decay feature (it was actually broken before, really bad oversight from me)
  • Weight decay: 0.02
  • Timestep sampling: Logit-Normal, shift 6, sigmoid scale 1.0

Additional training features

  • Tag dropout: 10%
  • Caption dropout: 5%
  • Mixed captions at 25/25/25/25 ratio (tags only, NL only, tags-nl, nl-tags)
  • Tag shuffle
  • Caption shuffle
  • Artist trigger attribution system

III. Recommended Settings

These are the settings used for the sample images above (ComfyUI):

  • Shift: 5.0 (or ComfyUI default)
  • Steps: 50 (or 30 step for faster gen)
  • CFG: 8 (or 5 for faster gen)
  • Sampler: er_sde or euler
  • Scheduler: simple or beta

These are just my usual settings and workflow β€” feel free to experiment.

Artist Trigger: This model use the Drawn by artistname trigger, if you want to use the model built-in artist tag, please always put one at the start of your prompt. (Currently V0.1-V0.3 barely support any artist or characters though)


IV. Dataset

Booru-Essence-2026 41k images

Originally created by Lodestone Rock, the dataset was updated to 2026 tag standard and captioned with SOTA API captioners, see dataset repo for details. Model training, Dataset and Captioning tooling lives in the utils folder of the training repo.


VI. Notes from the Training Diary

Full training diary: MageTrail-diary


VII. License

This model is a Derivative of Microsoft's MageFlow and is distributed under the same MIT License as the base model, with no additional restrictions.


VIII. Acknowledgments

Beeg thanks to:

  • Banodoco and their Discord β€” Their 88.77 dollar initial grant and further support on future requests made this project possible, the biggest thanks to them
  • Lodestone Rock β€” Creator of the original version of the booru essence dataset that this model is trained on
  • Motimalu β€” Inspiration behind finetuning practices and configs
  • Bluvoll β€” Creator of the modified 2.8B arch, diffusion-pipe fork derived from to use for training, and general training advice
  • Anzhc β€” general training advice
  • Nruaif β€” diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
  • Astromahdi β€” jupyter workspace and storage where I processed and store the dataset
  • Heato-Red β€” Designer of model page logo
  • animetimm/DeepGHS β€” Danbooru tagging model
  • RedRocket β€” E621 tagging model
  • Format inspired by Motimalu's Kirazuri diary
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for RicemanT/MageTrail

Finetuned
(1)
this model
Finetunes
1 model

Dataset used to train RicemanT/MageTrail