πŸ–ΌοΈ POCKET-Image-Zimage β€” Character-Perfect Text in Any Image

POCKET-Image is VIDRAFT's on-device text-in-image system. It places character-perfect text in any language β€” Korean Β· δΈ­ζ–‡ Β· ζ—₯本θͺž Β· Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ© Β· ΰΉ„ΰΈ—ΰΈ’ Β· Latin and more β€” onto any scene you describe. Built on the open, commercial-friendly Z-Image foundation and VIDRAFT's POCKET-Core engine.

Where global models smear non-Latin scripts ("μ•ˆλ…•ν•˜μ„Έμš”" β†’ "μ•ˆγ…κΈ°"), POCKET-Image renders every glyph exactly.

✨ Highlights

  • βœ… 100% accurate text, any language β€” RTL (Arabic) and complex shaping (Thai, Indic) handled
  • πŸ–ΌοΈ Any background from a prompt β€” text is optional (empty β‡’ pure image)
  • ⚑ No GPU, no NPU β€” runs on plain CPU + RAM
  • πŸͺΆ From 4.5 GB footprint Β· πŸ’» Windows Β· macOS Β· Linux (RTX or Apple Silicon)
  • πŸ”’ Fully local β€” no cloud, no account

🎨 Try it

πŸ’» On-device footprint (measured)

Machine Peak memory
RTX 3050 / 4060 8.6 GB
RTX 2060 (6 GB, offloaded) 4.5 GB
MacBook (M1 Pro–M4, 16 GB+) 13.4 GB

Base model: Tongyi-MAI/Z-Image (Apache-2.0). Engine: VIDRAFT POCKET-Core.


🧩 The POCKET Family β€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

πŸ“š Full POCKET collection

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FINAL-Bench/POCKET-Image-Zimage

Finetuned
(64)
this model

Space using FINAL-Bench/POCKET-Image-Zimage 1

Collection including FINAL-Bench/POCKET-Image-Zimage