
Photorealism
Natural skin texture, controlled light, and editorial detail in a close adult portrait.
Close beauty portrait, wet curls, opaque black fashion top, cobalt background, violet rim light.
A 6B-parameter, 8-NFE image model built for speed, photorealism, bilingual text, and practical hardware.




Source: official Z-Image repository and model card. Performance figures are vendor-reported under stated hardware conditions.
Model facts
Z-Image Turbo is the speed-focused, distilled member of Tongyi-MAI’s Z-Image family. It retains the family’s 6B scale while reducing generation to eight model evaluations.
The base model is the better fit when diversity, negative prompting, controllability, or fine-tuning matters most. Turbo is optimized for fast, polished generation.
Check the official model tableArchitecture
S3-DiT concatenates text, visual semantic, and image VAE tokens into one unified sequence. That single stream is designed to use parameters more efficiently than separate text and image branches.
“Sub-second” is the project’s H800 benchmark—not a promise for every GPU, resolution, or deployment.
Capability evidence
These are original outputs generated through this site’s Z-Image endpoint. Each example tests a distinct part of the official capability claims.

Natural skin texture, controlled light, and editorial detail in a close adult portrait.
Close beauty portrait, wet curls, opaque black fashion top, cobalt background, violet rim light.

The output preserves the requested counts: three lemons, two lilies, one ribbon, and one sphere.
Exactly 3 lemons · 2 calla lilies · 1 red ribbon · 1 silver sphere.

Chinese and English headline rendering is a stated strength; punctuation still deserves review.
Exact Chinese headline “造相未来” and English subheading “IMAGINE IN SECONDS”.

Large-scale architecture remains coherent across structure, weather, reflections, and human scale.
Coastal museum, white concrete curves, dark cliff, wet plaza, tiny visitors, misty ocean.
Recommended inference
The official example notes that setting 9 inference steps results in 8 DiT forward passes. It also recommends guidance scale 0 for Turbo.
Practical limits
FAQ
Z-Image Turbo is the distilled text-to-image variant in Alibaba Tongyi-MAI’s 6B-parameter Z-Image family. The official model table describes it as an 8-NFE generation model trained with pre-training, supervised fine-tuning, and reinforcement-learning post-training.
The project reports sub-second inference on an enterprise NVIDIA H800. That is a hardware-specific benchmark, not a universal promise. Local latency changes with GPU, precision, attention backend, compilation, offloading, and output resolution.
The official paper and repository say it fits within 16GB VRAM consumer hardware. Quantization or CPU offloading may reduce peak GPU memory further, but usually changes speed and sometimes output characteristics.
The base Z-Image model prioritizes diversity, controllability, and fine-tuning. Turbo is distilled for much faster generation and very high visual quality, but the official table rates its diversity as low and fine-tunability as not applicable.