Z-Image Turbo, explained.

A 6B-parameter, 8-NFE image model built for speed, photorealism, bilingual text, and practical hardware.

Photoreal close-up adult portrait generated with Z-Image
Transparent headphones product image generated with Z-Image
Coastal museum architecture generated with Z-Image
Fantasy whale in clouds generated with Z-Image
6B
parameters
Model-family size
8
NFEs
Distilled Turbo path
<16GB
VRAM
Official consumer target
Apache 2.0
license
Code and released weights

Source: official Z-Image repository and model card. Performance figures are vendor-reported under stated hardware conditions.

Model facts

What Turbo actually is.

Z-Image Turbo is the speed-focused, distilled member of Tongyi-MAI’s Z-Image family. It retains the family’s 6B scale while reducing generation to eight model evaluations.

The base model is the better fit when diversity, negative prompting, controllability, or fine-tuning matters most. Turbo is optimized for fast, polished generation.

Check the official model table
Official comparison
Z-Image
Z-Image Turbo
Official model-zoo steps
50
8
CFG
Yes
No · guidance 0
Visual quality
High
Very high
Output diversity
Medium
Low
Fine-tunability
Easy
N/A

Architecture

One stream, less overhead.

S3-DiT concatenates text, visual semantic, and image VAE tokens into one unified sequence. That single stream is designed to use parameters more efficiently than separate text and image branches.

“Sub-second” is the project’s H800 benchmark—not a promise for every GPU, resolution, or deployment.

Text tokens
Visual semantic tokens
Image VAE tokens
Concatenate
Unified sequence
S3-DiT
Image

Capability evidence

Where Turbo is strongest.

These are original outputs generated through this site’s Z-Image endpoint. Each example tests a distinct part of the official capability claims.

Close photorealistic portrait of an adult woman in rain light

Photorealism

Natural skin texture, controlled light, and editorial detail in a close adult portrait.

Close beauty portrait, wet curls, opaque black fashion top, cobalt background, violet rim light.

Still life with exactly three lemons, two white flowers, one red ribbon, and one silver sphere

Prompt adherence

The output preserves the requested counts: three lemons, two lilies, one ribbon, and one sphere.

Exactly 3 lemons · 2 calla lilies · 1 red ribbon · 1 silver sphere.

Black, cobalt, and violet poster reading 造相未来 and IMAGINE IN SECONDS

Bilingual text

Chinese and English headline rendering is a stated strength; punctuation still deserves review.

Exact Chinese headline “造相未来” and English subheading “IMAGINE IN SECONDS”.

Curved modern museum embedded into a coastal cliff after rain

Composition

Large-scale architecture remains coherent across structure, weather, reflections, and human scale.

Coastal museum, white concrete curves, dark cliff, wet plaza, tiny visitors, misty ocean.

Recommended inference

The official Turbo defaults.

Pipeline
ZImagePipeline
Precision
bfloat16
Inference steps
9 · 8 forwards
Guidance scale
0.0

The official example notes that setting 9 inference steps results in 8 DiT forward passes. It also recommends guidance scale 0 for Turbo.

Practical limits

Know the tradeoffs.

  1. 01The official model table rates Turbo output diversity as low; use the base model when exploration matters more than latency.
  2. 02Turbo is marked N/A for fine-tuning, while the base Z-Image checkpoint is designed for downstream development.
  3. 03Text, hands, anatomy, object counts, and crowded scenes should still be reviewed before production use.
  4. 04Local speed depends on GPU, precision, attention backend, compilation, offloading, and resolution.

FAQ

Direct answers.

What is Z-Image Turbo?+

Z-Image Turbo is the distilled text-to-image variant in Alibaba Tongyi-MAI’s 6B-parameter Z-Image family. The official model table describes it as an 8-NFE generation model trained with pre-training, supervised fine-tuning, and reinforcement-learning post-training.

How fast is Z-Image Turbo?+

The project reports sub-second inference on an enterprise NVIDIA H800. That is a hardware-specific benchmark, not a universal promise. Local latency changes with GPU, precision, attention backend, compilation, offloading, and output resolution.

How much VRAM does Z-Image Turbo need?+

The official paper and repository say it fits within 16GB VRAM consumer hardware. Quantization or CPU offloading may reduce peak GPU memory further, but usually changes speed and sometimes output characteristics.

What is the difference between Z-Image and Z-Image Turbo?+

The base Z-Image model prioritizes diversity, controllability, and fine-tuning. Turbo is distilled for much faster generation and very high visual quality, but the official table rates its diversity as low and fine-tunability as not applicable.

See what eight steps can make.

Z-Image Turbo: Model Facts, Speed, VRAM & Examples