Zequn Yang†
Yu Miao†
Haotian Ni†
Ziheng Chen†
Chengxiang Huang†
Dongzhan Zhou
Kai Chen
Qi Zhang
Ji-Rong Wen
Yake Wei‡
Di Hu‡,✉
† Equal contribution ‡ Team leader ✉ Corresponding author
Gestalt is a new paradigm of large multimodal model built around multimodal interplay. Guided by a multimodal interplay pyramid — from modality-specific modeling, through cross-modal alignment, to multimodal synergy — Gestalt adopts a unified discrete diffusion framework with an interplay-partitioned architecture, where learnable interplay tokens mediate cross-modal exchange and integration. The name is inspired by Gestalt psychology: the whole is greater than the sum of its parts.
- [2026-09-30] 🎉 We release the model weights and the training & inference code of Gestalt.
- 📄 The arXiv paper is coming soon.
- 📦 SFT data coming soon.
Multimodal intelligence grows from preserving what each modality knows to discovering what they can reveal together. We organize multimodal modeling as a bottom-up progression:
- Modality-Specific Information Modeling — preserve information available only in one modality, providing the foundation for multimodal learning.
- Cross-Modal Redundancy Alignment — align shared information across modalities to establish cross-modal correspondence.
- Multimodal Synergy — combine complementary cues to derive information that neither modality provides alone.
A case study: “I have two dogs. The larger one wears a red collar, while the smaller one wears a blue collar.” The image provides the collar color; the text links collar color to size. Together, they identify the smaller dog.
Gestalt brings vision and language into a shared discrete diffusion framework. The Transformer is divided into two successive zones, with learnable interplay tokens mediating cross-modal information flow throughout:
- Interplay Zone I — Bottleneck Interplay. Separate visual and textual FFNs preserve each modality’s information; direct cross-modal attention is restricted, and exchange is mediated by the interplay tokens.
- Interplay Zone II — Full Multimodal Interplay. The restriction is removed: image, text, and interplay tokens interact through full multimodal self-attention and a shared FFN, integrating complementary information across all tokens.
The two zones instantiate the pyramid as a progression from modality-specific processing, through controlled exchange, to full multimodal integration.
| Stage | Data | Objective |
|---|---|---|
| Multimodal Pretraining | 70M | Joint masked prediction |
| Continual Pretraining | 8M | Conditional masked prediction |
| Supervised Fine-Tuning | 13.7M (+≈2.5M T2I) | Four interplay categories, three-phase curriculum |
The SFT data are organized into redundant, text-unique, visual-unique, and synergistic categories. Across three curriculum phases, all four categories are retained while the sampling emphasis progressively shifts — from consolidating cross-modal alignment and modality-specific capability, to deeper multimodal synergy.
git clone https://github.com/GeWu-Lab/Gestalt.git
cd Gestalt
pip install -r requirements.txthuggingface-cli download GeWu-Lab/Gestalt --local-dir /path/to/checkpointTraining data are Parquet files with a unified schema covering I2T / MMU, T2I, and I2I tasks (see Data Format):
DATA_DIR=/path/to/parquet \
PRETRAINED=/path/to/checkpoint \
NPROC_PER_NODE=8 \
scripts/training/gestalt_sft.shThe checkpoint directory must contain tokenizer.json; training reads the tokenizer directly from PRETRAINED. Defaults: configs/training/gestalt_stage_3.yaml with DeepSpeed ZeRO-2 (DEEPSPEED=none to disable). Omit NPROC_PER_NODE for single-GPU runs. The number of Parquet files must be ≥ the number of distributed ranks.
Multimodal understanding (MMU) — image tokens + question → text:
python scripts/inference/run_mmu.py \
--model /path/to/checkpoint \
--image-tokens image_tokens.npy \
--question "What is in this image?" \
--token-h 32 --token-w 32Text-to-image (T2I) — text → VQVAE tokens:
python scripts/inference/run_t2i.py \
--model /path/to/checkpoint \
--prompt "A cat sitting on a windowsill" \
--output output_tokens.npy \
--lat-h 32 --lat-w 32Tip
T2I uses an empty system prompt by default, matching the training template. If your training data used a fixed non-empty system prompt, pass the same text via --system-prompt at inference.
Training input is a Parquet file (or a directory of Parquet files) with a unified schema:
| Field | I2T | T2I | I2I |
|---|---|---|---|
task_type |
i2t |
t2i |
i2i |
metadata_json |
JSON string | JSON string | JSON string |
conversations |
JSON string / turn list | — | — |
img_tokens |
raw VQVAE IDs | — | — |
system_prompt |
optional | string | string |
user_prompt |
— | string | string |
input_image_tokens |
— | — | raw VQVAE IDs |
answer_image_tokens |
— | raw VQVAE IDs | raw VQVAE IDs |
UniGenBench — Gestalt achieves the strongest overall result among the evaluated models, combining fine-grained semantic alignment with compositional generation.
TIIF-Bench — Gestalt leads across short and long instructions, with a larger advantage when longer prompts introduce more interdependent requirements.
The gains on vision-centric tasks reflect the value of preserving modality-specific information within a unified model.
Gestalt leads the evaluated diffusion-based unified models and achieves language performance comparable to autoregressive models such as Janus-Pro.
Visual and textual tokens are more interleaved in Gestalt, suggesting a more integrated representation space than the evaluated diffusion-based and autoregressive baselines.
Visual-specific and synergistic capability — strong visual-specific performance is paired with leading synergy results among the evaluated diffusion-based models.
Interplay-aware representations — interplay tokens form distinct visual-unique, text-unique, and synergistic structures; the synergy distribution partially bridges the other two, suggesting that these tokens adapt to different information demands.
If you find Gestalt useful for your research, please cite:
@article{gestalt2026,
title = {Gestalt: Large Multimodal Interplay Model},
author = {Yang, Zequn and Miao, Yu and Ni, Haotian and Chen, Ziheng and
Huang, Chengxiang and Zhou, Dongzhan and Chen, Kai and Zhang, Qi and
Wen, Ji-Rong and Wei, Yake and Hu, Di},
year = {2026},
url = {https://github.com/GeWu-Lab/Gestalt}
}This project is released under the Apache 2.0 license.
Zequn Yang, Yake Wei, and Di Hu drove the overall advancement of the project. Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Chengxiang Huang contributed equally to this work. Zequn Yang conducted model pretraining and supervised fine-tuning. Zequn Yang, Yu Miao, and Haotian Ni developed the model architecture and conducted the core experiments. Yu Miao, Ziheng Chen, and Chengxiang Huang contributed to data processing and organization. Yu Miao conducted the evaluation of image generation capabilities, Haotian Ni conducted the multimodal understanding evaluation, and Ziheng Chen conducted the text-only evaluation. Dongzhan Zhou, Kai Chen, Qi Zhang, and Ji-Rong Wen contributed to discussions on the technical design and methodology. Yake Wei, Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Di Hu contributed to writing and revising the manuscript. Di Hu initiated the project. Yake Wei and Di Hu supervised and advised the project.













