Skip to content

Latest commit

 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Gestalt

Gestalt: Large Multimodal Interplay Model

Website Paper Code Model Data

Zequn Yang†  Yu Miao†  Haotian Ni†  Ziheng Chen†  Chengxiang Huang†
Dongzhan Zhou  Kai Chen  Qi Zhang  Ji-Rong Wen  Yake Wei‡  Di Hu‡,✉

† Equal contribution ‡ Team leader ✉ Corresponding author

Gestalt is a new paradigm of large multimodal model built around multimodal interplay. Guided by a multimodal interplay pyramid — from modality-specific modeling, through cross-modal alignment, to multimodal synergy — Gestalt adopts a unified discrete diffusion framework with an interplay-partitioned architecture, where learnable interplay tokens mediate cross-modal exchange and integration. The name is inspired by Gestalt psychology: the whole is greater than the sum of its parts.

Gestalt teaser

📢 News

  • [2026-09-30] 🎉 We release the model weights and the training & inference code of Gestalt.
  • 📄 The arXiv paper is coming soon.
  • 📦 SFT data coming soon.

🧩 The Multimodal Interplay Pyramid

Multimodal intelligence grows from preserving what each modality knows to discovering what they can reveal together. We organize multimodal modeling as a bottom-up progression:

The multimodal interplay pyramid

  • Modality-Specific Information Modeling — preserve information available only in one modality, providing the foundation for multimodal learning.
  • Cross-Modal Redundancy Alignment — align shared information across modalities to establish cross-modal correspondence.
  • Multimodal Synergy — combine complementary cues to derive information that neither modality provides alone.

A case study: “I have two dogs. The larger one wears a red collar, while the smaller one wears a blue collar.” The image provides the collar color; the text links collar color to size. Together, they identify the smaller dog.

Case study: multimodal synergy

🏗️ Architecture: From Controlled Exchange to Full Interplay

Gestalt brings vision and language into a shared discrete diffusion framework. The Transformer is divided into two successive zones, with learnable interplay tokens mediating cross-modal information flow throughout:

Gestalt architecture

  • Interplay Zone I — Bottleneck Interplay. Separate visual and textual FFNs preserve each modality’s information; direct cross-modal attention is restricted, and exchange is mediated by the interplay tokens.
  • Interplay Zone II — Full Multimodal Interplay. The restriction is removed: image, text, and interplay tokens interact through full multimodal self-attention and a shared FFN, integrating complementary information across all tokens.

The two zones instantiate the pyramid as a progression from modality-specific processing, through controlled exchange, to full multimodal integration.

📚 Training: Multimodal Pretraining and Interplay Curriculum

Stage Data Objective
Multimodal Pretraining 70M Joint masked prediction
Continual Pretraining 8M Conditional masked prediction
Supervised Fine-Tuning 13.7M (+≈2.5M T2I) Four interplay categories, three-phase curriculum

The SFT data are organized into redundant, text-unique, visual-unique, and synergistic categories. Across three curriculum phases, all four categories are retained while the sampling emphasis progressively shifts — from consolidating cross-modal alignment and modality-specific capability, to deeper multimodal synergy.

Three-phase SFT curriculum

🔥 Quick Start

1️⃣ Environment

git clone https://github.com/GeWu-Lab/Gestalt.git
cd Gestalt
pip install -r requirements.txt

2️⃣ Download the Model

huggingface-cli download GeWu-Lab/Gestalt --local-dir /path/to/checkpoint

3️⃣ Training (Stage-III SFT)

Training data are Parquet files with a unified schema covering I2T / MMU, T2I, and I2I tasks (see Data Format):

DATA_DIR=/path/to/parquet \
PRETRAINED=/path/to/checkpoint \
NPROC_PER_NODE=8 \
scripts/training/gestalt_sft.sh

The checkpoint directory must contain tokenizer.json; training reads the tokenizer directly from PRETRAINED. Defaults: configs/training/gestalt_stage_3.yaml with DeepSpeed ZeRO-2 (DEEPSPEED=none to disable). Omit NPROC_PER_NODE for single-GPU runs. The number of Parquet files must be ≥ the number of distributed ranks.

4️⃣ Inference

Multimodal understanding (MMU) — image tokens + question → text:

python scripts/inference/run_mmu.py \
  --model /path/to/checkpoint \
  --image-tokens image_tokens.npy \
  --question "What is in this image?" \
  --token-h 32 --token-w 32

Text-to-image (T2I) — text → VQVAE tokens:

python scripts/inference/run_t2i.py \
  --model /path/to/checkpoint \
  --prompt "A cat sitting on a windowsill" \
  --output output_tokens.npy \
  --lat-h 32 --lat-w 32

Tip

T2I uses an empty system prompt by default, matching the training template. If your training data used a fixed non-empty system prompt, pass the same text via --system-prompt at inference.

📦 Data Format

Training input is a Parquet file (or a directory of Parquet files) with a unified schema:

Field I2T T2I I2I
task_type i2t t2i i2i
metadata_json JSON string JSON string JSON string
conversations JSON string / turn list — —
img_tokens raw VQVAE IDs — —
system_prompt optional string string
user_prompt — string string
input_image_tokens — — raw VQVAE IDs
answer_image_tokens — raw VQVAE IDs raw VQVAE IDs

📊 Benchmarks

🎨 Fine-Grained Text-to-Image Generation

Qualitative text-to-image comparisons

UniGenBench — Gestalt achieves the strongest overall result among the evaluated models, combining fine-grained semantic alignment with compositional generation.

UniGenBench results

TIIF-Bench — Gestalt leads across short and long instructions, with a larger advantage when longer prompts introduce more interdependent requirements.

TIIF-Bench results

👁️ Multimodal Understanding

The gains on vision-centric tasks reflect the value of preserving modality-specific information within a unified model.

Multimodal understanding results

📝 Language

Gestalt leads the evaluated diffusion-based unified models and achieves language performance comparable to autoregressive models such as Janus-Pro.

Language benchmark results

🔗 Multimodal Interplay

Visual and textual tokens are more interleaved in Gestalt, suggesting a more integrated representation space than the evaluated diffusion-based and autoregressive baselines.

Integrated vision-language representation space

Visual-specific and synergistic capability — strong visual-specific performance is paired with leading synergy results among the evaluated diffusion-based models.

Visual-specific and synergistic capability results

Interplay-aware representations — interplay tokens form distinct visual-unique, text-unique, and synergistic structures; the synergy distribution partially bridges the other two, suggesting that these tokens adapt to different information demands.

Interplay-token representations

✍️ Citation

If you find Gestalt useful for your research, please cite:

@article{gestalt2026,
  title   = {Gestalt: Large Multimodal Interplay Model},
  author  = {Yang, Zequn and Miao, Yu and Ni, Haotian and Chen, Ziheng and
             Huang, Chengxiang and Zhou, Dongzhan and Chen, Kai and Zhang, Qi and
             Wen, Ji-Rong and Wei, Yake and Hu, Di},
  year    = {2026},
  url     = {https://github.com/GeWu-Lab/Gestalt}
}

💬 Discussion Group (WeChat)

WeChat Group

📜 License

This project is released under the Apache 2.0 license.

👥 Author Contributions

Zequn Yang, Yake Wei, and Di Hu drove the overall advancement of the project. Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Chengxiang Huang contributed equally to this work. Zequn Yang conducted model pretraining and supervised fine-tuning. Zequn Yang, Yu Miao, and Haotian Ni developed the model architecture and conducted the core experiments. Yu Miao, Ziheng Chen, and Chengxiang Huang contributed to data processing and organization. Yu Miao conducted the evaluation of image generation capabilities, Haotian Ni conducted the multimodal understanding evaluation, and Ziheng Chen conducted the text-only evaluation. Dongzhan Zhou, Kai Chen, Qi Zhang, and Ji-Rong Wen contributed to discussions on the technical design and methodology. Yake Wei, Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Di Hu contributed to writing and revising the manuscript. Di Hu initiated the project. Yake Wei and Di Hu supervised and advised the project.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages