[📄 Paper] [🌐 Website] [🤗 Hugging Face]
The Official Implementation of "ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow"
- Clone this repository and create the environment.
git clone https://github.com/Dstate/ODEWorld.git
cd ODEWorld
conda create -n odeworld python=3.10 -y
conda activate odeworld- Install the remaining dependencies.
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt- Download the pretrained checkpoints into
assets/pretrained.
mkdir -p assets/pretrained
for model in \
ODEWorld-PT-Flow-LIBERO \
ODEWorld-PT-Flow-AgiBot \
ODEWorld-Goal-Predictor-LIBERO \
ODEWorld-RAE-LIBERO \
ODEWorld-RAE-AgiBot
do
hf download "ldxxx/${model}" --local-dir "assets/pretrained/${model}"
doneExample images and manifests are stored in assets/examples/<dataset>/.
Run the LIBERO examples:
python demo_infer.py --dataset liberoRun the AgiBot examples:
python demo_infer.py --dataset agibotUse --case-ids case_00 to run a single case. Results are written to outputs/<dataset>/<case_id>.
Run the following commands from the repository root to link your prepared HDF5 datasets into assets/data. Replace the source paths with your dataset locations:
mkdir -p assets/data
ln -s /path/to/libero assets/data/libero
ln -s /path/to/agibot_hdf5 assets/data/agibot_hdf5The linked directories must match the trajectory paths in assets/metas/*.json:
assets/data/
├── libero/
│ └── libero_90/<task>/demo_0.hdf5
└── agibot_hdf5/
└── demo_0.hdf5
The loaders expect per-trajectory HDF5 files with encoded image frames and the fields specified by the metadata. Use datasets prepared in this format. Each .hdf5 file contains one trajectory of T frames. The current dataset layouts are:
# LIBERO: demo_*.hdf5
observation/
third_image # (T,), variable-length uint8 arrays of JPEG bytes
wrist_image # (T,), variable-length uint8 arrays of JPEG bytes
language_instruction # scalar UTF-8 string describing the trajectory
action # (T, 7), float32
proprio # (T, 9), float32
# AgiBot: demo_*.hdf5
observation/
head_image # (T,), variable-length uint8 arrays of JPEG bytes
language_instruction # scalar UTF-8 string describing the trajectory
action # (T, 16), float32
proprio # (T, 16), float32
Image datasets have HDF5 dtype h5py.vlen_dtype(np.dtype("uint8")). Each element is a complete encoded image, decoded by cv2.imdecode(frame, cv2.IMREAD_COLOR).
Run from the repository root and choose the commands for your dataset.
-
Train the RAE image decoder.
bash scripts/train_dinov2rae_libero.sh bash scripts/train_dinov2rae_agibot.sh
-
Train the PT-Flow dynamics model.
bash scripts/train_dinov2ptflow_libero.sh bash scripts/train_dinov2ptflow_agibot.sh
-
Train the goal predictor for language-conditioned LIBERO inference.
bash scripts/train_dinov2goalpred_libero.sh
@article{liu-niu2026odeworld,
title={ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow},
author={Liu, Dongxiu and Niu, Haoyi and Cheng, Peng and Gao, Yuan and Kang, Xirui and Teng, Sangli and Sreenath, Koushil and Zhan, Xianyuan},
journal={Advances in Neural Information Processing Systems},
year={2026}
}