Track Any Anomalous Object: A Granular Video Anomaly Detection Framework (CVPR 2025)
Track Any Anomalous Object (TAO) introduces a Granular Video Anomaly Detection framework that, for the first time, unifies the detection and localization of multiple fine-grained anomalous objects within a single end-to-end pipeline.
Unlike conventional video anomaly detection methods that assign anomaly scores densely to every pixel at each time step, TAO reformulates anomaly detection as a pixel-level tracking problem of anomalous objects. By explicitly linking anomaly scores to downstream tasks such as image segmentation and video object tracking, our framework eliminates the need for heuristic threshold selection. This enables more accurate and robust anomaly localization, even in long and challenging video sequences.
Only the Ped2 subset is required for the default experimental setup.
- Dataset Reference: UCSD Anomaly Detection Dataset
Download the dataset:
cd datasets
wget http://www.svcl.ucsd.edu/projects/anomaly/UCSD_Anomaly_Dataset.tar.gzConvert video clips into frame sequences.
cd data
python extract_frames.py
Directory Structure:
path_videos = "./data/{dataset}/{training/testing}_videos/"
path_frames = "./data/{dataset}/{training/testing}/frames/"
Optical flow is extracted using FlowNet2.0. (Reference: FlowNet2 PyTorch)
(1) Install FlowNet2.0
cd pre_processing
bash install_flownet2.sh
cd ..(2) Download Pre-trained Weights
Download FlowNet2_checkpoint.pth.tar from here and place it in:
pre_processing/checkpoints/
(3) Extract Optical Flow
The extracted flows will be saved to ./data/ped2/{training/testing}/flows/.
- Test frames:
python flow.py --dataset_name=ped2
- Training frames:
python flow.py --dataset_name=ped2 --train
Bounding box annotations are stored as {dataset}_bboxes_train.npy and {dataset}_bboxes_test.npy.
Option 1: Use Provided Detections (Recommended) Download precomputed results from Google Drive. Place files as follows:
- Ped2 →
./data/ped2 - Avenue →
./data/avenue - ShanghaiTech →
./data/shanghaitech
Option 2: Generate Bounding Boxes Yourself
Note: The results are identical to Option 1.
- Install Detectron2:
python -m pip install 'git+https://github.com/facebookresearch/detectron2.git'
- Download ResNet50-FPN Weights:
wget https://dl.fbaipublicfiles.com/detectron2/COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x/137849600/model_final_f10217.pkl -P pre_processing/checkpoints/
- Run Extraction:
- Training Set (e.g., Ped2):
python pre_processing/bboxes.py --dataset_name=ped2 --train
# Output: ./data/ped2/ped2_bboxes_train.npy
- Test Set:
python pre_processing/bboxes.py --dataset_name=ped2
# Output: ./data/ped2/ped2_bboxes_test.npy
Compute the number of frames per video clip and the cumulative frame index.
cd data
python count_frames.py
Extract motion velocity and deep appearance features. Pose features (pose.npy) are already provided.
python feature_extraction.py --dataset_name={dataset}
Compute calibration parameters for each feature representation.
python score_calibration.py --dataset_name={dataset}
Run anomaly evaluation. Recommended sigma values:
- Ped2 / Avenue:
sigma = 3 - ShanghaiTech:
sigma = 7
python evaluate.py --dataset_name={dataset} --sigma={sigma}
python getbox_ped2.py
Please follow the instructions in the official SAM2 repository.
python robust_filtering.py
python sam2_inference.py
Configuration Paths:
- Frame data:
video_dirs = ./data/{dataset}/testing/frames - Anomalous regions:
pkl_dir = ./outputs/getbox_results/
Navigate to the benchmark folder:
cd Benchmark_Metrics/pixel_level
1. Process results after robust filtering: Computes AUROC, AP, AUPRO, and F1-score.
python vsresult_process.py
2. Evaluate SAM2 outputs:
python pixle_level_evaluate.py
- Extract detected anomaly trajectories:
python get_anomalies_path.py.py
- Extract ground-truth tracks:
Run the notebook:
get_tracks_path.ipynb - Compute TBDC and RBDC:
python compute_tbdc_rbdc.py \
--tracks-path=PATH_TO_GT_TRACKS \
--anomalies-path=PATH_TO_DETECTIONS \
--num-frames=NUM_FRAMES
Data Formats:
- Track:
track_id, frame_id, x_min, y_min, x_max, y_max - Detection:
frame_id, x_min, y_min, x_max, y_max, anomaly_score
We recommend using two separate environments to avoid dependency conflicts.
Used for sections 1 through 3.
- Python: 3.7
- PyTorch: 1.12.0+cu102
Used for Section 4 onwards.
- Follow the official SAM2 installation instructions.
If you find this work useful, please cite our paper:
@article{tao2025track,
title = {Track Any Anomalous Object: A Granular Video Anomaly Detection Framework},
author = {Author Names},
journal = {arXiv preprint arXiv:2506.05175},
year = {2025}
}