Instructions to use CRUISEResearchGroup/AnyMo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CRUISEResearchGroup/AnyMo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="CRUISEResearchGroup/AnyMo", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("CRUISEResearchGroup/AnyMo", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
NeurIPS 2026
[Paper] [Project Page] [Code] [AnyMo-Bench]
AnyMo is a geometry-aware framework for setup-agnostic wearable IMU motion understanding. It simulates IMU signals over dense body-surface placements, learns setup-stable full-body motion representations from sparse observations, discretizes them into compact IMU tokens, and aligns those tokens with a language model. The released model supports zero-shot activity recognition, IMU-text retrieval, and wearable motion captioning.
Model Components
This repository packages the paper model and its three motion components together:
CRUISEResearchGroup/AnyMo/
├── model.safetensors AnyMo motion-language model
├── anymo_encoder/ Geometry-aware ST-GCN encoder
├── anymo_tokenizer/ Product-quantized motion tokenizer
└── anymo_imu_codebook/ Standalone IMU codebook
The motion-language model uses Qwen2.5-0.5B as its language backbone. The motion tokenizer contains two 2,048-entry codebooks with 64-dimensional code vectors.
Using the Released Model
Set trust_remote_code=True to load the released raw-IMU pipeline.
import numpy as np
import torch
from transformers import pipeline
anymo = pipeline(
"anymo",
model="CRUISEResearchGroup/AnyMo",
trust_remote_code=True,
device=0,
dtype=torch.bfloat16,
)
# [time, sensors, channels] = [T, S, 6]
# Channel order: [acc_x, acc_y, acc_z, gyro_x, gyro_y, gyro_z]
imu = np.load("<PATH_TO_IMU_ARRAY>")
locations = ["Head", "L_Forearm", "R_Forearm"]
# Task 1: zero-shot human activity recognition
recognition = anymo.classify(
imu,
sensor_locations=locations,
sampling_rate=60,
candidate_labels=["walking", "sitting", "running"],
)
# Task 2: cross-modal IMU-to-text retrieval
candidate_texts = [
"A person is walking forward.",
"A person is sitting still.",
"A person is running.",
]
imu_embedding = anymo.encode_imu(imu, locations, sampling_rate=60)
text_embeddings = anymo.encode_text(candidate_texts)
similarities = imu_embedding @ text_embeddings.T
retrieved_text = candidate_texts[similarities[0].argmax().item()]
# Task 3: wearable IMU motion captioning
caption = anymo.caption(imu, locations, sampling_rate=60)
print(recognition)
print(retrieved_text)
print(caption)
Input Format
For one sample, imu must have shape [T, S, 6]. The sensor axis and sensor_locations are positionally aligned:
imu[:, 0, :] <-> sensor_locations[0]
imu[:, 1, :] <-> sensor_locations[1]
...
imu[:, S-1, :] <-> sensor_locations[S-1]
Sensor locations may be supplied in any order because the processor maps each name to its canonical node in the 23-node body graph. If the sensor axis is reordered, sensor_locations must be reordered in exactly the same way. Multiple sensors cannot map to the same graph node.
For batched input, use [B, T, S, 6] and provide either one shared list of S locations or a nested [B][S] list for samples with different wearable setups. All samples in a batch must have the same length after resampling.
- Accelerometer unit:
m/s² - Gyroscope unit:
rad/s - Released-model target rate: 60 Hz
- Recommended window: five seconds (
T=300at 60 Hz), matching the paper
The processor resamples the temporal axis to 60 Hz but does not otherwise crop or pad it. Canonical names follow the 23 body segments used in the paper. Common aliases such as left wrist, right wrist, waist, and chest are accepted.
Transformers Interface
The motion-language model and processor can also be loaded through standard Transformers auto classes:
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
processor = AutoProcessor.from_pretrained(
"CRUISEResearchGroup/AnyMo",
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
"CRUISEResearchGroup/AnyMo",
trust_remote_code=True,
dtype=torch.bfloat16,
)
Use the custom pipeline("anymo", ...) interface above for end-to-end inference from raw IMU because it additionally loads the released encoder, motion tokenizer, and standalone IMU codebook. The auto classes expose the motion-language model and processor separately for advanced use.
Intended Use and Limitations
AnyMo is intended for research on wearable motion representation learning, zero-shot activity recognition, motion-language retrieval, and motion captioning. The released checkpoint was trained with five-second windows and a 60 Hz target rate. Its body-location interface maps sensors to 23 anatomical segments and therefore does not represent multiple simultaneous sensors assigned to the same segment. Performance can decrease for motions that are poorly observable from the supplied sensors, context-dependent activity labels, substantially different hardware or mounting conditions, and domains far from the training distribution. It should not be used as the sole basis for safety-critical, clinical, or high-stakes decisions.
Citation
@article{chen2026anymo,
title = {AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild},
author = {Chen, Baiyu and Li, Zechen and Wongso, Wilson and Li, Lihuan and Lin, Xiachong and Xue, Hao and Tag, Benjamin and Salim, Flora},
journal = {arXiv preprint arXiv:2605.22715},
year = {2026}
}
License
The released model includes the Qwen2.5-0.5B backbone and is distributed under the Apache 2.0 license. AnyMo-specific source code in the GitHub repository is released under the MIT License. Dataset use remains subject to the licenses and access terms of the corresponding source datasets.
- Downloads last month
- 9
Model tree for CRUISEResearchGroup/AnyMo
Base model
Qwen/Qwen2.5-0.5B