AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

NeurIPS 2026

[Paper] [Project Page] [Code] [AnyMo-Bench]

AnyMo is a geometry-aware framework for setup-agnostic wearable IMU motion understanding. It simulates IMU signals over dense body-surface placements, learns setup-stable full-body motion representations from sparse observations, discretizes them into compact IMU tokens, and aligns those tokens with a language model. The released model supports zero-shot activity recognition, IMU-text retrieval, and wearable motion captioning.

Model Components

This repository packages the paper model and its three motion components together:

CRUISEResearchGroup/AnyMo/
├── model.safetensors              AnyMo motion-language model
├── anymo_encoder/                Geometry-aware ST-GCN encoder
├── anymo_tokenizer/              Product-quantized motion tokenizer
└── anymo_imu_codebook/            Standalone IMU codebook

The motion-language model uses Qwen2.5-0.5B as its language backbone. The motion tokenizer contains two 2,048-entry codebooks with 64-dimensional code vectors.

Using the Released Model

Set trust_remote_code=True to load the released raw-IMU pipeline.

import numpy as np
import torch
from transformers import pipeline

anymo = pipeline(
    "anymo",
    model="CRUISEResearchGroup/AnyMo",
    trust_remote_code=True,
    device=0,
    dtype=torch.bfloat16,
)

# [time, sensors, channels] = [T, S, 6]
# Channel order: [acc_x, acc_y, acc_z, gyro_x, gyro_y, gyro_z]
imu = np.load("<PATH_TO_IMU_ARRAY>")
locations = ["Head", "L_Forearm", "R_Forearm"]

# Task 1: zero-shot human activity recognition
recognition = anymo.classify(
    imu,
    sensor_locations=locations,
    sampling_rate=60,
    candidate_labels=["walking", "sitting", "running"],
)

# Task 2: cross-modal IMU-to-text retrieval
candidate_texts = [
    "A person is walking forward.",
    "A person is sitting still.",
    "A person is running.",
]
imu_embedding = anymo.encode_imu(imu, locations, sampling_rate=60)
text_embeddings = anymo.encode_text(candidate_texts)
similarities = imu_embedding @ text_embeddings.T
retrieved_text = candidate_texts[similarities[0].argmax().item()]

# Task 3: wearable IMU motion captioning
caption = anymo.caption(imu, locations, sampling_rate=60)

print(recognition)
print(retrieved_text)
print(caption)

Input Format

For one sample, imu must have shape [T, S, 6]. The sensor axis and sensor_locations are positionally aligned:

imu[:, 0, :]   <-> sensor_locations[0]
imu[:, 1, :]   <-> sensor_locations[1]
...
imu[:, S-1, :] <-> sensor_locations[S-1]

Sensor locations may be supplied in any order because the processor maps each name to its canonical node in the 23-node body graph. If the sensor axis is reordered, sensor_locations must be reordered in exactly the same way. Multiple sensors cannot map to the same graph node.

For batched input, use [B, T, S, 6] and provide either one shared list of S locations or a nested [B][S] list for samples with different wearable setups. All samples in a batch must have the same length after resampling.

  • Accelerometer unit: m/s²
  • Gyroscope unit: rad/s
  • Released-model target rate: 60 Hz
  • Recommended window: five seconds (T=300 at 60 Hz), matching the paper

The processor resamples the temporal axis to 60 Hz but does not otherwise crop or pad it. Canonical names follow the 23 body segments used in the paper. Common aliases such as left wrist, right wrist, waist, and chest are accepted.

Transformers Interface

The motion-language model and processor can also be loaded through standard Transformers auto classes:

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

processor = AutoProcessor.from_pretrained(
    "CRUISEResearchGroup/AnyMo",
    trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
    "CRUISEResearchGroup/AnyMo",
    trust_remote_code=True,
    dtype=torch.bfloat16,
)

Use the custom pipeline("anymo", ...) interface above for end-to-end inference from raw IMU because it additionally loads the released encoder, motion tokenizer, and standalone IMU codebook. The auto classes expose the motion-language model and processor separately for advanced use.

Intended Use and Limitations

AnyMo is intended for research on wearable motion representation learning, zero-shot activity recognition, motion-language retrieval, and motion captioning. The released checkpoint was trained with five-second windows and a 60 Hz target rate. Its body-location interface maps sensors to 23 anatomical segments and therefore does not represent multiple simultaneous sensors assigned to the same segment. Performance can decrease for motions that are poorly observable from the supplied sensors, context-dependent activity labels, substantially different hardware or mounting conditions, and domains far from the training distribution. It should not be used as the sole basis for safety-critical, clinical, or high-stakes decisions.

Citation

@article{chen2026anymo,
  title   = {AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild},
  author  = {Chen, Baiyu and Li, Zechen and Wongso, Wilson and Li, Lihuan and Lin, Xiachong and Xue, Hao and Tag, Benjamin and Salim, Flora},
  journal = {arXiv preprint arXiv:2605.22715},
  year    = {2026}
}

License

The released model includes the Qwen2.5-0.5B backbone and is distributed under the Apache 2.0 license. AnyMo-specific source code in the GitHub repository is released under the MIT License. Dataset use remains subject to the licenses and access terms of the corresponding source datasets.

Downloads last month
9
Safetensors
Model size
0.5B params
Tensor type
I64
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CRUISEResearchGroup/AnyMo

Finetuned
(728)
this model

Dataset used to train CRUISEResearchGroup/AnyMo

Collection including CRUISEResearchGroup/AnyMo

Paper for CRUISEResearchGroup/AnyMo