SMAD CRNN

SMAD is an audio classification model for identifying speech, music, singing, and non-vocal content in short audio segments. It is designed for lightweight audio analysis workflows where fast, practical content categorization is needed.

Load by model id

import torch
from transformers import AutoFeatureExtractor, AutoModelForAudioClassification

model_id = "duclvQ/smad"

feature_extractor = AutoFeatureExtractor.from_pretrained(
    model_id,
    trust_remote_code=True,
)
model = AutoModelForAudioClassification.from_pretrained(
    model_id,
    trust_remote_code=True,
).eval()

# `audio` must be mono float32 audio sampled at 16 kHz. For longer files, run
# this over 4-second windows.
inputs = feature_extractor(audio, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits / model.config.temperature, dim=-1)

label_id = int(probs.argmax(-1)[0])
label = model.config.id2label[label_id]
confidence = float(probs[0, label_id])

For arbitrary file paths, load/resample with librosa:

import librosa

audio, _ = librosa.load("clip.mp3", sr=16000, mono=True)
Downloads last month
-
Safetensors
Model size
835k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support