Post
1437
Most deepfake audio detectors are quietly cheating.
They don’t really listen to the speech — they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.
AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.
The result is a detector that actually has to learn the artifacts instead of gaming the feature space.
Model: Modotte/AIRealNet-Audio
They don’t really listen to the speech — they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.
AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.
The result is a detector that actually has to learn the artifacts instead of gaming the feature space.
Model: Modotte/AIRealNet-Audio