Parakeet TDT 0.6B V2 — Core AI

Apple Core AI conversion of NVIDIA Parakeet V2, pinned to ae9ad07059c7c739ffaf932226a8fe64ae2620b0. Converted weights retain the upstream Creative Commons Attribution 4.0 license. See LICENSE and NOTICE. This conversion is not endorsed by NVIDIA.

Choose one self-contained bundle:

  • fast/: full finite Fourier relative attention, 192 encoder positions, fixed 15-second chunks, and an FP16 ANE decoder batching up to 128 independent chunks.
  • quality/: full sinusoidal relative attention, shared weights for 576 and 3,008 encoder positions (45-second and four-minute inputs), exact tiled convolution subsampling, fixed 240-second chunks, and an FP32 CPU decoder batching up to four independent chunks. Select the smallest shape that fits each chunk; short recordings use the 45-second shape.

Both use W8A16 encoder weights with selected sensitive projections retained in FP16, and four concurrent encoder requests. Neither truncates attention within its chunk or drops input audio. metadata.json supplies available input shapes; runtime.json records the qualified scheduling policy. Weights are shared across static functions in each asset. The labels describe a measured speed/context tradeoff, not a universal quality ranking.

Requires physical Apple silicon on macOS 27 or iOS 27 and a model-specific host runtime. Graphs accept model features and recurrent states, not audio files. The host must implement the frontend, greedy TDT loop and tokenizer described by the metadata and sidecars. Vocabulary: 1,024 nonblank tokens, blank ID 1,024, durations 0/1/2/3/4, and at most ten symbols per frame. Recurrent state is independent between chunks. TDT emission frames and predicted durations support token/word timings at an 80 ms frame step; they are native alignments, not forced alignment.

Measured performance

M3 MacBook Air (16 GB), macOS 27 build 26A428. Medians after warmup; frontend, encoder and decoding are included. Preparation, file I/O and chunk planning are excluded. Audio is JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the 18-minute 15-second recording.

Bundle Audio duration Transcription Audio / elapsed
fast 20 seconds 0.150 s 133.4×
fast 18 min 15 s 1.876 s 584.0×
quality 20 seconds 0.153 s 130.9×
quality 18 min 15 s 6.257 s 175.0×

Quality measurements use three timed runs after warmup, with stable tokens and native word timings and nominal thermal state. The full recording uses 74 fast chunks or five quality chunks. These contexts and execution policies differ.

Whisper-normalized WER against the supplied long-recording reference:

Bundle / reference Word errors WER
fast 79/2220 3.56%
quality 45/2220 2.03%
FP32 source, same four-minute cuts as quality 46/2220 2.07%

With the benchmark's simpler normalization, quality scores 54/2219 (2.43%) and the matching source 53/2219 (2.39%). This is one English recording, not a general accuracy ranking. Quantization and floating-point operation order can change token decisions. Both quality shapes pass the short-fixture projected encoder check at 2.04% relative RMS versus the independent FP32 source and produce identical valid outputs to each other. Word timings are ordered, bounded and text-preserving; human word-boundary accuracy was not measured.

A bounded full-recording quality trace contains ANE predictions within all ten encoder/subsampling calls and zero target GPU intervals. The decoder uses the CPU. Compiler manifests mark both functions in each quality asset fully placed on ANE; this is placement evidence, not an arithmetic-utilization measurement. Sampled peak client-plus-attributed-neural memory is approximately 1.07 GB, excluding unattributed compiler, driver and system memory.

Measured iPhone performance

iPhone 15 Pro Max, iOS 27 build 24A437, Release runtime. Three-run medians after warmup, using the same Moon speech. Includes bounded WAV reading/conversion, chunk planning, frontend, encoder, decoding and output delivery. Preparation, warmup and report writing are excluded. Hardware traces were captured separately.

Bundle Audio Transcription Audio / elapsed
fast 20 seconds 0.166 s 120.7×
fast 18 min 15 s 2.418 s 453.0×
quality 20 seconds 0.171 s 117.3×
quality 18 min 15 s 7.189 s 152.4×
Bundle Observed preparation with specialization Cached preparation
fast 47.7 s 0.115 s
quality encoder 446.9 s; subsampling 41.4 s (separate loads) 0.063 s

Preparation includes loading and any specialization requested by Core AI. Underlying driver cache reuse is opaque; these are observed histories, not controlled fresh-install measurements or evidence of an empty system cache. The quality inference runs follow a cooldown after compilation. All reported passes remained foregrounded; nominal and fair thermal states were observed.

Fast's complete recording trace contains 240 unique ANE predictions across 74 encoder calls, 74 subsampling calls and the batched decoder loop, with zero target-process GPU intervals. Fast token IDs and transcript text match the qualified Mac outputs for both input lengths.

Quality's encoder and subsampling calls contain ANE predictions for both shapes, with zero target-process GPU intervals. Quality decoding uses the CPU.

Quality's short output matches the Mac. Full-recording tokens differ but remain stable across phone runs: 46/2,220 errors (2.07% Whisper-normalized WER), versus 45/2,220 (2.03%) on the Mac and 46/2,220 for the matching FP32 source. With the benchmark's simpler normalization, the phone scores 53/2,219 and the Mac 54/2,219. These small device differences do not establish a general quality ranking.

Placement traces establish execution, not ALU utilization or energy use. Concurrent scopes can overlap the same hardware event; aggregate counts use unique intervals.

Device specialization and loading

Source .aimodel files specialize on the device. On the same Mac:

Bundle First observed preparation with specialization Subsequent cached preparation
fast 33 s 0.060 s
quality 6 min 2 s 0.041 s

Preparation includes Core AI model initialization and function loading; it excludes host sidecar reads, audio I/O, chunk planning, warmup and transcription. These are observed cache histories, not guaranteed fresh-install times. The underlying ANE cache state is not fully observable, and OS/device/application changes can require specialization again.

AoT compiled artifacts are omitted because no significant load-time benefit has been demonstrated. Authoring debug locations were removed while preserving graph signatures and operation counts. SHA256.json lists distributed payload checksums.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/parakeet-v2-coreai

Finetuned
(46)
this model

Collection including coder543/parakeet-v2-coreai