Prosodia
⚠️ Why not a model
Audio modality specialist — not a text-based System One peer; different input domain entirely.
Technical specs
- Base LLMs
- whisper-encoder
- Decision types
- choice, score, noul
- Features
- audio-native, no-asr, whisper-encoder
- License
- Research
Links
Prosodia (audio-jevlike) is an audio-native specialist that provides Jev-shaped decisions (choice/score/noul) directly from audio input, without running ASR text decode.
How it works
Instead of:
- Audio → Whisper transcription → Text → Decision model
Prosodia does:
- Audio → Whisper encoder → Decision head → Typed decision
The Whisper decoder never runs — decisions come directly from the encoder representations.
Why it’s listed here (not in Models)
Prosodia is a modality specialist:
- Different input domain (audio, not text)
- Not comparable to text-based System One models
- Research prototype status
- Would distort the quadrant if included alongside text models
Research tone
The project has a research framing:
- IIA comparisons vs Jev are UnverifiedClaim
- Phase 1 ablations are author-reported
- Early-stage prototype
