kcode-speaker-v3-llm-protocol 0.1.0

Deterministic Speaker V3 LLM protocol construction and decoding
Documentation
Use the attached audio and transcript to estimate speaker features.

Use the target speaker's high-quality, non-overlapping speech. Estimate each requested feature from the audio. Use “insufficient evidence” when the available speech does not support a feature. Return clearly labelled values.

Transcript:
{{TRANSCRIPT}}

Requested features:

low_vowel_f2_hz: median F2 of the closest low or open /a/-like vowel, in Hz
high_back_vowel_f1_hz: median F1 of the closest high-back /u/-like vowel, in Hz
mean_formant_dispersion_hz: mean adjacent F1–F4 formant spacing across clear modal vowels, in Hz
creaky_phonation_percent: percentage of voiced speech showing habitual creaky phonation
hypernasality_0_to_4: persistent hypernasality from 0 for none to 4 for severe
sibilant_center_of_gravity_hz: median spectral center of gravity of clear /s/-like sibilants, in Hz
consonant_cluster_reduction_percent: percentage of eligible consonant clusters substantially simplified
perceived_vocal_age_years: perceived vocal age based on the voice, in years

Target speaker:
{{TARGET_SPEAKER}}