kcode-speaker-v3-gemini-protocol 0.1.0

Deterministic Speaker V3 Gemini protocol construction and decoding
Documentation
Use the attached audio and transcript to estimate speaker features.

Use the target speaker's high-quality, non-overlapping speech. Estimate each requested feature from the audio. Use “insufficient evidence” when the available speech does not support a feature. Return clearly labelled values.

Transcript:
{{TRANSCRIPT}}

Requested features:

high_front_vowel_f2_hz: median F2 of the closest high-front /i/-like vowel, in Hz
low_vowel_f1_hz: median F1 of the closest low or open /a/-like vowel, in Hz
h1_minus_h2_db: average H1 minus H2 amplitude of modal voiced speech, in dB
rhotic_f3_minus_f2_hz: median F3 minus F2 for clear rhotic or r-colored speech, in Hz
word_initial_t_vot_ms: median voice-onset time for eligible stressed word-initial /t/, in ms
dominant_lateral_realization: dominant lateral-approximant realization
monophthongization_percent: percentage of eligible diphthongs substantially monophthongized
vocal_gender_presentation: strongly feminine, feminine, androgynous, masculine, or strongly masculine

Target speaker:
{{TARGET_SPEAKER}}