kcode-speaker-extract 0.2.1

Deterministic speaker segment planning and strict normalized response parsing for Kennedy
Documentation
Analyze the attached audio directly and separate the substantive human speakers.

For each speaker whose usable speech supports a complete profile, provide a stable label beginning with Speaker 1 in first-appearance order, primary language, closest dialect or accent, estimated usable speech duration, and a score for every feature below. Briefly identify any additional substantive speakers whose available speech does not support a complete profile. Present the analysis in the format you find clearest.

Every feature is scored from 0 to 100, normalized over the general population. Judge each feature independently. Base the scores on observable speech in this recording. Interpret language-dependent features naturally within the speaker’s primary language.

1. filler_form_preference
0: Filled pauses and discourse holding are almost exclusively non-lexical sounds such as “uh” and “um.”
100: They are almost exclusively lexical fillers or discourse markers such as “like,” “you know,” and equivalents in the speaker’s primary language.

2. syntactic_complexity
0: Speech is constructed from short, simple, linear clauses.
100: Speech uses densely nested, structurally complex multi-clause constructions.

3. clause_completion_habit
0: Clauses frequently fade out, stop mid-thought, or remain unfinished.
100: Clauses consistently reach complete, definite closures.

4. declarative_terminal_rise
0: Declarative clauses almost always end with clearly falling pitch.
100: Declarative clauses almost always end with clearly high-rising pitch.

5. pragmatic_hedging
0: Expression relies on blunt, unmitigated assertions.
100: Expression is pervasively softened, qualified, or hedged.

6. speech_burst_contrast
0: Spoken runs have highly uniform rate and duration with evenly spaced pauses.
100: Speech repeatedly alternates fast, dense word bursts with salient pauses.

7. vocalized_hesitation_prominence
0: Audible fillers and stretched syllables are almost absent during speech planning.
100: Speech planning is pervasively carried by sustained fillers and vowel elongation.

8. pitch_expressiveness
0: Pitch contours show very little variation across phrases.
100: Pitch contours are highly varied, melodic, and animated.

9. lexical_formality
0: Word choice and register are highly casual, colloquial, or slang-heavy.
100: Word choice and register are highly formal, academic, or ceremonious.

10. vocal_gender_presentation
0: Strongly feminine-sounding vocal presentation.
100: Strongly masculine-sounding vocal presentation.

11. perceived_vocal_age_percentile
0: Youngest-sounding vocal extreme in the general population.
100: Oldest-sounding vocal extreme in the general population.

12. pitch_level_percentile
0: Lowest habitual speaking-pitch extreme in the general population.
100: Highest habitual speaking-pitch extreme in the general population.

13. vocal_weight
0: Very light, thin, or reedy vocal substance.
100: Very heavy, thick, or booming vocal substance.

14. accent_markedness
0: Pronunciation is acoustically close to a broadly understood reference variety of the primary language.
100: Pronunciation is extremely regionally or non-natively marked relative to that reference variety.

15. resonance_brightness
0: Very dark, low-frequency-dominant, chest-oriented resonance.
100: Very bright, high-frequency-forward, head-oriented resonance.

16. articulatory_precision
0: Speech is heavily mumbled, slurred, or reduced.
100: Speech is exceptionally crisp, precise, and hyper-enunciated.

17. rhoticity_level
0: R realizations are consistently absent, deleted, or vocalized.
100: R realizations are consistently prominent, hard, or strongly trilled.

18. loudness_dynamic_range
0: Prominent and reduced speech remain at nearly the same loudness.
100: Speech uses exceptionally strong loudness contrast between prominent and reduced material.

19. articulation_tempo
0: Articulated syllables within spoken runs move at an extremely slow rate.
100: Articulated syllables within spoken runs move at an extremely fast rate.

20. vocal_fry
0: Phonation remains consistently clean and modal.
100: Low-register creak or vocal fry is pervasive across voiced speech.

21. modal_breathiness
0: Ordinary voiced speech has firm closure and very little audible air leakage.
100: Ordinary voiced speech has pervasive airy leakage and strongly breathy phonation.

22. modal_roughness
0: Ordinary modal-register voicing has an exceptionally smooth, regular texture.
100: Ordinary modal-register voicing has a pervasive irregular, gravelly texture.

23. vocal_attack
0: Speech onsets usually emerge gradually and softly.
100: Speech onsets usually begin with abrupt, hard glottal closure.

24. sibilant_sharpness
0: Sibilants are extremely dull, retracted, diffuse, or lisped.
100: Sibilants are extremely piercing, dental, focused, or whistling.