kcode-speaker-extract 0.1.1

Deterministic speaker extraction planning and strict response parsing for Kennedy
Documentation
Analyze the attached audio directly. Separate substantive human speakers without guessing identity. A substantive speaker has enough intelligible speech to support every rating below; ignore synthetic/GPS voices, laughter-only voices, and incidental fragments that are too short to rate.

Return exactly one JSON object and no Markdown or commentary. Use one of these two shapes.

Scored:
{"status":"scored","recordingQuality":82,"speakers":[{"speakerOrdinal":0,"primaryLanguage":"eng","closestDialect":"General American English","usableSpeechMs":43120,"features":[35 integers]}]}

Unscorable:
{"status":"unscorable","reason":"Short concrete reason."}

Use status scored only when the recording contains at least one substantive human speaker and you can give all 35 values for every substantive speaker. Otherwise return whole-recording status unscorable. Never return null, N/A, unknown, a range, a float, a stringified number, a partial row, per-field confidence, or per-field abstention. Do not mix scored speakers with unscorable status.

For scored output:
- Assign speakerOrdinal values uniquely and contiguously from 0 in first-appearance order.
- primaryLanguage is the lowercase ISO 639-3 code for the language containing the largest share of that speaker's speech.
- closestDialect is one concise nonempty best-guess language/dialect/accent label, without identity claims.
- usableSpeechMs is a positive integer estimate of usable speech from that speaker and cannot exceed clip duration. Speakers may overlap; do not force their durations to sum to clip duration.
- recordingQuality is one integer 0-100 for acoustic suitability of the whole clip: 0 unusable, 50 workable with material noise/compression/distance, 100 exceptionally clean close speech.
- features is exactly 35 integers in the fixed order below.

Except perceived_age, every feature uses 0-100. Use 0 and 100 only for rare extremes and 50 for the comparison-population midpoint, not for uncertainty. Interpret language-dependent behavior in the speaker's primary language, including natural regional equivalents. Rate audible behavior in this clip, not biography, anatomy, personality, intent, health, or immutable identity. If the evidence is not sufficient to rate every field for every substantive speaker, abstain on the whole recording.

Fixed feature order and anchors:
0 filler_preference — 0 almost exclusively non-lexical filled pauses such as uh/um; 50 balanced non-lexical and lexical fillers; 100 almost exclusively lexical fillers or discourse markers such as like/you know and primary-language equivalents. Count fillers, not planning silence.
1 syntactic_complexity — 0 short simple linear clauses; 50 ordinary mixed coordination/subordination; 100 densely nested multi-clause constructions.
2 phrase_ending_habit — 0 frequently trails off, leaves clauses unfinished, or signals weak completion; 50 mixed/neutral completion; 100 consistently crisp, hard, definitive closure. Judge completion independently of terminal pitch direction.
3 uptalk_hrt — 0 strongly falling terminal pitch; 50 level or mixed terminals; 100 strongly high-rising terminals on declaratives.
4 pragmatic_hedging — 0 blunt unmitigated declaratives; 50 ordinary contextual mitigation; 100 heavily hedged, softened, or qualified phrasing.
5 burstiness — 0 evenly metronomic pacing; 50 moderate variation; 100 rapid speech bursts separated by salient pauses.
6 verbosity — 0 telegraphic minimum wording; 50 ordinary elaboration; 100 highly circumlocutory or rambling expression for the conveyed content.
7 hesitation_strategy — 0 planning expressed mainly through silence; 50 mixed silence and vocalized delay; 100 planning expressed mainly through aggressive vowel elongation or drawn-out fillers.
8 pitch_expressiveness — 0 monotone pitch contour; 50 ordinary conversational melody; 100 highly melodic or animated pitch movement.
9 self_amused_phonation — 0 consistently dry/serious delivery; 50 occasional audible smile; 100 pervasive speech-through-smile, suppressed laughter, or chuckling.
10 vocal_gender_presentation — 0 strongly feminine-sounding presentation; 50 androgynous presentation; 100 strongly masculine-sounding presentation. This is an acoustic presentation estimate, not gender identity.
11 perceived_age — one integer 0-120 representing the single best guessed vocal age in years; this is an estimate, not biography.
12 vocal_weight — 0 very thin, reedy, or light voice; 50 medium weight; 100 very heavy, booming, or thick voice.
13 accentedness — 0 acoustically close to a broadly understood standard/reference variety of the primary language; 50 clearly regionally or foreign marked; 100 extremely marked relative to that reference. Do not infer origin or nativeness from the number.
14 resonance_placement — 0 very dark/chest-dominant resonance; 50 balanced resonance; 100 very bright, head-dominant, or nasal-forward resonance.
15 baseline_energy — 0 lethargic or very low-energy delivery; 50 ordinary alertness; 100 intensely energetic or chipper delivery.
16 relative_pitch_level — 0 unusually low, 50 typical, 100 unusually high relative to adult speakers with comparable vocal gender presentation speaking the primary language.
17 articulatory_precision — 0 heavily slurred, mumbled, or reduced articulation; 50 ordinary conversational precision; 100 hyper-enunciated, exceptionally crisp articulation.
18 rhoticity_level — 0 absent/deleted/vocalized r realizations; 50 moderate or mixed rhotic realization; 100 consistently hard, prominent, or strongly trilled r realizations, interpreted within the primary language.
19 formality — 0 highly casual/slang-heavy register; 50 ordinary conversational register; 100 highly academic, stiff, or formal register.
20 dynamic_range — 0 nearly flat loudness; 50 ordinary stress-linked loudness contrast; 100 very strong stressed/unstressed volume contrast.
21 tempo — 0 extremely slow baseline speech rate; 50 typical conversational rate; 100 extremely fast baseline speech rate. Exclude silent pauses when judging articulation tempo.
22 vocal_fry — 0 clean modal phonation with no audible fry; 50 frequent fry; 100 pervasive creak/fry across voiced speech.
23 hypernasality — 0 no audible hypernasality; 50 clearly moderate hypernasality; 100 extremely severe nasal resonance.
24 breathiness — 0 no breathiness or strongly pressed voice; 50 clearly breathy; 100 severely breathy phonation.
25 roughness — 0 smooth/clear voice; 50 clearly rough/gravelly; 100 extremely rough or heavily gravelled voice.
26 vocal_attack — 0 consistently soft/breathy onset; 50 mixed or neutral onset; 100 consistently hard glottal attack.
27 pitch_register_stability — 0 frequent cracks, breaks, or squeaks; 50 moderate stability; 100 exceptionally stable register.
28 emphatic_stress_strategy — 0 emphasis is predominantly volume-led; 50 balanced volume and pitch; 100 emphasis is predominantly pitch-led.
29 inspiratory_prominence — 0 silent or nearly inaudible pre-speech breathing; 50 regularly audible inhalation; 100 very loud, frequent pre-speech inhalation.
30 jaw_articulation — 0 audibly restricted/clenched articulation; 50 ordinary opening; 100 audibly wide/open articulation. Rate the sound only, not physical anatomy.
31 unstressed_vowel_centralization — 0 unstressed vowels remain unusually pure/full; 50 ordinary language-relative reduction; 100 pervasive heavy centralization toward schwa or the primary-language equivalent.
32 sibilant_sharpness — 0 very dull, lisped, or retracted sibilants; 50 ordinary sibilants; 100 extremely piercing, dental, or sharp sibilants.
33 pre_speech_articulatory_noise — 0 consistently clean speech starts; 50 regularly audible lip smacks/wet clicks; 100 very frequent, prominent pre-speech articulatory noises.
34 aspiration_intensity — 0 very weak or unaspirated stop releases; 50 ordinary language-relative aspiration; 100 extremely strong aspiration on relevant plosives.

Base every value only on observable evidence in this clip. Return the JSON object only.