Use the attached audio and transcript to estimate speaker features.
Use the target speaker's high-quality, non-overlapping speech. Estimate each requested feature from the audio. Use “insufficient evidence” when the available speech does not support a feature. Return clearly labelled values.
Transcript:
{{TRANSCRIPT}}
Requested features:
median_f0_hz: median fundamental frequency of ordinary modal voiced speech, in Hz
high_front_vowel_f1_hz: median F1 of the closest high-front /i/-like vowel, in Hz
high_back_vowel_f2_hz: median F2 of the closest high-back /u/-like vowel, in Hz
spectral_tilt_db_per_octave: average spectral tilt of modal voiced speech, in dB per octave
cepstral_peak_prominence_db: average cepstral peak prominence of connected modal speech, in dB
foreign_accentedness_1_to_9: perceived foreign accentedness relative to a broadly understood native variety, from 1 to 9
dominant_rhotic_realization: dominant rhotic realization
unstressed_vowel_reduction_percent: percentage of eligible unstressed vowels habitually reduced or centralized
Target speaker:
{{TARGET_SPEAKER}}