Skip to main content

Crate normalizer_tr

Crate normalizer_tr 

Source
Expand description

A local, synchronous Turkish text-to-speech normalizer.

The default preserves unresolved spans and reports them. Use AmbiguityPolicy::Reject when partial speech is not acceptable. Original source ranges are UTF-8 byte coordinates, not character indices.

use normalizer_tr::{Normalizer, NormalizeOptions};
let normalizer = Normalizer::new()?;
let result = normalizer.normalize("25 TL", &NormalizeOptions::default())?;
assert_eq!(result.normalized_text(), "yirmi beş Türk lirası");
assert!(result.complete());

§normalizer-tr

CI

Turkish text normalization for text-to-speech, with a Rust core and a typed Python binding. Turn written numbers, measurements, money and other supported expressions into spoken Turkish without rewriting ordinary prose.

Input:  saat 09:30'da 5 kg malzeme ve %12,5'lik fark
Output: saat dokuz otuzda beş kilogram malzeme ve yüzde on iki virgül beşlik fark

The same Rust engine powers both APIs. It is synchronous, works offline and requires no model, network service, Torch or async runtime. Exact decimal and money arithmetic uses integers, not floating point.

Status: early-development 0.3.0, not a stable 1.0 API or a universal pronunciation guarantee. The Rust crate and Python distribution share the name normalizer-tr; both use normalizer_tr in code. The internal Rust/Python companion is not a separate crates.io product.

§Choose your API

Use caseComponentDependencies
Rust applicationsnormalizer-tr / import normalizer_trRust only; optional Serde
Python applicationsnormalizer-tr / import normalizer_trThe compiled Rust binding; no speech model

bindings/python is a separate Cargo workspace member and wheel, not part of the core .crate. The binding depends on the Rust core, never the reverse.

§Rust quickstart

Use Rust 1.94 or newer / edition 2024. Development builds use Rust 1.99.0. From your application’s Cargo.toml:

[dependencies]
normalizer-tr = "0.3"

Commit your application’s Cargo.lock to pin resolved dependencies. For an unpublished checkout, use a reviewed Git revision or a local path such as normalizer-tr = { path = '..\normalizer-tr' }.

use normalizer_tr::{Normalizer, NormalizeOptions};

let normalizer = Normalizer::new()?;
let result = normalizer.normalize(
    "saat 09:30'da 5 kg malzeme ve %12,5'lik fark",
    &NormalizeOptions::default(),
)?;

assert_eq!(
    result.normalized_text(),
    "saat dokuz otuzda beş kilogram malzeme ve yüzde on iki virgül beşlik fark",
);
assert!(result.complete());

Keep and reuse a Normalizer: its compiled resources are immutable and shared. It is cloneable and Send + Sync; inputs and results are not cached. Enable the optional serde feature to serialize results, segments and issues.

§Python quickstart

The Python bridge calls Rust directly rather than duplicating language rules. The release targets ordinary CPython 3.11–3.14 on Windows/Linux x64 and macOS x64/arm64. Linux wheels require glibc 2.28 or newer; macOS wheels target 12.0 or newer. No PyPy, free-threaded Python or other architectures are claimed.

python -m pip install normalizer-tr

Compatible wheels require no Rust compiler. To build from source instead:

Install Rust and the MSVC C++ build tools, then run in PowerShell. On Windows, use a short checkout path (for example C:\src\normalizer-tr) to avoid path-length limits during installation.

git clone https://github.com/erdemtuna/normalizer-tr.git
Set-Location normalizer-tr
py -3.13 -m venv .venv
.\.venv\Scripts\python.exe -m pip install .\bindings\python
.\.venv\Scripts\python.exe -X utf8 .\bindings\python\examples\showcase.py

The source installation builds a native wheel and requires Rust/C++ build tools. For a local wheel, use python -m pip install --no-deps <wheel-path>.

from normalizer_tr import Normalizer

normalizer = Normalizer()
result = normalizer.normalize("25 TL; 5 kg")
print(result.normalized_text)  # yirmi beş Türk lirası; beş kilogram
assert result.complete
assert result.issues == ()

-X utf8 avoids Turkish stdout encoding errors on Windows shells configured with a legacy encoding. It does not alter normalization. See the Python API and build guide for hints, cancellation, exceptions and installed-wheel testing.

§Ambiguity is explicit

By default, supported spans normalize and unresolved spans remain exactly as written. Check complete before passing a result to a speech model:

use normalizer_tr::{Normalizer, NormalizeOptions};

let result = Normalizer::new()?.normalize("25 TL; 1.234", &NormalizeOptions::default())?;
assert_eq!(result.normalized_text(), "yirmi beş Türk lirası; 1.234");
assert!(!result.complete());
assert_eq!(result.issues().len(), 1);

Bare 1.234 has insufficient reading intent. Use a whole-span hint if your application knows how to interpret it, or select strict rejection:

from normalizer_tr import Hint, Normalizer, NormalizationError

normalizer = Normalizer()
assert normalizer.normalize(
    "00042", hints=(Hint(0, 5, "digits"),)
).normalized_text == "sıfır sıfır sıfır dört iki"

try:
    normalizer.normalize("1.234", ambiguity_policy="reject")
except NormalizationError as error:
    assert error.code == "unresolved"
    assert error.issues  # No partial result is returned in strict mode.

All ranges are half-open original UTF-8 byte offsets, not character positions, Python indices or UTF-16 offsets. Hints must cover a whole expression, be grapheme-safe and not overlap. For a prefix in Python, compute its byte length with len(prefix.encode("utf-8")).

Invalid input/hints, cancellation, deadlines, resource limits and internal failures are errors in either policy, not successful preservation.

§Supported expressions

WrittenSpoken
12,05on iki virgül sıfır beş
%3,25'tenyüzde üç virgül iki beşten
25 TL'denyirmi beş Türk lirasından
€14,05 / $40 / 25 GBPon dört avro beş sent / kırk dolar / yirmi beş sterlin
1.'nin / 4.'yebirincinin / dördüncüye
5 kg'dan / 2 sa'ten / 5 m³'ebeş kilogramdan / iki saatten / beş metrekübe
90 km/sa / 5 m/ssaatte doksan kilometre / saniyede beş metre
10-15 kişion ila on beş kişi
tarih 01.02.2026tarih bir Şubat iki bin yirmi altı
Prof. / TBMM / KDVprofesör / te be me me / katma değer vergisi
0850 222 33 44sıfır sekiz yüz elli iki yüz yirmi iki otuz üç kırk dört
II. Dünya Savaşıikinci Dünya Savaşı
info@ornek.cominfo et ornek nokta kom

Coverage is deliberately bounded, not a universal Turkish pronunciation engine. Normalization reference documents the exact grammars, aliases, hints, boundaries and unsupported cases. Unknown or malformed identifiers/composites are protected as whole spans, not rewritten in fragments. Ordinary word case is retained; the library does not globally lowercase or adapt output to a particular voice/model vocabulary.

§Develop and contribute

cargo test --locked
cargo run --locked --example normalize
cargo run --locked --example general
cargo doc --locked --no-deps --open

Default Cargo operations build the standalone core. --workspace also includes the Python binding. CI exercises the Rust configurations and the actual installed Python wheel without speech models.

See CONTRIBUTING.md for environment setup, architecture, regression tests and the full scripts\verify.ps1 command. That serial command produces clean packages, external-consumer results, latency reports and one checksum/provenance manifest in a new output directory.

The measured warm release Rust p95 was below 1 ms separately for representative short and medium cohorts on the inspected Windows host. This is not an all-input, cold-start, Python, concurrent-call or end-to-end audio guarantee. See PERFORMANCE.md for the measurement protocol and limits.

§Acknowledgements

Thanks to @canberk7 for practical TTS feedback and suggestions on making the Python package easier to distribute.

§License and boundaries

Owned code and independently authored tables are Apache-2.0. Third-party notices preserve dependency and Unicode rights. This repository contains the normalizer and Python binding only: no speech-engine integration, external model source, weights, credentials or audio.

Structs§

Hint
Caller intent at original grapheme-safe byte coordinates.
Issue
Non-sensitive diagnostic for an unresolved original-source range.
NormalizeOptions
Per-call options. No implicit guessing policy is provided.
NormalizeResult
Owned immutable result. Completeness concerns TN work, not voice quality.
Normalizer
Immutable, thread-safe normalizer. Clones share validated built-in resources.
Segment
One member of an ordered, contiguous original-source partition.
SourceRange
Original-source half-open UTF-8 byte range.
WorkControl
Runtime-neutral cooperative control. Clones share the cancellation signal.

Enums§

AmbiguityPolicy
How unresolved linguistic expressions are handled.
HintKind
Explicit interpretation of a whole original-source span.
IssueCategory
Machine-readable reason for preserved linguistic work.
LimitKind
Resource limit which was exceeded.
NormalizeError
Explicit input, policy, resource, control, or engine failure.
SegmentKind
Semantic kind of a result segment.

Constants§

MAX_CANDIDATES
Maximum number of semiotic candidate records, including unresolved spans.
MAX_HINTS
Maximum number of caller hints.
MAX_INPUT_BYTES
Maximum original UTF-8 input length in bytes.
MAX_RESULT_BYTES
Maximum logical bytes of owned result and diagnostic allocations.
NORMALIZER_ID
Diagnostic identity of this built-in normalizer, not a selectable behavior profile.