Speed and Accuracy
This page reports a reproducible Node/WASM speed run and an in-sample boundary-agreement measurement. Each figure states what it covers so you can compare it with your own workload.
WASM speed under Node
The script measures the public JavaScript API, including result decoding. It first creates shared-runtime handles, so the creation row is measured after the WASM module cache is warm; it does not measure a cold download or module startup.
| Median | |
|---|---|
| Create an analyzer in a loaded/shared WASM runtime | 2.386 ms |
| First analysis after creation | 0.426 ms |
| Steady-state analysis, per text | 0.300291 ms |
| Steady-state throughput | 13,320 tokens/sec |
make build
make wasm
(cd bindings/wasm && yarn install --immutable && yarn build:js)
node scripts/measure_wasm_metrics.mjs --instances=3 --iterations=500 --samples=5 --warmup=1This run was measured on October 9, 2026, using Suzume 0.9.14 (commit b0eaf611) on an Apple M5 Max (arm64) with Node v24.21.0 and the script's three built-in short texts. Pass --corpus=/path/to/corpus.txt to measure another input set. These are Node measurements; they do not predict browser or phone timings.
The playground on Getting Started and How It Works runs its own browser measurement on your device and prints the result below the output. The native CLI has a separate benchmark command:
suzume-cli test benchmark --iterations=500 --samples=5 --warmup=1These checks use separate implementations and measurement conditions; compare results within the environment you are testing. The WASM binary is 260KiB gzipped, loaded once and cached thereafter. This size includes the embedded dictionaries but excludes the JavaScript loader and API files.
Boundary agreement on a fixture subset
The script scores cases under tests/data/tokenization whose expected surfaces reassemble to the raw single-line input. It invokes the native CLI with its default dictionary loading. The full universal_tokenization_test suite uses skip_user_dictionary=true, while this script skips cases whose input would be normalized or contains a newline. Treat the result as an in-sample regression measure for this script's subset, not as full-suite health.
Suzume is not measured by agreement with MeCab, since the two do not aim to produce interchangeable output — see Differences from MeCab.
| Score | |
|---|---|
| Boundary F1 | 0.9998 |
| Boundary precision / recall | 0.9995 / 1.0000 |
| Token F1 | 0.9996 |
| Token precision / recall | 0.9995 / 0.9998 |
| Sentences segmented exactly | 0.9992 (6,628 / 6,633) |
Scored on October 9, 2026, using Suzume 0.9.14 (commit b0eaf611), over 6,633 cases and 24,144 tokens.
make dict
python3 scripts/measure_segmentation_accuracy.py --per-categoryScope of this score
These cases are Suzume's own test suite, and the tokenizer is fixed until it passes them. The score does not estimate how Suzume handles text it has never seen, and the script's case filtering and dictionary configuration differ from the full native fixture suite.
Do not use it to compare Suzume against another tokenizer or as an expected accuracy for your own corpus. Run your own text through the live demo or the CLI instead.
Boundary scores count agreement on interior boundaries between adjacent tokens; document edges are excluded. Token scores require both edges for each token, and sentence exactness requires the complete segmentation to match. Use the --per-category output to inspect the current fixture breakdown.
What is not measured here
- Comparative accuracy. There is no table putting Suzume's F1 next to another tokenizer's on a shared corpus. Doing that fairly needs an annotation standard both tools target, and Suzume deliberately does not target MeCab's. See Differences from MeCab.
- Held-out accuracy. The reported cases are in-sample; a held-out estimate needs text that was never used to fix a bug here.
- Memory under adversarial input. Allocation counts require an instrumented build; the released artifact does not report them.