Skip to content

Differences from MeCab ​

MeCab and Suzume can assign different word boundaries and POS labels to the same text. Choose a topic below to see concrete comparisons.

The source of truth for intentional differences is the Python MeCab normalization pipeline used to generate test expectations.

Find a topic ​

What you want to compareDetails
Compounds, numbers and units, place names, suffix boundariesCompounds, Quantities, and Names
Compound verbs and auxiliaries, particles, negatives, classical endingsVerbs and Grammar
Colloquial spelling, prolonged sounds, identifiers, URLs, symbolsText and Notation
POS labels, lemmas, and extended POS on the same tokensPOS Classification

Comparison Baseline ​

Every MeCab boundary and displayed feature in this comparison guide was recorded from MeCab 0.996 with mecab-ipadic 2.7.0-20070801 (the UTF-8 IPA dictionary). The labels abbreviate MeCab's comma-separated feature fields for readability; the comparison does not rewrite its token boundaries. It invokes mecab with the default configuration and no user dictionary. You can inspect the active setup with:

bash
mecab --version
mecab -D

MeCab's output depends on the selected dictionary, its costs, and user dictionaries. UniDic, NEologd, or a customized IPA dictionary can produce different boundaries and labels from these examples. The normalization pipeline linked above starts from this IPA-dictionary output and applies Suzume's generalized comparison rules when generating test expectations.

Reading the comparisons

In each comparison the MeCab row uses MeCab's Japanese POS names, and the Suzume row uses the public API tags (NOUN, VERB, ADJ, …) with the Japanese POS name shown next to each code. A colored underline marks every token, so where the two tokenizers split or merge is visible at a glance.

Design Philosophy ​

MeCab analyzes morphemes using the vocabulary, taxonomy, and connection costs of the selected dictionary. Suzume uses compact dictionaries and character, grammar, and connection rules to produce tokens for search and display.

MeCabSuzume
ApproachDictionary-drivenFeature-driven
DictionaryExternal dictionary required; size and output depend on the selected dictionaryCompact dictionaries embedded in the WASM binary (~260KiB gzipped)
Unknown wordsFalls back to character typesPattern-based candidate generation
Compound handlingBoundaries follow the selected dictionaryDictionary and structural rules; unknown runs may merge
TargetDetailed dictionary-based analysisCompact search/display tokenization across browser, edge, and native runtimes

Constraints ​

These are known limitations arising from Suzume's feature-based architecture.

Cannot Split Merged Compounds ​

Suzume cannot infer arbitrary lexical boundaries inside an otherwise unknown same-script compound. It can still split boundaries licensed by grammatical rules or compact-dictionary entries.

Input東京都庁前
MeCab
東京都庁前
3 tokens
Suzume
東京都庁前名詞NOUN
1 token

Suzume cannot determine the internal boundaries; MeCab splits them from its dictionary

Workaround: Use the runtime-loading examples in the user-dictionary guide to register boundaries required by your application. That page covers JavaScript, Python, Go, C++, C, and the native CLI.

When to Use Which ​

Use CaseRecommendation
Browser / client-side appsSuzume — no server required
Search indexing / tag extractionSuzume — compound merging is often desirable
Compatibility with a particular MeCab dictionary/corpusMeCab — preserves that dictionary's boundaries and taxonomy
Real-time UI (input-as-you-type)Suzume — fast, no network latency
Dictionary-defined compound word splittingMeCab — boundaries come from the selected dictionary
Pattern candidates for words absent from the dictionarySuzume — can analyze character sequences without a lexical entry