Skip to content

Compounds, Quantities, and Names ​

Suzume can keep a compound or quantity together while leaving a productive suffix or administrative boundary visible. The examples below group those decisions by the kind of word being analyzed.

All examples use the comparison baseline. For labels on the same tokens, see POS Classification.

Compound words ​

Kanji Compounds ​

An unregistered kanji run with no dictionary or grammatical evidence for an internal boundary is normally kept as one noun candidate. Structural rules can still split forms such as 神奈川県 / 横浜市 and 会議 / 中.

Input経済成長
MeCab
経済名詞成長名詞
2 tokens
Suzume
経済成長名詞NOUN
1 token
Input開始予定
MeCab
開始名詞予定名詞
2 tokens
Suzume
開始予定名詞NOUN
1 token

Katakana Compounds ​

An unknown ordinary katakana run is normally kept as one noun candidate. Dictionary and shape rules can assign another POS, as with mimetic ドキドキ(ADV).

Inputサンプルデータ
MeCab
サンプル名詞データ名詞
2 tokens
Suzume
サンプルデータ名詞NOUN
1 token
Inputセットリスト
MeCab
セット名詞リスト名詞
2 tokens
Suzume
セットリスト名詞NOUN
1 token

Mixed-Script Compounds ​

An alphabetic term followed by a katakana term is one compound noun.

InputAIブーム
MeCab
AIブーム
2 tokens
Suzume
AIブーム名詞NOUN
1 token

Deverbal Compound Nouns ​

A verb continuative plus 会, and the destination suffix 行き, form single event/route nouns.

Input飲み会
MeCab
飲み会
2 tokens
Suzume
飲み会名詞NOUN
1 token
Input東京行き
MeCab
東京行き
2 tokens
Suzume
東京行き名詞NOUN
1 token

A compound verb's continuative used as a noun can remain separate from a following noun: 取り扱い方法 becomes 取り扱い(NOUN) / 方法(NOUN).

Noun + Single-Character Suffixes ​

The following closed set of noun + single-character suffix combinations is merged.

Input報告書
MeCab
報告書
2 tokens
Suzume
報告書名詞NOUN
1 token
Input成功率
MeCab
成功率
2 tokens
Suzume
成功率名詞NOUN
1 token

Applies to suffixes: 書, 誌, 時, 率, 性

Verb Stem + 方 ​

When the formal noun 方 follows a short verb stem, Suzume merges the expression into one search unit denoting a method.

Input走り方
MeCab
走り方
2 tokens
Suzume
走り方名詞NOUN
1 token

Only short stems merge: the continuative must be at most two characters (走り方, やり方). Longer continuatives keep the boundary — 打ち合わせ方 stays 打ち合わせ / 方. The same boundary applies to the deverbal noun in 取り扱い方, which becomes 取り扱い(NOUN) / 方(SUFFIX).

Quantities and dates ​

Numbers and Units ​

Cardinal numbers followed by counters or units are normally merged into one quantity token. This includes large number units (万, 億, 兆), decimal numbers, percentages, and alphabetic units. Ordinal and structural suffix rules can split forms such as 第三 / 回.

Input3人
MeCab
3人
2 tokens
Suzume
3人名詞NOUN
1 token
Input100円
MeCab
100円
2 tokens
Suzume
100円名詞NOUN
1 token
Input3.14
MeCab
3.14
3 tokens
Suzume
3.14名詞NOUN
1 token

Suzume also preserves search-unit boundaries around quantities.

Input徒歩五分
MeCab
徒歩五分
3 tokens
Suzume
徒歩五分名詞NOUN
2 tokens
Input三ヶ月間入院
MeCab
三ヶ月間入院
4 tokens
Suzume
三ヶ月間名詞NOUN入院
2 tokens

Quantity phrases such as 3種類, 3人分, and 3ページ目 stay whole as NOUN tokens. A counter does not cut through a following word, and quantity suffixes such as 分 and the ordinal 目 remain part of the quantity.

Comma-grouped numerals stay whole. A following counter remains a separate SUFFIX, while a currency amount stays one search unit.

Input1,000人
MeCab
1,000人
4 tokens
Suzume
1,000名詞NOUN人接尾辞SUFFIX
2 tokens
Input1,000円
MeCab
1,000円
4 tokens
Suzume
1,000円名詞NOUN
1 token

The same merging applies beyond Arabic numerals.

Kana-spelled quantities:

Inputよんにん
MeCab
よ形容詞ん名詞に助詞ん助詞
4 tokens
Suzume
よんにん名詞NOUN
1 token

Distributive quantities:

Input一語一語
MeCab
一語名詞一名詞語名詞
3 tokens
Suzume
一語一語名詞NOUN
1 token

Address and lot numbers:

Input1-2-3
MeCab
1-2-3
5 tokens
Suzume
1-2-3名詞NOUN
1 token

Ordinal 第 versus approximate 約: the ordinal prefix 第 merges with its number, while the following counter stays a separate SUFFIX token. The approximation prefix 約 instead stays a separate PREFIX, and the number merges with its counter.

Input第三回
MeCab
第三回
3 tokens
Suzume
第三名詞NOUN回接尾辞SUFFIX
2 tokens
Input約三人
MeCab
約三人
3 tokens
Suzume
約接頭辞PREFIX三人名詞NOUN
2 tokens

Dates ​

Full date expressions are merged into a single token.

Input2024年12月23日
MeCab
2024年12月23日
6 tokens
Suzume
2024年12月23日名詞NOUN
1 token

Names and suffixes ​

Proper Nouns and Place Names ​

Many place-name components with region suffixes are merged. The structural 県+市 rule is an exception and splits the prefecture from the city.

Input東京都新宿区
MeCab
東京都新宿区
4 tokens
Suzume
東京都新宿区名詞NOUN
1 token

place name

Prefecture + City ​

Prefecture-city compound nouns are split at administrative boundaries.

Input神奈川県横浜市
MeCab
神奈川県横浜市
4 tokens
Suzume
神奈川県横浜市
2 tokens

split at the 県 / 市 boundary

Note: This split rule applies only to the 県+市 pattern. Other combinations like 都+区 (東京都新宿区) or 府+市 (大阪府大阪市) are merged into single tokens by the Proper Nouns and Place Names rule.

Honorific Suffixes and Hiragana Nicknames ​

Honorific suffixes are split from names.

Applies to suffixes: さん, ちゃん, くん, 様, さま. The kanji forms 君 and 殿 split after a host of at least two kanji, as in 佐藤 / 君 and 先生 / 殿.

The runtime POS can vary among these written titles: 佐藤様 becomes 佐藤(NOUN) / 様(NOUN).

Input佐藤君
MeCab
佐藤名詞・固有名詞君名詞・接尾
2 tokens
Suzume
佐藤名詞NOUN君接尾辞SUFFIX
2 tokens

After a single kanji, Suzume cannot distinguish a name from an ordinary compound such as 主君. 林君 therefore stays whole; this is a lexical limitation.

Input林君
MeCab
林名詞・固有名詞君名詞・接尾
2 tokens
Suzume
林君名詞NOUN
1 token

Exceptions: family terms like お兄ちゃん and お母さん, and the collective forms 皆様 / 皆さん, are kept as single tokens.

Short nicknames made from a two- or three-character hiragana stem followed by ちゃん or くん, along with lexicalized family terms, can merge as a search unit. Ordinary さん remains a separate suffix, including after names written in kanji.

Inputわんちゃん
MeCab
わんちゃん
2 tokens
Suzume
わんちゃん名詞NOUN
1 token

Verb Stem + Productive Suffix ​

Productive suffixes after a verb stem — がち (tendency), たて (freshness), っぱなし (left as-is) — keep their boundary and are tagged SUFFIX, so the verb stem stays searchable with its lemma.

Input忘れがち
MeCab
忘れ動詞がち名詞・接尾
2 tokens
Suzume
忘れ動詞VERB→忘れるがち接尾辞SUFFIX
2 tokens
Inputできたて
MeCab
でき動詞た助動詞て助詞
3 tokens
Suzume
でき動詞VERB→できるたて接尾辞SUFFIX
2 tokens
Input開けっぱなし
MeCab
開けっぱなし名詞
1 token
Suzume
開け動詞VERB→開けるっぱなし接尾辞SUFFIX
2 tokens

The negative construction 読みっこない follows the same principle: 読み (VERB) / っこ (SUFFIX) / ない (ADJ).

Quantity and State Suffixes ​

Even though Suzume merges kanji compounds aggressively, productive quantity/state suffixes keep their boundary and are tagged SUFFIX.

Input二階建て
MeCab
二階建て
3 tokens
Suzume
二階名詞NOUN建て接尾辞SUFFIX
2 tokens
Input砂糖抜き
MeCab
砂糖名詞抜き名詞
2 tokens
Suzume
砂糖名詞NOUN抜き接尾辞SUFFIX
2 tokens
Input会議中に
MeCab
会議名詞中名詞に助詞
3 tokens
Suzume
会議名詞NOUN中接尾辞SUFFIXに助詞PARTICLE
3 tokens

例年並み splits as 例年(NOUN) / 並み(SUFFIX), and 汗まじり as 汗(NOUN) / まじり(SUFFIX). 中 depends on its host: after an attributive adjective, 忙しい中 gives 忙しい(ADJ) / 中(NOUN), with 中 as a formal noun rather than a suffix.

Pure relabeling suffixes with no boundary change, such as the nominalizer さ, are covered in POS Classification.