Compounds, Quantities, and Names
Suzume can keep a compound or quantity together while leaving a productive suffix or administrative boundary visible. The examples below group those decisions by the kind of word being analyzed.
All examples use the comparison baseline. For labels on the same tokens, see POS Classification.
Compound words
Kanji Compounds
An unregistered kanji run with no dictionary or grammatical evidence for an internal boundary is normally kept as one noun candidate. Structural rules can still split forms such as 神奈川県 / 横浜市 and 会議 / 中.
Katakana Compounds
An unknown ordinary katakana run is normally kept as one noun candidate. Dictionary and shape rules can assign another POS, as with mimetic ドキドキ(ADV).
Mixed-Script Compounds
An alphabetic term followed by a katakana term is one compound noun.
Deverbal Compound Nouns
A verb continuative plus 会, and the destination suffix 行き, form single event/route nouns.
A compound verb's continuative used as a noun can remain separate from a following noun: 取り扱い方法 becomes 取り扱い(NOUN) / 方法(NOUN).
Noun + Single-Character Suffixes
The following closed set of noun + single-character suffix combinations is merged.
Applies to suffixes: 書, 誌, 時, 率, 性
Verb Stem + 方
When the formal noun 方 follows a short verb stem, Suzume merges the expression into one search unit denoting a method.
Only short stems merge: the continuative must be at most two characters (走り方, やり方). Longer continuatives keep the boundary — 打ち合わせ方 stays 打ち合わせ / 方. The same boundary applies to the deverbal noun in 取り扱い方, which becomes 取り扱い(NOUN) / 方(SUFFIX).
Quantities and dates
Numbers and Units
Cardinal numbers followed by counters or units are normally merged into one quantity token. This includes large number units (万, 億, 兆), decimal numbers, percentages, and alphabetic units. Ordinal and structural suffix rules can split forms such as 第三 / 回.
Suzume also preserves search-unit boundaries around quantities.
Quantity phrases such as 3種類, 3人分, and 3ページ目 stay whole as NOUN tokens. A counter does not cut through a following word, and quantity suffixes such as 分 and the ordinal 目 remain part of the quantity.
Comma-grouped numerals stay whole. A following counter remains a separate SUFFIX, while a currency amount stays one search unit.
The same merging applies beyond Arabic numerals.
Kana-spelled quantities:
Distributive quantities:
Address and lot numbers:
Ordinal 第 versus approximate 約: the ordinal prefix 第 merges with its number, while the following counter stays a separate SUFFIX token. The approximation prefix 約 instead stays a separate PREFIX, and the number merges with its counter.
Dates
Full date expressions are merged into a single token.
Names and suffixes
Proper Nouns and Place Names
Many place-name components with region suffixes are merged. The structural 県+市 rule is an exception and splits the prefecture from the city.
place name
Prefecture + City
Prefecture-city compound nouns are split at administrative boundaries.
split at the 県 / 市 boundary
Note: This split rule applies only to the 県+市 pattern. Other combinations like 都+区 (東京都新宿区) or 府+市 (大阪府大阪市) are merged into single tokens by the Proper Nouns and Place Names rule.
Honorific Suffixes and Hiragana Nicknames
Honorific suffixes are split from names.
Applies to suffixes: さん, ちゃん, くん, 様, さま. The kanji forms 君 and 殿 split after a host of at least two kanji, as in 佐藤 / 君 and 先生 / 殿.
The runtime POS can vary among these written titles: 佐藤様 becomes 佐藤(NOUN) / 様(NOUN).
After a single kanji, Suzume cannot distinguish a name from an ordinary compound such as 主君. 林君 therefore stays whole; this is a lexical limitation.
Exceptions: family terms like お兄ちゃん and お母さん, and the collective forms 皆様 / 皆さん, are kept as single tokens.
Short nicknames made from a two- or three-character hiragana stem followed by ちゃん or くん, along with lexicalized family terms, can merge as a search unit. Ordinary さん remains a separate suffix, including after names written in kanji.
Verb Stem + Productive Suffix
Productive suffixes after a verb stem — がち (tendency), たて (freshness), っぱなし (left as-is) — keep their boundary and are tagged SUFFIX, so the verb stem stays searchable with its lemma.
The negative construction 読みっこない follows the same principle: 読み (VERB) / っこ (SUFFIX) / ない (ADJ).
Quantity and State Suffixes
Even though Suzume merges kanji compounds aggressively, productive quantity/state suffixes keep their boundary and are tagged SUFFIX.
例年並み splits as 例年(NOUN) / 並み(SUFFIX), and 汗まじり as 汗(NOUN) / まじり(SUFFIX). 中 depends on its host: after an attributive adjective, 忙しい中 gives 忙しい(ADJ) / 中(NOUN), with 中 as a formal noun rather than a suffix.
Pure relabeling suffixes with no boundary change, such as the nominalizer さ, are covered in POS Classification.