Text and Notation
Word boundaries also depend on how text is written. These examples cover everyday kana, expressive spelling, identifiers, and symbols.
All examples use the comparison baseline. For labels on the same tokens, see POS Classification.
Everyday and expressive text
Everyday Hiragana Words
Common words normally written in hiragana are kept intact through pattern rules and the compact L2 dictionary. Rules cover forms such as おととい, ひこうき, みっつ, and calendar compounds such as 翌営業日. Lexical evidence resolves ambiguous all-hiragana nouns whose characters can also be particles or inflectional endings: みず, てがみ, ひらがな, にわ, いりぐち, はにわ, あけぼの, and くだもの.
Slang and Modern Words
Modern colloquial adjectives and verbs are recognized natively.
Recognized examples include エモい, キモい, ウザい, ダサい, イタい, ヤバい, and compound i-adjectives. Hiragana variants can require context: 頭がいたい gives 頭(NOUN) / が(PARTICLE) / いたい(ADJ), while standalone いたい gives い(VERB) / たい(AUX).
Recognized verb examples include バズる, ググる, and パクる.
Nai-Adjectives
Certain adjectives ending in ない are treated as single lexical units rather than being split.
Main examples handled as one token: だらしない, つまらない, もったいない, くだらない, いたたまれない, ものたりない, こころもとない
This is a closed word list — productive "stem + ない" combinations still split; see Productive Negative Splitting.
Prolonged Sound Marks
Prolonged sound marks (ー) are merged with the preceding token. For recognized colloquial i-adjectives, Suzume keeps the marks in the surface while normalizing the lemma to the ordinary dictionary form.
Kanji adjective stems also retain expressive spelling: 高ーい and clipped 高っ are each one ADJ with lemma 高い. An emphatic internal っ can remain in the lemma: すっごい is one ADJ with lemma すっごい.
A small vowel extending a particle stays with that particle: のにぃ is one PARTICLE with lemma のに. A clipped greeting such as ありがとっ is one INTJ with lemma ありがとう.
Emphatic Colloquial Particles
Colloquial emphatic particles are split as single units instead of being fragmented.
ったら topic particle:
ってば emphatic particle:
Filler Decomposition
Fixed conversational phrases lexicalized as fillers are decomposed into their grammatical parts.
Technical notation
Technical Text
Technical identifiers are merged into single tokens.
Snake_case identifiers:
Version numbers:
ASCII name + number:
ASCII dot notation:
ASCII word-internal separators:
The same rule covers apostrophes, ampersands, and slashes when they occur between ASCII word characters.
URLs, Mentions, and Hashtags
URLs, @mentions, and #hashtags are merged into single tokens. A hashtag scanner accepts Japanese text, including hiragana that would otherwise be particle-like, and stops at whitespace or punctuation.
Symbols and punctuation
Content Symbols and Punctuation
Currency and unit signs, arrows, mathematical or technical marks, and emoji remain in the default output as OTHER. They carry text content and keep the token offsets covering that content. Punctuation-like characters are SYMBOL tokens and are omitted by default; enable preserveSymbols to keep them.
| Options | 価格は€50🎉。 |
|---|---|
| Default | 価格(NOUN) / は(PARTICLE) / €(OTHER) / 50(NOUN) / 🎉(OTHER) |
preserveSymbols: true | 価格(NOUN) / は(PARTICLE) / €(OTHER) / 50(NOUN) / 🎉(OTHER) / 。(SYMBOL) |