Skip to content

Text and Notation ​

Word boundaries also depend on how text is written. These examples cover everyday kana, expressive spelling, identifiers, and symbols.

All examples use the comparison baseline. For labels on the same tokens, see POS Classification.

Everyday and expressive text ​

Everyday Hiragana Words ​

Common words normally written in hiragana are kept intact through pattern rules and the compact L2 dictionary. Rules cover forms such as おととい, ひこうき, みっつ, and calendar compounds such as 翌営業日. Lexical evidence resolves ambiguous all-hiragana nouns whose characters can also be particles or inflectional endings: みず, てがみ, ひらがな, にわ, いりぐち, はにわ, あけぼの, and くだもの.

Slang and Modern Words ​

Modern colloquial adjectives and verbs are recognized natively.

Inputエモい
MeCab
エモ名詞い動詞
2 tokens
Suzume
エモい形容詞ADJ
1 token

Recognized examples include エモい, キモい, ウザい, ダサい, イタい, ヤバい, and compound i-adjectives. Hiragana variants can require context: 頭がいたい gives 頭(NOUN) / が(PARTICLE) / いたい(ADJ), while standalone いたい gives い(VERB) / たい(AUX).

Recognized verb examples include バズる, ググる, and パクる.

Nai-Adjectives ​

Certain adjectives ending in ない are treated as single lexical units rather than being split.

Inputだらしない
MeCab
だらし名詞・ナイ形容詞語幹ない助動詞
2 tokens
Suzume
だらしない形容詞ADJ
1 token

Main examples handled as one token: だらしない, つまらない, もったいない, くだらない, いたたまれない, ものたりない, こころもとない

This is a closed word list — productive "stem + ない" combinations still split; see Productive Negative Splitting.

Prolonged Sound Marks ​

Prolonged sound marks (ー) are merged with the preceding token. For recognized colloquial i-adjectives, Suzume keeps the marks in the surface while normalizing the lemma to the ordinary dictionary form.

Inputそうー
MeCab
そう副詞ー名詞
2 tokens
Suzume
そうー副詞ADV
1 token
Inputすごーーい
MeCab
すご形容詞ーー名詞い名詞
3 tokens
Suzume
すごーーい形容詞ADJ→すごい
1 token

Kanji adjective stems also retain expressive spelling: 高ーい and clipped 高っ are each one ADJ with lemma 高い. An emphatic internal っ can remain in the lemma: すっごい is one ADJ with lemma すっごい.

A small vowel extending a particle stays with that particle: のにぃ is one PARTICLE with lemma のに. A clipped greeting such as ありがとっ is one INTJ with lemma ありがとう.

Emphatic Colloquial Particles ​

Colloquial emphatic particles are split as single units instead of being fragmented.

ったら topic particle:

Inputあなたったら
MeCab
あな名詞たっ動詞たら助動詞
3 tokens
Suzume
あなた代名詞PRONったら助詞PARTICLE
2 tokens

ってば emphatic particle:

Inputもうってば
MeCab
も助詞うっ動詞て助詞ば助詞
4 tokens
Suzume
もう副詞ADVってば助詞PARTICLE
2 tokens

Filler Decomposition ​

Fixed conversational phrases lexicalized as fillers are decomposed into their grammatical parts.

Inputそうですね
MeCab
そうですねフィラー
1 token
Suzume
そう副詞ADVです助動詞AUXね助詞PARTICLE
3 tokens

Technical notation ​

Technical Text ​

Technical identifiers are merged into single tokens.

Snake_case identifiers:

Inputuser_name
MeCab
user_name
3 tokens
Suzume
user_name
1 token

Version numbers:

Inputv1.2.3
MeCab
v1.2.3
6 tokens
Suzume
v1.2.3
1 token

ASCII name + number:

InputModel15
MeCab
Model15
2 tokens
Suzume
Model15
1 token

ASCII dot notation:

Inputconsole.log
MeCab
console.log
3 tokens
Suzume
console.log
1 token

ASCII word-internal separators:

Inputdata-driven
MeCab
data-driven
3 tokens
Suzume
data-driven
1 token

The same rule covers apostrophes, ampersands, and slashes when they occur between ASCII word characters.

URLs, Mentions, and Hashtags ​

URLs, @mentions, and #hashtags are merged into single tokens. A hashtag scanner accepts Japanese text, including hiragana that would otherwise be particle-like, and stops at whitespace or punctuation.

Inputhttps://example.com にアクセス
MeCab
https://example.comにアクセス
7 tokens
Suzume
https://example.comにアクセス
3 tokens
Input@user_name に送信
MeCab
@user_nameに送信
6 tokens
Suzume
@user_nameに送信
3 tokens
Input#topicについて
MeCab
#topicについて
3 tokens
Suzume
#topicについて名詞NOUN
1 token
Input#日本語タグ
MeCab
#日本語タグ
3 tokens
Suzume
#日本語タグ名詞NOUN
1 token

Symbols and punctuation ​

Content Symbols and Punctuation ​

Currency and unit signs, arrows, mathematical or technical marks, and emoji remain in the default output as OTHER. They carry text content and keep the token offsets covering that content. Punctuation-like characters are SYMBOL tokens and are omitted by default; enable preserveSymbols to keep them.

Options価格は€50🎉。
Default価格(NOUN) / は(PARTICLE) / €(OTHER) / 50(NOUN) / 🎉(OTHER)
preserveSymbols: true価格(NOUN) / は(PARTICLE) / €(OTHER) / 50(NOUN) / 🎉(OTHER) / 。(SYMBOL)