Japanese Text Normalizer & Invisible Character Detector
A tool that unifies variations in Japanese text — full-width/half-width characters, half-width katakana, symbols, hyphens and tildes, punctuation, spaces, and line breaks — and finds and removes invisible characters such as zero-width spaces, BOMs, and bidi controls, plus platform-dependent characters such as circled numbers and ㈱. You can review a diff and a per-item count of changes, and everything runs in your browser.
How to use
- Paste text into the "Before" box, or drag and drop a text file.
- Choose a preset (Excel/CSV import, database entry, proofreading, or copy-paste cleanup), or set each checkbox and option individually.
- The result appears in the "After" box. Check the per-item counts under "Changes" and the changed spots under "Differences".
- "Special characters found" lists the position and type of invisible and platform-dependent characters and marks them in the text. "Remove invisible characters" strips them from the input in one click.
- Save the result with "Copy to clipboard" or "Download".
About this tool
The Japanese Text Normalizer cleans up variations common in Japanese text, such as letters and digits mixing full-width and half-width forms, half-width katakana, and different kinds of hyphens and wave dashes. It's useful for preparing data before importing it into CSV files or databases, preprocessing documents for proofreading, and improving search and duplicate detection. Pair it with the Character Count tool when you also need to count characters.
Text copied from web pages, chats, or PDFs can contain invisible characters such as zero-width spaces, BOMs, no-break spaces, and controls that change text direction. These can make matching fail in programs or Excel and cause unexpected display problems. This tool lists their positions and types and marks them in red in the text.
When unifying hyphens and dashes, those right after katakana or hiragana are treated as the prolonged sound mark "ー", so "コンピュ-タ" becomes "コンピュータ" while "03-1234" becomes "03-1234". Half-width ",", ".", and "~" are converted only when next to Japanese text or digits, so numbers, English, and URLs stay intact. All conversions use built-in tables, and your text is never sent to a server.
Frequently asked questions
What is NFKC?
It's one of Unicode's normalization forms; it converts full-width letters and digits, half-width katakana, and compatibility characters such as ① and ㌔ into standard characters. Because it converts so broadly, meaning can change (for example, ① becomes 1). For finer control, choose NFC and set each item individually.
Can removing invisible characters change emoji?
Some emoji, such as family emoji, join several emoji with zero-width joiners (ZWJ) to display as one. Removing ZWJ splits them into separate emoji, so check the result when your text contains emoji.
What are variation selectors (IVS)?
They are invisible characters placed right after a kanji such as 葛 to specify its glyph. Because they're sometimes used intentionally, for example in names, this tool only detects and shows them and does not remove them with the invisible character cleanup.
Is my text sent anywhere?
No. All conversion and detection happen in your browser.