Skip to content

to_entropy() cannot read the Japanese mnemonics to_mnemonic() writes - #145

Open
fametrano wants to merge 1 commit into
trezor:masterfrom
fametrano:japanese_delimiter_read_back
Open

fametrano wants to merge 1 commit into
trezor:masterfrom
fametrano:japanese_delimiter_read_back

Conversation

@fametrano

@fametrano fametrano commented Aug 5, 2026 •

Copy link
Copy Markdown

to_entropy() cannot read the Japanese mnemonics to_mnemonic() writes. Mnemonic("japanese") joins words with U+3000, the ideographic space, as the BIP39 vectors do, but to_entropy and expand split on " " and so see a single word: to_entropy raises and expand returns the sentence unchanged. On vectors.json, to_entropy fails on all 24 Japanese mnemonics and on none of the other 264.

#110 split on self.delimiter and was reverted for failing CI. That cannot work in check, which applies NFKD first and so turns U+3000 into a plain space before splitting. This PR splits on any run of whitespace, as detect_language already does. It also makes to_seed collapse whitespace the same way, so that input check rejects, such as a trailing newline or a tab, no longer gives a different seed from the same words; if you would rather not change that, drop the to_seed line and the last assertion of test_whitespace_runs. All 288 vectors are unchanged, no input check accepts today gets a different seed, and the passphrase is not touched. Locally the tests pass on Python 3.8 to 3.14 and PyPy 3.11, and black, isort, flake8 and pyright report nothing.

Written with LLM, a computer and the Internet; the mistakes, as usual, are all mine.

@fametrano

Copy link
Copy Markdown
Author

@prusnak could you approve the workflow run here? This is my first pull request from this fork, so CI has never started, and the local run in the description is the only evidence in the thread.

The branch is a single commit now. Since filing I added one test for expand, the fourth line the patch changes and the one that had no coverage.

@fametrano

Copy link
Copy Markdown
Author

@prusnak the workflow runs expired before approval, so CI still hasn't run. Could you re-run them, or should I push a new commit?

@fametrano fametrano changed the title to_entropy() cannot read the japanese mnemonics to_mnemonic() writes to_entropy() cannot read the Japanese mnemonics to_mnemonic() writes Sep 29, 2026
@fametrano
fametrano force-pushed the japanese_delimiter_read_back branch from bd50ae2 to b5e1375 Compare September 29, 2026 21:09
to_mnemonic joins a Japanese mnemonic with U+3000, while to_entropy,
check and expand split on " ". So the library cannot read the sentences
it writes: to_entropy raises on all 24 Japanese vectors of vectors.json,
and expand returns the sentence unchanged.

check splits after NFKD, which maps U+3000 to U+0020. So it cannot
split on self.delimiter: that was trezor#110, reverted in df3e150 for failing
CI. Split on any run of whitespace instead, as detect_language already
does.

to_seed follows the same rule. Before, input that check rejects gave a
different seed from the same words separated by single spaces.

The tests fail without the fix: to_entropy raises on every Japanese
vector, check rejects a sentence separated by anything but one space,
and expand returns a tab- or U+3000-separated sentence unexpanded.
The existing round-trip test hides the first failure, because it splits
the sentence itself before passing it in.
@fametrano
fametrano force-pushed the japanese_delimiter_read_back branch from b5e1375 to 66c1898 Compare October 2, 2026 16:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant