Conversation
This was referenced Aug 5, 2026
fametrano
force-pushed
the
japanese_delimiter_read_back
branch
from
August 13, 2026 07:59
2cb21cc to
bd50ae2
Compare
Author
|
@prusnak could you approve the workflow run here? This is my first pull request from this fork, so CI has never started, and the local run in the description is the only evidence in the thread. The branch is a single commit now. Since filing I added one test for |
Author
|
@prusnak the workflow runs expired before approval, so CI still hasn't run. Could you re-run them, or should I push a new commit? |
fametrano
force-pushed
the
japanese_delimiter_read_back
branch
from
September 29, 2026 21:09
bd50ae2 to
b5e1375
Compare
to_mnemonic joins a Japanese mnemonic with U+3000, while to_entropy, check and expand split on " ". So the library cannot read the sentences it writes: to_entropy raises on all 24 Japanese vectors of vectors.json, and expand returns the sentence unchanged. check splits after NFKD, which maps U+3000 to U+0020. So it cannot split on self.delimiter: that was trezor#110, reverted in df3e150 for failing CI. Split on any run of whitespace instead, as detect_language already does. to_seed follows the same rule. Before, input that check rejects gave a different seed from the same words separated by single spaces. The tests fail without the fix: to_entropy raises on every Japanese vector, check rejects a sentence separated by anything but one space, and expand returns a tab- or U+3000-separated sentence unexpanded. The existing round-trip test hides the first failure, because it splits the sentence itself before passing it in.
fametrano
force-pushed
the
japanese_delimiter_read_back
branch
from
October 2, 2026 16:35
b5e1375 to
66c1898
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
to_entropy()cannot read the Japanese mnemonicsto_mnemonic()writes.Mnemonic("japanese")joins words with U+3000, the ideographic space, as the BIP39 vectors do, butto_entropyandexpandsplit on" "and so see a single word:to_entropyraises andexpandreturns the sentence unchanged. Onvectors.json,to_entropyfails on all 24 Japanese mnemonics and on none of the other 264.#110 split on
self.delimiterand was reverted for failing CI. That cannot work incheck, which applies NFKD first and so turns U+3000 into a plain space before splitting. This PR splits on any run of whitespace, asdetect_languagealready does. It also makesto_seedcollapse whitespace the same way, so that inputcheckrejects, such as a trailing newline or a tab, no longer gives a different seed from the same words; if you would rather not change that, drop theto_seedline and the last assertion oftest_whitespace_runs. All 288 vectors are unchanged, no inputcheckaccepts today gets a different seed, and the passphrase is not touched. Locally the tests pass on Python 3.8 to 3.14 and PyPy 3.11, and black, isort, flake8 and pyright report nothing.Written with LLM, a computer and the Internet; the mistakes, as usual, are all mine.