Skip to content

gh-159037: Improve performance of str.join pre-pass - #159038

Open
eendebakpt wants to merge 3 commits into
python:mainfrom
eendebakpt:unicode-join-prepass-v2
Open

eendebakpt wants to merge 3 commits into
python:mainfrom
eendebakpt:unicode-join-prepass-v2

Conversation

@eendebakpt

@eendebakpt eendebakpt commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

The pre-pass of _PyUnicode_JoinArray() did more work per item than needed. In this PR we now ORs the kinds and ANDs the ASCII flags of the items; the maximum character and use_memcpy are derived from these after the loop. One separator is counted per item and the extra one is subtracted after the loop. Tests for separators and items of mixed kinds are added.

@vstinner

Benchmark main PR
sep '', 10 short items 63.6 ns 59.8 ns: 1.06x faster
sep '', 10**3 short items 3.13 us 2.81 us: 1.11x faster
sep '', 10**6 short items 3.12 ms 2.77 ms: 1.13x faster
sep '', 10 long items 92.7 ns 89.8 ns: 1.03x faster
sep '', 10**2 long items 432 ns 398 ns: 1.08x faster
sep '', 10**4 long items 56.4 us 51.8 us: 1.09x faster
sep '.', 10 short items 83.0 ns 80.1 ns: 1.04x faster
sep '.', 10**3 short items 4.79 us 4.39 us: 1.09x faster
sep '.', 10**6 short items 4.58 ms 4.26 ms: 1.07x faster
sep '.', 10 long items 117 ns 114 ns: 1.03x faster
sep '.', 10**2 long items 583 ns 559 ns: 1.04x faster
sep '.', 10**4 long items 60.2 us 57.6 us: 1.05x faster
Geometric mean (ref) 1.07x faster
Benchmark main PR
latin1 sep '', 10 short items 77.9 ns 66.9 ns: 1.17x faster
latin1 sep '', 10**3 short items 3.39 us 2.82 us: 1.20x faster
latin1 sep '.', 10**3 short items 4.98 us 4.16 us: 1.20x faster
latin1 sep '.', 10**2 long items 596 ns 542 ns: 1.10x faster
ucs2 sep '', 10 short items 105 ns 86.9 ns: 1.21x faster
ucs2 sep '', 10**3 short items 5.79 us 5.19 us: 1.11x faster
ucs2 sep '.', 10**3 short items 10.9 us 10.6 us: 1.03x faster
ucs4 sep '', 10**3 short items 5.45 us 4.93 us: 1.10x faster
ucs4 sep '.', 10**3 short items 12.5 us 12.3 us: 1.02x faster
ucs4 sep '.', 10**2 long items 1.56 us 1.50 us: 1.04x faster
latin1 sep, ascii items, 10**3 items 4.72 us 4.69 us: 1.01x faster
ucs2 sep, ascii items, 10 items 125 ns 120 ns: 1.05x faster
mixed ascii+latin1 items, 10**3 items 4.78 us 4.29 us: 1.11x faster
mixed ascii+ucs2 items, 10**3 items 10.5 us 10.2 us: 1.04x faster
mixed latin1+ucs4 items, 10**3 items 6.73 us 5.94 us: 1.13x faster
Geometric mean (34 non-ASCII and ASCII cases) (ref) 1.08x faster

Callgrind confirms the direction: the join function executes 8-18% fewer instructions per call, mostly from dropping the per-item kind comparison.

Benchmark scripts

First table (from the PR 158693):

import pyperf
runner = pyperf.Runner()
for sep in ('', '.'):
    for list_size in ('10', '10**3', '10**6'):
        runner.timeit(
            f'sep {sep!r}, {list_size} short items',
            setup=f"data=['x']*{list_size}; sep={sep!r}",
            stmt="sep.join(data)")

    for list_size in ('10', '10**2', '10**4'):
        runner.timeit(
            f'sep {sep!r}, {list_size} long items',
            setup=f"data=['x' * 100]*{list_size}; sep={sep!r}",
            stmt="sep.join(data)")

Second table:

import pyperf
runner = pyperf.Runner()
KINDS = {'ascii': 'x', 'latin1': '\xe9', 'ucs2': '€', 'ucs4': '\U0001f600'}
for kname, ch in KINDS.items():
    for sep in ('', '.'):
        for list_size in ('10', '10**3'):
            runner.timeit(
                f'{kname} sep {sep!r}, {list_size} short items',
                setup=f"data=[{ch!r}]*{list_size}; sep={sep!r}",
                stmt="sep.join(data)")
        runner.timeit(
            f'{kname} sep {sep!r}, 10**2 long items',
            setup=f"data=[{ch!r} * 100]*10**2; sep={sep!r}",
            stmt="sep.join(data)")
for name, sep, items in (
        ("latin1 sep, ascii items", '\xe9', "['x']"),
        ("ucs2 sep, ascii items", '€', "['x']"),
        ("mixed ascii+latin1 items", '.', "['x', '\\xe9']"),
        ("mixed ascii+ucs2 items", '.', "['x', '\\u20ac']"),
        ("mixed latin1+ucs4 items", '', "['\\xe9', '\\U0001f600']"),
        ):
    for list_size in ('10', '10**3'):
        runner.timeit(
            f'{name}, {list_size} items',
            setup=f"data=({items}*{list_size})[:{list_size}]; sep={sep!r}",
            stmt="sep.join(data)")

eendebakpt and others added 2 commits October 8, 2026 20:45
Compute the maximum character and whether memcpy() can be used from the
OR of the item kinds and the AND of their ASCII flags after the pre-pass,
instead of per item. Count one separator per item and subtract the extra
one after the loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed kinds

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@bedevere-app bedevere-app Bot added the type-feature A feature request or enhancement label Oct 8, 2026
@vstinner

vstinner commented Oct 9, 2026

Copy link
Copy Markdown
Member

Interesting optimization! I will try to review it next week.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting core review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants