Skip to content

gh-87613: Use AC vectorcall for bytes and bytearray - #158891

Open
cmaloney wants to merge 1 commit into
python:mainfrom
cmaloney:ac_vc_bytes_take
Open

cmaloney wants to merge 1 commit into
python:mainfrom
cmaloney:ac_vc_bytes_take

Conversation

@cmaloney

@cmaloney cmaloney commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Speeds up common construction of both objects. Enables optimizing construction off uniquely referenced temporary objects.

Previously the argument tuples / dictionaries required by the call protocol resulted in arguments never being unique temporaries. With vectorcall they can be. Avoid copying by adopting the existing allocation when possible.

Benchmark

Generally faster. The large speedups are avoiding copies, just changing call protocol gives moderate speedup.

Run on my linux dev machine: PGO+LTO, clang 23, --with-tail-call-interp
(note: AI Written, but I think sufficiently representative)

"""bytes() and bytearray() construction.

temp/ passes a unique temporary (adopted, no copy); named/ binds the same
source to a name first so it is always copied.
"""
import pyperf

runner = pyperf.Runner()

runner.timeit('call/bytes()', 'bytes()')
runner.timeit('call/bytes(int)', 'bytes(16)')
runner.timeit('call/bytes(bytes)', "bytes(b'abcdefghabcdefgh')")
runner.timeit("call/bytes(str, 'ascii')", "bytes('abcdefgh', 'ascii')")
runner.timeit('call/bytearray()', 'bytearray()')
runner.timeit('call/bytearray(int)', 'bytearray(16)')
runner.timeit('call/bytearray(bytes)', "bytearray(b'abcdefghabcdefgh')")
runner.timeit("call/bytearray(str, 'ascii')",
              "bytearray('abcdefgh', 'ascii')")

for size_name, size in (('16B', 16), ('4KiB', 4096), ('1MiB', 1 << 20)):
    # Slicing in the statement makes a fresh unique temporary.
    setup = f"b = b'x' * {size}; ba = bytearray(b)"
    runner.timeit(f'temp/bytes(bytearray) {size_name}',
                  'bytes(ba[1:])', setup)
    runner.timeit(f'named/bytes(bytearray) {size_name}',
                  'x = ba[1:]; bytes(x); del x', setup)
    runner.timeit(f'temp/bytearray(bytes) {size_name}',
                  'bytearray(b[1:])', setup)
    runner.timeit(f'named/bytearray(bytes) {size_name}',
                  'x = b[1:]; bytearray(x); del x', setup)
    runner.timeit(f'temp/bytearray(bytearray) {size_name}',
                  'bytearray(ba[1:])', setup)
    runner.timeit(f'named/bytearray(bytearray) {size_name}',
                  'x = ba[1:]; bytearray(x); del x', setup)
Benchmark main patch
temp/bytearray(bytearray) 1MiB 515 us 14.9 us: 34.53x faster
temp/bytearray(bytes) 1MiB 508 us 14.8 us: 34.27x faster
temp/bytes(bytearray) 1MiB 507 us 14.9 us: 34.05x faster
temp/bytearray(bytes) 4KiB 144 ns 79.5 ns: 1.81x faster
temp/bytearray(bytearray) 4KiB 155 ns 94.7 ns: 1.63x faster
temp/bytearray(bytes) 16B 56.4 ns 36.1 ns: 1.56x faster
temp/bytes(bytearray) 4KiB 150 ns 96.4 ns: 1.55x faster
temp/bytearray(bytearray) 16B 64.1 ns 47.3 ns: 1.35x faster
call/bytes() 14.7 ns 11.1 ns: 1.32x faster
call/bytes(int) 31.5 ns 24.3 ns: 1.29x faster
call/bytes(str, 'ascii') 36.9 ns 28.7 ns: 1.28x faster
call/bytes(bytes) 34.3 ns 27.3 ns: 1.26x faster
temp/bytes(bytearray) 16B 59.4 ns 49.5 ns: 1.20x faster
call/bytearray(str, 'ascii') 45.6 ns 39.1 ns: 1.17x faster
call/bytearray(bytes) 44.3 ns 38.9 ns: 1.14x faster
call/bytearray(int) 41.7 ns 37.3 ns: 1.12x faster
named/bytearray(bytes) 16B 59.6 ns 54.4 ns: 1.10x faster
named/bytes(bytearray) 16B 61.3 ns 56.2 ns: 1.09x faster
named/bytearray(bytearray) 16B 65.3 ns 60.9 ns: 1.07x faster
named/bytes(bytearray) 4KiB 154 ns 148 ns: 1.04x faster
call/bytearray() 22.1 ns 21.4 ns: 1.03x faster
named/bytearray(bytearray) 4KiB 156 ns 151 ns: 1.03x faster
named/bytearray(bytes) 4KiB 146 ns 143 ns: 1.02x faster
Geometric mean (ref) 1.77x faster

Benchmark hidden because not significant (3): named/bytes(bytearray) 1MiB, named/bytearray(bytes) 1MiB, named/bytearray(bytearray) 1MiB

Speeds up common construction of both objects. Enables optimizing
construction off uniquely referenced temporary objects.

Previously the argument tuples / dictionaries required by the call
protocol resulted in arguments never being unique temporaries. With
vectorcall they can be. Avoid copying by adopting the existing
allocation when possible.

Benchmark
Run on my linux dev machine: PGO+LTO, clang 23, --with-tail-call-interp

(note: AI Written, but I think sufficiently representative)
```python
"""bytes() and bytearray() construction.

temp/ passes a unique temporary (adopted, no copy); named/ binds the same
source to a name first so it is always copied.
"""
import pyperf

runner = pyperf.Runner()

runner.timeit('call/bytes()', 'bytes()')
runner.timeit('call/bytes(int)', 'bytes(16)')
runner.timeit('call/bytes(bytes)', "bytes(b'abcdefghabcdefgh')")
runner.timeit("call/bytes(str, 'ascii')", "bytes('abcdefgh', 'ascii')")
runner.timeit('call/bytearray()', 'bytearray()')
runner.timeit('call/bytearray(int)', 'bytearray(16)')
runner.timeit('call/bytearray(bytes)', "bytearray(b'abcdefghabcdefgh')")
runner.timeit("call/bytearray(str, 'ascii')",
              "bytearray('abcdefgh', 'ascii')")

for size_name, size in (('16B', 16), ('4KiB', 4096), ('1MiB', 1 << 20)):
    # Slicing in the statement makes a fresh unique temporary.
    setup = f"b = b'x' * {size}; ba = bytearray(b)"
    runner.timeit(f'temp/bytes(bytearray) {size_name}',
                  'bytes(ba[1:])', setup)
    runner.timeit(f'named/bytes(bytearray) {size_name}',
                  'x = ba[1:]; bytes(x); del x', setup)
    runner.timeit(f'temp/bytearray(bytes) {size_name}',
                  'bytearray(b[1:])', setup)
    runner.timeit(f'named/bytearray(bytes) {size_name}',
                  'x = b[1:]; bytearray(x); del x', setup)
    runner.timeit(f'temp/bytearray(bytearray) {size_name}',
                  'bytearray(ba[1:])', setup)
    runner.timeit(f'named/bytearray(bytearray) {size_name}',
                  'x = ba[1:]; bytearray(x); del x', setup)
```

| Benchmark                       | main    | patch                  |
|---------------------------------|:-------:|:----------------------:|
| temp/bytearray(bytearray) 1MiB  | 515 us  | 14.9 us: 34.53x faster |
| temp/bytearray(bytes) 1MiB      | 508 us  | 14.8 us: 34.27x faster |
| temp/bytes(bytearray) 1MiB      | 507 us  | 14.9 us: 34.05x faster |
| temp/bytearray(bytes) 4KiB      | 144 ns  | 79.5 ns: 1.81x faster  |
| temp/bytearray(bytearray) 4KiB  | 155 ns  | 94.7 ns: 1.63x faster  |
| temp/bytearray(bytes) 16B       | 56.4 ns | 36.1 ns: 1.56x faster  |
| temp/bytes(bytearray) 4KiB      | 150 ns  | 96.4 ns: 1.55x faster  |
| temp/bytearray(bytearray) 16B   | 64.1 ns | 47.3 ns: 1.35x faster  |
| call/bytes()                    | 14.7 ns | 11.1 ns: 1.32x faster  |
| call/bytes(int)                 | 31.5 ns | 24.3 ns: 1.29x faster  |
| call/bytes(str, 'ascii')        | 36.9 ns | 28.7 ns: 1.28x faster  |
| call/bytes(bytes)               | 34.3 ns | 27.3 ns: 1.26x faster  |
| temp/bytes(bytearray) 16B       | 59.4 ns | 49.5 ns: 1.20x faster  |
| call/bytearray(str, 'ascii')    | 45.6 ns | 39.1 ns: 1.17x faster  |
| call/bytearray(bytes)           | 44.3 ns | 38.9 ns: 1.14x faster  |
| call/bytearray(int)             | 41.7 ns | 37.3 ns: 1.12x faster  |
| named/bytearray(bytes) 16B      | 59.6 ns | 54.4 ns: 1.10x faster  |
| named/bytes(bytearray) 16B      | 61.3 ns | 56.2 ns: 1.09x faster  |
| named/bytearray(bytearray) 16B  | 65.3 ns | 60.9 ns: 1.07x faster  |
| named/bytes(bytearray) 4KiB     | 154 ns  | 148 ns: 1.04x faster   |
| call/bytearray()                | 22.1 ns | 21.4 ns: 1.03x faster  |
| named/bytearray(bytearray) 4KiB | 156 ns  | 151 ns: 1.03x faster   |
| named/bytearray(bytes) 4KiB     | 146 ns  | 143 ns: 1.02x faster   |
| Geometric mean                  | (ref)   | 1.77x faster           |

Benchmark hidden because not significant (3): named/bytes(bytearray) 1MiB,
named/bytearray(bytes) 1MiB, named/bytearray(bytearray) 1MiB
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant