feat: Stream request bodies from files, iterables, and responses - #1060
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #1060 +/- ##
==========================================
+ Coverage 95.22% 95.30% +0.08%
==========================================
Files 59 60 +1
Lines 5548 5753 +205
==========================================
+ Hits 5283 5483 +200
- Misses 265 270 +5
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
… or __aiter__ `StreamedRequestBody.is_source` accepted only an `Iterator` or `AsyncIterator`, so a class that yields chunks from a generator method - one with `__iter__` or `__aiter__` but no `__next__` or `__anext__` - fell through to the JSON path and was serialized with `default=str`, sending its repr as the body. The detection now accepts any `Iterable` or `AsyncIterable` that is not a `Collection` or a pydantic model. That keeps every value the client uploads whole or serializes as JSON - `str`, `bytes`, `bytearray`, `memoryview`, `list`, `tuple`, `set`, `dict` and its views, `deque`, `range`, and a model whose `__iter__` walks its fields - out of the streaming path, while admitting the `__iter__`-only and `__aiter__`-only wrappers. An object offering both protocols is iterated synchronously, so it works with both clients. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GnzSJqRQkQRtQVeoDuwYQx
…e to is_streamable
…st-bodies # Conflicts: # src/apify_client/types.py
Pijukatel
left a comment
There was a problem hiding this comment.
I ran some tests with long-lasting streams. Currently, the API will return 408 (treated as not retryable) after around 300s-360s of streaming duration. There is no setting in the client to prevent this as far as I know.
This fails the upload and, in some cases, it can leave the source stream open.
Should there be an update to the API code?
This is not an edge case, as streaming makes a lot of sense for large objects or some incoming data source - so for it to last long is not that unusual.
The least we can do is to document this limit and clearly show it in the logs when it happens, but this limit makes the feature far less useful.
Summary
set_record, therun_inputofstart/call/metamorph, andHttpClient.call(data=...)accept anio.IOBasestream such as an open file, an iterable ofbytesorstrchunks, or a streamedHttpResponse. The client uploads it in chunks without holding it in memory, so one Actor'sOUTPUTrecord can become another Actor's input. The async client also accepts async iterables andaiofiles-style readers.Streaming is experimental. The docs say so, and the first streamed body in a process emits a
UserWarning.Behavior
io.IOBasesource is read in 64 KiB chunks. Any other object with areadmethod is read whole and sent as one chunk.io.IOBasesource is retried, rewound before each attempt. Other sources get a single attempt.set_recordvalue used to be compressed, so uploading from a file handle now sends more bytes;file.read()keeps the old behavior.send_requestof a custom transport receivesbytes | Iterator[bytes] | None(async:AsyncIterator[bytes]). It only sees an iterator when a caller passes a streamable value, so existing transports keep working.Docs and tests
A "Streaming uploads" section on the streaming concept page, plus updates to the compression and HTTP client pages, the custom transport examples, and the README. Unit tests cover
StreamedRequestBody, the retry pipeline, andset_recordover Impit and HTTPX2; two integration tests cover a 3 MiB upload and a record piped between stores.Issues
✍️ Drafted by Claude Code