You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds a benchmark report for the Jetson Orin NX 16 GB (Super dev kit, JetPack 6.2, MAXN_SUPER). All 12 models that fit ran; OpenVLA-OFT doesn't fit in 16 GB. Every model lands between the Orin Nano Super and the AGX Orin.
It was measured at f7e0f7f rather than c93ca0a. The two differ only by docs, release tooling and OpenVINO-only scripts, so I added a line to the README like the one for the Core Ultra X7. Happy to re-run on another commit if you prefer.
Latency looks right but I am not sure about memory report:
The L4T version probably isn't the cause. The AGX Orin report ran on the same R36.4.3, and its Peak RSS does include the CUDA buffers (5982 MiB for π0). So something else must differ in your setup. Were you running inside a container, or reading RSS from a wrapper process instead of vla-bench itself?
Board RAM doesn't give the footprint either. The report says to subtract the 2083 MB idle reading to get the model's footprint. For several models that leaves less than the GGUF file itself: Octo comes out at 132 MB against a 547 MB file, and π0 at 4760 MB against a 6.49 GB file.
Action items:
Remove the L4T sentence.
Re-measure with VmHWM from /proc/<pid>/status for the vla-bench process. Or you can drop both memory columns and say the footprint should match the Orin Nano and AGX reports.
When memory report is resolved, this PR is good to merge.
Thanks for the review! You're right on both points.
No container. The sweep ran vla-bench directly on the host. It did go through a small Python wrapper per run, but the wrapper read RUSAGE_CHILDREN, so the number is vla-bench's peak, not the wrapper's. So I don't know yet why the CUDA buffers don't show up on this board, and the L4T guess was wrong.
Agreed that tegrastats "RAM used" isn't a valid footprint.
Since VmHWM reads the same peak RSS counter as getrusage, I expect it would give the same numbers. So I went with your second option: I removed the L4T sentence, dropped both memory columns, and noted that the footprint should match the Orin Nano and AGX reports. If I find out what's different on this board, I'll follow up with memory numbers in a separate PR.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a benchmark report for the Jetson Orin NX 16 GB (Super dev kit, JetPack 6.2, MAXN_SUPER). All 12 models that fit ran; OpenVLA-OFT doesn't fit in 16 GB. Every model lands between the Orin Nano Super and the AGX Orin.
It was measured at
f7e0f7frather thanc93ca0a. The two differ only by docs, release tooling and OpenVINO-only scripts, so I added a line to the README like the one for the Core Ultra X7. Happy to re-run on another commit if you prefer.Real-arm numbers from the same board are in #31.