Skip to content

Add Jetson Orin NX 16 GB benchmark report - #36

Merged
khanhnd61-vr merged 2 commits into
VinRobotics:mainfrom
ravediamond:docs/benchmark-jetson-orin-nx
Oct 7, 2026
Merged

khanhnd61-vr merged 2 commits into
VinRobotics:mainfrom
ravediamond:docs/benchmark-jetson-orin-nx

Conversation

@ravediamond

@ravediamond ravediamond commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Adds a benchmark report for the Jetson Orin NX 16 GB (Super dev kit, JetPack 6.2, MAXN_SUPER). All 12 models that fit ran; OpenVLA-OFT doesn't fit in 16 GB. Every model lands between the Orin Nano Super and the AGX Orin.

It was measured at f7e0f7f rather than c93ca0a. The two differ only by docs, release tooling and OpenVINO-only scripts, so I added a line to the README like the one for the Core Ultra X7. Happy to re-run on another commit if you prefer.

Real-arm numbers from the same board are in #31.

@khanhnd61-vr

Copy link
Copy Markdown
Collaborator

Thanks for the benchmark!

Latency looks right but I am not sure about memory report:

  • The L4T version probably isn't the cause. The AGX Orin report ran on the same R36.4.3, and its Peak RSS does include the CUDA buffers (5982 MiB for π0). So something else must differ in your setup. Were you running inside a container, or reading RSS from a wrapper process instead of vla-bench itself?
  • Board RAM doesn't give the footprint either. The report says to subtract the 2083 MB idle reading to get the model's footprint. For several models that leaves less than the GGUF file itself: Octo comes out at 132 MB against a 547 MB file, and π0 at 4760 MB against a 6.49 GB file.

Action items:

  • Remove the L4T sentence.
  • Re-measure with VmHWM from /proc/<pid>/status for the vla-bench process. Or you can drop both memory columns and say the footprint should match the Orin Nano and AGX reports.

When memory report is resolved, this PR is good to merge.

Thanks again!

@ravediamond

ravediamond commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for the review! You're right on both points.

No container. The sweep ran vla-bench directly on the host. It did go through a small Python wrapper per run, but the wrapper read RUSAGE_CHILDREN, so the number is vla-bench's peak, not the wrapper's. So I don't know yet why the CUDA buffers don't show up on this board, and the L4T guess was wrong.

Agreed that tegrastats "RAM used" isn't a valid footprint.

Since VmHWM reads the same peak RSS counter as getrusage, I expect it would give the same numbers. So I went with your second option: I removed the L4T sentence, dropped both memory columns, and noted that the footprint should match the Orin Nano and AGX reports. If I find out what's different on this board, I'll follow up with memory numbers in a separate PR.

@khanhnd61-vr
khanhnd61-vr merged commit 1adf078 into VinRobotics:main Oct 7, 2026
12 of 14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants