Skip to content
View hungho77's full-sized avatar

Block or report hungho77

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hungho77/README.md
Portfolio Engineering notes LinkedIn

I am Ho Thinh Hung, a Senior Edge AI Engineer working where multimodal models meet physical machines. I optimize Vision-Language-Action and Edge LLM systems, build compact VLA models, and turn research ideas into measurable deployment trade-offs.

model quality  ×  memory  ×  latency  ×  power  →  useful intelligence at the edge

What I am building

  • Small VLA models for responsive, resource-constrained robotic systems.
  • Inference optimization across BF16, FP8, NVFP4, INT8, and INT4/AWQ.
  • Edge deployment paths using TensorRT, TensorRT Edge-LLM, ONNX, CUDA, and Jetson.
  • Reproducible benchmarks that expose the real memory, latency, throughput, and accuracy trade-offs.
  • Learning in public through experiment reports, paper notes, and practical implementation guides.

Engineering proof points

Signal What it represents
100K+ model downloads Quantized models adopted through a company Hugging Face organization.
TensorRT Edge-LLM investigations Isolated numerical and export failures with controlled layer-by-layer experiments.
NVIDIA maintainer confirmation Findings and fix direction acknowledged in the upstream repository.
Production AI systems Experience spanning GPU inference, real-time multimodal pipelines, and robotics.

Read the full experiment: TensorRT Edge-LLM — four fixes from controlled experiments
Upstream evidence: issue #151 · issue #105
Feature contribution: PR #193 — InternVLA-N1-DualVLN support — submitted model support for vision-language navigation in TensorRT Edge-LLM.

Selected work

Project Focus
vla.cpp A unified C++ inference runtime for vision-language-action models across heterogeneous hardware.
FoldQuantVLA Native low-bit VLA inference through consistent calibration, weight rounding, and INT4/INT8 execution, without policy retraining.
Model Quantization Recipes Practical ModelOpt recipes and benchmarks for BF16, FP8, NVFP4, INT8 SmoothQuant, and INT4 AWQ.
Digital Human A real-time AI virtual receptionist combining perception, dialogue, and avatar interaction over WebRTC.
EdgeVLA-TRT Edge VLA/WAM inference debugging and upstream feature development, including InternVLA-N1-DualVLN model support.

Technical toolkit

Python C++ PyTorch CUDA TensorRT ONNX vLLM Jetson Docker Linux

More about the systems I work on
  • VLA & multimodal: vision encoders, language backbones, action heads, policy inference, asynchronous execution.
  • Optimization: quantization, mixed precision, calibration, KV-cache and activation memory, kernel/runtime profiling.
  • Serving: vLLM, Triton Inference Server, streaming APIs, multi-GPU inference, latency and throughput analysis.
  • Perception: DeepStream, YOLO, tracking, face recognition, OCR, and real-time video analytics.
  • Systems: CUDA, WebRTC, Kafka, Redis, Docker, REST, WebSocket, and production observability.

Small models. Big machines.
I care about the numbers between a paper result and a reliable deployed system.

Pinned Loading

  1. cair-vinuni/FoldQuantVLA cair-vinuni/FoldQuantVLA Public

    FoldQuant: offline folding for W4A4 / W8A8 vision-language-action inference on edge GPUs

    Python 23 2

  2. VinRobotics/vla.cpp VinRobotics/vla.cpp Public

    A unified inference runtime for VLA models.

    C++ 222 38

  3. EdgeVLA-TRT EdgeVLA-TRT Public

    Forked from NVIDIA/TensorRT-Edge-LLM

    EdgeVLA-TRT: Vision-Language-Action inference on NVIDIA Jetson, built on NVIDIA TensorRT Edge-LLM

    Python 2

  4. VinRobotics/model-quantization-recipes VinRobotics/model-quantization-recipes Public

    Practical quantization recipes for large language models and speech models, from model preparation through deployment-oriented validation.

    Python 36 7

  5. transformer transformer Public

    Pytorch implement of transformer

    Python 3

  6. Isaac-GR00T Isaac-GR00T Public

    Forked from NVIDIA/Isaac-GR00T

    NVIDIA Isaac GR00T N1.5 - A Foundation Model for Generalist Robots.

    Python 1 1