ncnn is a high-performance neural network inference framework optimized for the mobile platform
-
Updated
Sep 24, 2026 - C++
ncnn is a high-performance neural network inference framework optimized for the mobile platform
SIMD-accelerated distances, dot products, matrix ops, geospatial & geometric kernels for 16 numeric types — from 6-bit floats to 64-bit complex — across x86, Arm, RISC-V, and WASM, with bindings for Python, Rust, C, C++, Swift, JS, and Go 📐
FeatherCNN is a high performance inference engine for convolutional neural networks.
🚀 Fast prime number generator
A modern C++17 glTF 2.0 library focused on speed, correctness, and usability
🚀 Fast prime counting function library
Heterogeneous Run Time version of Caffe. Added heterogeneous capabilities to the Caffe, uses heterogeneous computing infrastructure framework to speed up Deep Learning on Arm-based heterogeneous embedded platform. It also retains all the features of the original Caffe architecture which users deploy their applications seamlessly.
benchmark for embededded-ai deep learning inference engines, such as NCNN / TNN / MNN / TensorFlow Lite etc.
RV: A Unified Region Vectorizer for LLVM
Hardware-accelerated, zero-copy systems file scanner for Node.js. Powered by AVX-512/AVX2/NEON SIMD, kernel mmap, and persistent thread pools for 50+ GB/s throughput with 0 MB V8 heap.
Heterogeneous Run Time version of MXNet. Added heterogeneous capabilities to the MXNet, uses heterogeneous computing infrastructure framework to speed up Deep Learning on Arm-based heterogeneous embedded platform. It also retains all the features of the original MXNet architecture which users deploy their applications seamlessly.
Single Header Quite Fast QOI(Quite OK Image Format) Implementation written in C++20
Heterogeneous Run Time version of TensorFlow. Added heterogeneous capabilities to the TensorFlow, uses heterogeneous computing infrastructure framework to speed up Deep Learning on Arm-based heterogeneous embedded platform. It also retains all the features of the original TensorFlow architecture which users deploy their applications seamlessly.
NZ1 - NanoZip 1: ultra-fast, dependency-free, portable C compression library optimized for embedded and high-performance use. Full docs: https://ferki-git-creator.github.io/nz1-site/
Custom llama.cpp fork with character intelligence engine: control vectors, attention bias, head rescaling, attention temperature, fast weight memory
To associate your repository with the arm-neon topic, visit your repo's landing page and select "manage topics."