LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
-
Updated
May 19, 2025 - Python
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
A Survey of Spoken Dialogue Models (60 pages)
AAAI 2025: Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
[ICCV 2025] Official Implementation for "Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition"
SlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"
Code for DeSTA2.5-Audio, general-purpose LALM
A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.
Code and model for ICASSP 2025 Paper "Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data"
Streamable Text-to-Speech model using a language modeling approach, without vector quantization
Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications
A high-performance, universal serving framework for any-to-any models.
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
Survey of audio language models
The official code for the SALMon🍣 benchmark (ICASSP 2025 - Oral)
[IEEE OJSP'26, IEEE SLT'24] "Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization"
A collections of audio codecs with a standardized API
a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.
[EMNLP 2026] FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates
End-to-end speech model research: waveform encoding, causal decoding, codec tokens, joint losses and streaming.
To associate your repository with the speech-language-model topic, visit your repo's landing page and select "manage topics."