AlphaAvatar is a self-hostable Omni Personal Assistant framework designed to evolve into an intelligent personal butler β a continuous, personalized, and proactive assistant that can remember, understand, plan, and act on behalf of the user.
It is built around a plugin-based real-time Agent architecture, combining:
- π§ Memory for long-term user, assistant, and tool interaction history
- 𧬠Persona for user understanding, identity continuity, and personalization
- π‘ Reflection for self-improvement and long-term behavioral adaptation
- π Planning for task decomposition, reminders, and future-oriented actions
- βοΈ Behavior for response style, workflow policy, and proactive assistance
- π§° Tools through MCP, RAG, DeepResearch, and external integrations
- π Virtual Character for real-time voice/avatar interaction
β¨ Fully self-hostable and privacy-first β AlphaAvatar can run locally or on your own infrastructure, giving you control over your data, memory, tools, and behavior.
AlphaAvatar follows a layered realtime multimodal architecture:
- ποΈ User & Channels: voice, text, camera, screen, files, and messaging platforms.
- π RTC Adapter: connects LiveKit and other realtime communication backends.
- ποΈ Core Perception: normalizes typed multimodal observations, source and segment lineage, ordered perception events, annotations, timelines, and historical snapshots.
- βοΈ Agent & Runtime: manages sessions, context, semantic addressing, conversation focus, multimodal turn taking, shared inference access, and runtime lifecycle.
- π§© Plugin Ecosystem: adds Memory, Persona, RAG, MCP, Character, and other capabilities.
- π§ Provider & Infrastructure: connects models, embeddings, routing, tracing, and structured output.
- πΎ Storage & Data: stores identity, memory, vectors, traces, artifacts, and media.
- π€ Assistant Outputs: delivers voice, text, avatar responses, tool actions, and status updates.
|
|
|
|
|
|
π‘ AlphaAvatar is not just a chatbot. It is a foundation for building stateful, proactive, multimodal, and self-evolving personal AI assistants.
AlphaAvatar organizes plugins into three layers: shared foundation services, perception and interaction capabilities, and tools for external tasks.
Available means implemented in this source tree, not necessarily enabled by default or API-stable. Planned entries are not yet available as standalone plugins.
Shared components that other plugins use through FoundationRuntime.
| Plugin | Capabilities | Status |
|---|---|---|
| Voice | Provides voice activity detection, speech recognition, and speech synthesis. Shares reusable voice components with Router and other consumers. |
|
| Provider | Will provide shared model access and task execution. Will unify structured responses and streaming output through the Foundation layer. |
|
| Context | Will make context construction and composition independently pluggable. Will expose shared context services to other plugins. |
Provider and Context functionality already exists; their standalone Foundation plugins are planned. Voice is currently the only service attached to FoundationRuntime.
Capabilities for understanding users, maintaining memory, and coordinating interaction.
| Plugin | Capabilities | Status |
|---|---|---|
| Router | Coordinates multimodal routing, addressing, turn taking, and interruption. Combines speech, transcripts, and attention evidence for realtime interaction decisions. |
|
| Memory | Stores durable memories from conversations, tools, and environment observations. Retrieves relevant information to maintain continuity across interactions. |
|
| Persona | Builds user profiles and recognizes speakers and faces. Maintains identity continuity and supplies context for personalization. |
|
| Character | Connects AlphaAvatar to realtime virtual-character interfaces. Synchronizes character presentation with voice-based interaction. |
|
| Status | Publishes progress feedback while AlphaAvatar thinks or uses tools. Keeps intermediate activity visible during longer-running workflows. |
|
| Reflection | Will review memories, behavior, and interaction history. Will derive insights to support ongoing adaptation and improvement. |
|
| Planning | Will turn goals into coordinated, multi-step plans. Will track task progress and support longer-horizon execution. |
|
| Behavior | Will manage response style, workflow policies, and proactive behavior. Will adapt interaction strategies to user preferences and context. |
Integrations for research, document access, and external actions.
| Plugin | Capabilities | Status |
|---|---|---|
| MCP | Discovers, retrieves, and invokes tools exposed by MCP servers. Connects the assistant to external services and real-world operations. |
|
| DeepResearch | Combines web search and content extraction for multi-step research. Builds informed responses from external information. |
|
| RAG | Ingests documents and retrieves relevant knowledge for model responses. Connects uploaded content to retrieval-augmented workflows. |
|
| Sandbox | Will provide isolated execution for external actions. Will support controlled interaction with tools and other agents. |
See Plugin Layout and Boundaries for development details and the Roadmap for planned work.
-
[2026/09] Released AlphaAvatar version 0.6.7: Added AlphaAvatar-owned multimodal turn taking, local Semantic Addressing with per-speaker conversation focus, typed addressing evidence and fusion, exact annotation-driven turn snapshots, full-duplex interruption, and removed LiveKit text turn detection and the legacy acoustic invocation path.
-
[2026/08] Released AlphaAvatar version 0.6.6: Added a unified time-aligned multimodal runtime, adaptive audiovisual ENV Memory, provider-neutral model input, runtime capability awareness, time-based perception retention, participant-scoped timezone context, and date-grouped session storage.
-
[2026/07] Released AlphaAvatar version 0.6.4: Added a transport-agnostic perception runtime with typed multimodal streams, shared timelines, annotated payload views, and online ENV memory extraction from live visual observations.
- Released AlphaAvatar version 0.6.5: Added shared realtime audio perception, the Interaction Router, AlphaAvatar-native VAD and STT, isolated per-runner inference processes, and migrated all AlphaAvatar VDB workloads away from LiveKitβs shared inference executor.
-
[2026/06] Released AlphaAvatar version 0.6.0: Added the Status plugin, sampled visual input support, and status-aware DeepResearch / RAG / MCP tool feedback.
- Released AlphaAvatar version 0.6.1: Added visual identity support for Persona, including face detection, face vector matching, speaker-face identity fusion, and several bug fixes.
- Released AlphaAvatar version 0.6.2: introduced the unified provider layer, nested configuration, provider tracing, and a cleaner session/runtime foundation for future multi-user multimodal memory.
- Released AlphaAvatar version 0.6.3: Added the first graph-aware Memory foundation, including multi-object memory items, session-scoped graph node mentions, LanceDB graph-node retrieval, alias-ready graph lookup, and cleaner session-content memory extraction prompts.
-
[2026/05] Released AlphaAvatar version 0.5.4:
- Added LanceDB-backed MCP tool retrieval, enabling AlphaAvatar to semantically search relevant MCP tools from Agent queries.
- Refactored system prompt and runtime prompt composition, improved Persona runtime state tracking, added temporary-user to real-user identity merging, and improved RAG runtime behavior.
- Released AlphaAvatar version 0.5.5: Fixed the inference runner registration lifecycle for production
startmode, ensuring plugins runners are registered after config parsing and before LiveKit creates the inference executor.
-
[2026/04] Released AlphaAvatar version 0.5.3:
- Added localized Markdown backup for the Memory plugin.
- Added LanceDB as the default local VDB option when Qdrant credentials are not provided.
-
[2026/03] Released AlphaAvatar version 0.5.0:
-
[2026/02] Released AlphaAvatar version 0.4.0:
- Added RAG support through RAG-Anything.
- Optimized the Memory and DeepResearch modules.
- Released AlphaAvatar version 0.4.1: Fixed Persona plugin bugs and added a new MCP plugin.
-
[2026/01] Released AlphaAvatar version 0.3.0:
- Added DeepResearch support through the Tavily API.
- Released AlphaAvatar version 0.3.1: Added tool-call memory extraction during userβassistant interactions.
2025 Release History
-
[2025/12] Released AlphaAvatar version 0.2.0:
- Added AIRI Live2D-based virtual character display.
-
[2025/11] Released AlphaAvatar version 0.1.0:
- Added automatic memory extraction.
- Added automatic user persona extraction and matching.
Install stable AlphaAvatar version from PyPI:
uv venv .my-env --python 3.11
source .my-env/bin/activate
pip install alpha-avatar-agentsInstall latest AlphaAvatar version from GitHub:
git clone --recurse-submodules https://github.com/AlphaAvatar/AlphaAvatar.git
cd AlphaAvatar
uv venv .venv --python 3.11
source .venv/bin/activate
uv sync --all-packagesStart your agent in dev mode to connect it to LiveKit and make it available from anywhere on the internet.
π§© Step 1. Configure Environment Variables
cd AlphaAvatar
# Copy template
cp .env.template .env.devEdit .env.dev and set required environment variables.
β Step 2. Run the Agent
Enabled inference runners resolve their model files during initialization. Missing files are downloaded automatically; no separate download command is required.
Set ALPHAAVATAR_MODEL_CACHE to choose the model cache directory. Set ALPHAAVATAR_MODEL_OFFLINE=1 to require locally cached, validated files. See Model Loading for cache and integrity details.
ENV_FILE=.env.dev alphaavatar dev examples/agent_configs/voice/pipeline_openai_tools_minimal.yaml
# or
ENV_FILE=.env.dev alphaavatar dev examples/agent_configs/mm/pipeline_openai_tools.yamlTo see more supported modes, please refer to the LiveKit doc.
To see more examples, please refer to the Examples README
AlphaAvatar supports multiple Access Channels, allowing different types of users β from end users to developers β to interact with the system.
AlphaAvatar now provides a browser-based realtime demo interface built on LiveKit.
π Try the Web Demo: https://www.alphaavatar.ai/demo
The Web Demo supports:
- ποΈ Real-time voice interaction
- π¬ Text chat with the Avatar
- π· Camera preview and video-ready interaction
- π Agent audio playback
- π Virtual character / avatar stage
- π§ Full plugin support, including Memory, Persona, RAG, MCP, and DeepResearch
- π Browser timezone metadata, enabling AlphaAvatar to understand local login time
The Web Demo is the recommended way to try AlphaAvatar with a full realtime multimodal experience.
Interact with AlphaAvatar directly inside messaging platforms.
Capabilities:
- π¬ Text-based conversation
- π€ Voice message interaction
- π§° Tool invocation via chat interface
π¦ Channel introduction: README
Make sure AlphaAvatar Agent is already running (see Quick Start above).
ENV_FILE=.env.dev sh examples/channels/start_whatsapp.shπ‘ The WhatsApp channel runs as an independent bridge process and connects to the Agent runtime.
A dedicated AlphaAvatar mobile application providing:
- ποΈ Real-time voice communication
- π Live2D / Virtual character visualization
- π§ Persistent memory & persona
Developers can immediately access AlphaAvatar via the LiveKit Playground.
π https://agents-playground.livekit.io/
After starting your AlphaAvatar server:
- Connect to your LiveKit instance
- Configure the Agent name in the Playground (must match
avatar_name, default:Assistant) to enable Explicit Dispatch. - Connect to the agent room
- Start testing real-time interaction
Supported capabilities:
- ποΈ Voice interaction
- π§ Memory extraction
- π RAG retrieval
- π§° MCP tool invocation
- π Virtual character display
π‘ AlphaAvatar is currently developer-first, with a Web Demo available for realtime interaction.
More user-facing web and mobile experiences are under active development.



