Skip to content
View AlphaAvatar's full-sized avatar
🎯
Focus
🎯
Focus

Block or report AlphaAvatar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AlphaAvatar/README.md
AlphaAvatar logo and banner

PRs Welcome GitHub last commit License

GitHub watchers GitHub forks GitHub stars

Learnable, configurable, and pluggable Omni Personal Assistant for everyone

Website Docs Demo Roadmap

Discord Members GitHub Discussions


AlphaAvatar Introduction

AlphaAvatar is a self-hostable Omni Personal Assistant framework designed to evolve into an intelligent personal butler β€” a continuous, personalized, and proactive assistant that can remember, understand, plan, and act on behalf of the user.

It is built around a plugin-based real-time Agent architecture, combining:

  • 🧠 Memory for long-term user, assistant, and tool interaction history
  • 🧬 Persona for user understanding, identity continuity, and personalization
  • πŸ’‘ Reflection for self-improvement and long-term behavioral adaptation
  • πŸ“… Planning for task decomposition, reminders, and future-oriented actions
  • βš™οΈ Behavior for response style, workflow policy, and proactive assistance
  • 🧰 Tools through MCP, RAG, DeepResearch, and external integrations
  • 😊 Virtual Character for real-time voice/avatar interaction

✨ Fully self-hostable and privacy-first β€” AlphaAvatar can run locally or on your own infrastructure, giving you control over your data, memory, tools, and behavior.


Runtime Architecture 🧠

AlphaAvatar Runtime Architecture

AlphaAvatar follows a layered realtime multimodal architecture:

  • πŸŽ™οΈ User & Channels: voice, text, camera, screen, files, and messaging platforms.
  • πŸ”Œ RTC Adapter: connects LiveKit and other realtime communication backends.
  • πŸ‘οΈ Core Perception: normalizes typed multimodal observations, source and segment lineage, ordered perception events, annotations, timelines, and historical snapshots.
  • βš™οΈ Agent & Runtime: manages sessions, context, semantic addressing, conversation focus, multimodal turn taking, shared inference access, and runtime lifecycle.
  • 🧩 Plugin Ecosystem: adds Memory, Persona, RAG, MCP, Character, and other capabilities.
  • 🧠 Provider & Infrastructure: connects models, embeddings, routing, tracing, and structured output.
  • πŸ’Ύ Storage & Data: stores identity, memory, vectors, traces, artifacts, and media.
  • πŸ“€ Assistant Outputs: delivers voice, text, avatar responses, tool actions, and status updates.

What AlphaAvatar Is Designed For

1️⃣ Personal Data & Life Metrics Management

  • πŸ“Š Track and analyze personal metrics such as health, fitness, sleep, and study progress
  • πŸ“ˆ Provide long-term insights and trend analysis
  • 🎯 Suggest improvements based on historical patterns

2️⃣ Knowledge & Notes Management

  • πŸ“– Organize personal notes, documents, and knowledge
  • πŸ” Retrieve relevant information through RAG
  • 🧠 Build a personal knowledge base over time

3️⃣ Task & Event Management

  • πŸ“… Schedule tasks and reminders
  • ⏰ Proactively notify based on context and priority
  • πŸ”„ Break down long-term goals into actionable steps

4️⃣ Autonomous Planning & Execution

  • 🧠 Plan multi-step workflows such as learning plans, projects, and research
  • πŸ”§ Call tools automatically to complete tasks
  • πŸ“Œ Maintain consistency across long time horizons

5️⃣ Personalized Companion & Context Awareness

  • 🧬 Understand user preferences, habits, and personality
  • πŸ’¬ Provide highly personalized responses
  • 🀝 Maintain continuity across conversations and modalities

6️⃣ External World Interaction

  • 🌐 Search, research, and summarize real-world information
  • 🧰 Integrate with tools such as email, databases, APIs, and messaging apps
  • πŸ”— Act as a bridge between user intent and external systems

πŸ’‘ AlphaAvatar is not just a chatbot. It is a foundation for building stateful, proactive, multimodal, and self-evolving personal AI assistants.


AlphaAvatar Plugins

AlphaAvatar organizes plugins into three layers: shared foundation services, perception and interaction capabilities, and tools for external tasks.

Available means implemented in this source tree, not necessarily enabled by default or API-stable. Planned entries are not yet available as standalone plugins.

βš™οΈ Foundation

Shared components that other plugins use through FoundationRuntime.

Plugin Capabilities Status
Voice Provides voice activity detection, speech recognition, and speech synthesis.
Shares reusable voice components with Router and other consumers.
Available
Provider Will provide shared model access and task execution.
Will unify structured responses and streaming output through the Foundation layer.
Planned
Context Will make context construction and composition independently pluggable.
Will expose shared context services to other plugins.
Planned

Provider and Context functionality already exists; their standalone Foundation plugins are planned. Voice is currently the only service attached to FoundationRuntime.

🧠 Perception

Capabilities for understanding users, maintaining memory, and coordinating interaction.

Plugin Capabilities Status
Router Coordinates multimodal routing, addressing, turn taking, and interruption.
Combines speech, transcripts, and attention evidence for realtime interaction decisions.
Available
Memory Stores durable memories from conversations, tools, and environment observations.
Retrieves relevant information to maintain continuity across interactions.
Available
Persona Builds user profiles and recognizes speakers and faces.
Maintains identity continuity and supplies context for personalization.
Available
Character Connects AlphaAvatar to realtime virtual-character interfaces.
Synchronizes character presentation with voice-based interaction.
Available
Status Publishes progress feedback while AlphaAvatar thinks or uses tools.
Keeps intermediate activity visible during longer-running workflows.
Available
Reflection Will review memories, behavior, and interaction history.
Will derive insights to support ongoing adaptation and improvement.
Planned
Planning Will turn goals into coordinated, multi-step plans.
Will track task progress and support longer-horizon execution.
Planned
Behavior Will manage response style, workflow policies, and proactive behavior.
Will adapt interaction strategies to user preferences and context.
Planned

🧰 Tools

Integrations for research, document access, and external actions.

Plugin Capabilities Status
MCP Discovers, retrieves, and invokes tools exposed by MCP servers.
Connects the assistant to external services and real-world operations.
Available
DeepResearch Combines web search and content extraction for multi-step research.
Builds informed responses from external information.
Available
RAG Ingests documents and retrieves relevant knowledge for model responses.
Connects uploaded content to retrieval-augmented workflows.
Available
Sandbox Will provide isolated execution for external actions.
Will support controlled interaction with tools and other agents.
Planned

See Plugin Layout and Boundaries for development details and the Roadmap for planned work.


Latest News πŸ”₯

  • [2026/09] Released AlphaAvatar version 0.6.7: Added AlphaAvatar-owned multimodal turn taking, local Semantic Addressing with per-speaker conversation focus, typed addressing evidence and fusion, exact annotation-driven turn snapshots, full-duplex interruption, and removed LiveKit text turn detection and the legacy acoustic invocation path.

  • [2026/08] Released AlphaAvatar version 0.6.6: Added a unified time-aligned multimodal runtime, adaptive audiovisual ENV Memory, provider-neutral model input, runtime capability awareness, time-based perception retention, participant-scoped timezone context, and date-grouped session storage.

  • [2026/07] Released AlphaAvatar version 0.6.4: Added a transport-agnostic perception runtime with typed multimodal streams, shared timelines, annotated payload views, and online ENV memory extraction from live visual observations.

    • Released AlphaAvatar version 0.6.5: Added shared realtime audio perception, the Interaction Router, AlphaAvatar-native VAD and STT, isolated per-runner inference processes, and migrated all AlphaAvatar VDB workloads away from LiveKit’s shared inference executor.
  • [2026/06] Released AlphaAvatar version 0.6.0: Added the Status plugin, sampled visual input support, and status-aware DeepResearch / RAG / MCP tool feedback.

    • Released AlphaAvatar version 0.6.1: Added visual identity support for Persona, including face detection, face vector matching, speaker-face identity fusion, and several bug fixes.
    • Released AlphaAvatar version 0.6.2: introduced the unified provider layer, nested configuration, provider tracing, and a cleaner session/runtime foundation for future multi-user multimodal memory.
    • Released AlphaAvatar version 0.6.3: Added the first graph-aware Memory foundation, including multi-object memory items, session-scoped graph node mentions, LanceDB graph-node retrieval, alias-ready graph lookup, and cleaner session-content memory extraction prompts.
  • [2026/05] Released AlphaAvatar version 0.5.4:

    • Added LanceDB-backed MCP tool retrieval, enabling AlphaAvatar to semantically search relevant MCP tools from Agent queries.
    • Refactored system prompt and runtime prompt composition, improved Persona runtime state tracking, added temporary-user to real-user identity merging, and improved RAG runtime behavior.
    • Released AlphaAvatar version 0.5.5: Fixed the inference runner registration lifecycle for production start mode, ensuring plugins runners are registered after config parsing and before LiveKit creates the inference executor.
  • [2026/04] Released AlphaAvatar version 0.5.3:

    • Added localized Markdown backup for the Memory plugin.
    • Added LanceDB as the default local VDB option when Qdrant credentials are not provided.
  • [2026/03] Released AlphaAvatar version 0.5.0:

    • Added the MCP plugin, enabling retrieval and concurrent invocation of MCP tools.
    • Released AlphaAvatar version 0.5.1: Added WhatsApp channel support via Baileys.
    • Released AlphaAvatar version 0.5.2: Added the AlphaAvatar Voice plugin with Voice.ai TTS support.
  • [2026/02] Released AlphaAvatar version 0.4.0:

    • Added RAG support through RAG-Anything.
    • Optimized the Memory and DeepResearch modules.
    • Released AlphaAvatar version 0.4.1: Fixed Persona plugin bugs and added a new MCP plugin.
  • [2026/01] Released AlphaAvatar version 0.3.0:

    • Added DeepResearch support through the Tavily API.
    • Released AlphaAvatar version 0.3.1: Added tool-call memory extraction during user–assistant interactions.
2025 Release History
  • [2025/12] Released AlphaAvatar version 0.2.0:

    • Added AIRI Live2D-based virtual character display.
  • [2025/11] Released AlphaAvatar version 0.1.0:

    • Added automatic memory extraction.
    • Added automatic user persona extraction and matching.

Installation βš™οΈ

Install stable AlphaAvatar version from PyPI:

uv venv .my-env --python 3.11
source .my-env/bin/activate
pip install alpha-avatar-agents

Install latest AlphaAvatar version from GitHub:

git clone --recurse-submodules https://github.com/AlphaAvatar/AlphaAvatar.git
cd AlphaAvatar

uv venv .venv --python 3.11
source .venv/bin/activate
uv sync --all-packages

Quick Start ⚑️

Start your agent in dev mode to connect it to LiveKit and make it available from anywhere on the internet.


🧩 Step 1. Configure Environment Variables

cd AlphaAvatar

# Copy template
cp .env.template .env.dev

Edit .env.dev and set required environment variables.

βœ… Step 2. Run the Agent

Enabled inference runners resolve their model files during initialization. Missing files are downloaded automatically; no separate download command is required.

Set ALPHAAVATAR_MODEL_CACHE to choose the model cache directory. Set ALPHAAVATAR_MODEL_OFFLINE=1 to require locally cached, validated files. See Model Loading for cache and integrity details.

ENV_FILE=.env.dev alphaavatar dev examples/agent_configs/voice/pipeline_openai_tools_minimal.yaml
# or
ENV_FILE=.env.dev alphaavatar dev examples/agent_configs/mm/pipeline_openai_tools.yaml

To see more supported modes, please refer to the LiveKit doc.

To see more examples, please refer to the Examples README


Usage πŸš€

AlphaAvatar supports multiple Access Channels, allowing different types of users β€” from end users to developers β€” to interact with the system.


🌐 Web Access

AlphaAvatar now provides a browser-based realtime demo interface built on LiveKit.

πŸ‘‰ Try the Web Demo: https://www.alphaavatar.ai/demo

The Web Demo supports:

  • πŸŽ™οΈ Real-time voice interaction
  • πŸ’¬ Text chat with the Avatar
  • πŸ“· Camera preview and video-ready interaction
  • πŸ”Š Agent audio playback
  • 😊 Virtual character / avatar stage
  • 🧠 Full plugin support, including Memory, Persona, RAG, MCP, and DeepResearch
  • 🌍 Browser timezone metadata, enabling AlphaAvatar to understand local login time

AlphaAvatar Web Demo Screenshot

The Web Demo is the recommended way to try AlphaAvatar with a full realtime multimodal experience.


πŸ’¬ Social & Messaging Platforms

Interact with AlphaAvatar directly inside messaging platforms.

Capabilities:

  • πŸ’¬ Text-based conversation
  • 🎀 Voice message interaction
  • 🧰 Tool invocation via chat interface

WhatsApp

πŸ“¦ Channel introduction: README

▢️ Start WhatsApp Channel

Make sure AlphaAvatar Agent is already running (see Quick Start above).

ENV_FILE=.env.dev sh examples/channels/start_whatsapp.sh

πŸ’‘ The WhatsApp channel runs as an independent bridge process and connects to the Agent runtime.

WeChat

Slack


πŸ“² Native Mobile App

A dedicated AlphaAvatar mobile application providing:

  • πŸŽ™οΈ Real-time voice communication
  • 😊 Live2D / Virtual character visualization
  • 🧠 Persistent memory & persona

πŸ§ͺ Developer Playground

Developers can immediately access AlphaAvatar via the LiveKit Playground.

πŸ‘‰ https://agents-playground.livekit.io/

After starting your AlphaAvatar server:

  1. Connect to your LiveKit instance
  2. Configure the Agent name in the Playground (must match avatar_name, default: Assistant) to enable Explicit Dispatch.
  3. Connect to the agent room
  4. Start testing real-time interaction

Supported capabilities:

  • πŸŽ™οΈ Voice interaction
  • 🧠 Memory extraction
  • πŸ” RAG retrieval
  • 🧰 MCP tool invocation
  • 😊 Virtual character display

playground airi screenshot


πŸ’‘ AlphaAvatar is currently developer-first, with a Web Demo available for realtime interaction.

More user-facing web and mobile experiences are under active development.

Pinned Loading

  1. AlphaAvatar AlphaAvatar Public

    An Omni Real-Time Assistant for Everyone

    Python 816 63

  2. AlphaAvatar-web AlphaAvatar-web Public

    Official website for AlphaAvatar.

    TypeScript 3

  3. AlphaAvatar-zero AlphaAvatar-zero Public

    3

  4. AlphaAvatar-distill AlphaAvatar-distill Public

    Agentic distillation for transforming frontier-scale models into real-time, edge-device models.

    Python 5