The production-grade agentic metaframework, built on Pydantic AI.
Build with Firefly agents, tools, memory and typed model options — gain lifecycle hooks, delegation, reasoning patterns, validation loops, RAG, and DAG pipelines, all protocol-driven and swappable.
📘 The Tutorial · Quick Start · Models and APIs · Why · Architecture · Features · Docs · Ecosystem · Changelog
Copyright 2026 Firefly Software Foundation · Licensed under the Apache License 2.0
Table of contents
Pydantic AI provides an excellent foundation: type-safe, model-agnostic agents with structured output. But a production agentic system demands far more than a single agent call. You need to orchestrate multi-step reasoning, validate and retry LLM outputs against schemas, manage conversation memory across turns, observe every call with traces and metrics, and run A/B experiments to compare models — all without coupling your domain logic to infrastructure concerns.
fireflyframework-agentic is the production framework built on top of Pydantic AI. It extends the engine with composable layers — from core configuration through agent management, intelligent reasoning, experimentation, and pipeline orchestration — so that every concern has a dedicated, protocol-driven module. You write your business logic; the framework provides the architecture.
What "metaframework" means in practice:
- Use
FireflyAgent,@firefly_tool,MemoryManager, andModelOptionsas the application interface; native Pydantic AI APIs remain available for advanced integrations. - The framework wraps them with lifecycle hooks, registries, delegation routers, reasoning patterns, validation loops, and DAG pipelines — all optional, all composable, all swappable through Python protocols.
- Keep provider selection in configuration and validate options against each model's capabilities. Swap memory backends or other protocol implementations independently.
Your application uses FireflyAgent + Firefly tools + MemoryManager +
ModelOptions. Pydantic AI handles the underlying model calls. Select the provider,
model, and API in configuration; keep the same agent code, tools, and memory API.
| Model selector | API used by Firefly | Compatibility |
|---|---|---|
openai: or openai-chat: |
OpenAI Chat Completions | Existing openai: agents keep their API |
openai-responses: |
OpenAI Responses | Explicit opt-in for Responses features |
azure: or azure-chat: |
Azure Chat Completions | Existing azure: agents keep their API |
azure-responses: |
Azure Responses | Requires a deployment that supports Responses |
| Other supported provider prefixes | The selected provider's API | Same Firefly agent, tool, memory, and options interface |
Set FIREFLY_AGENTIC_DEFAULT_MODEL, or pass model= to an agent. Firefly preserves
the legacy meaning of openai: even though Pydantic AI 2.x changed its default
routing. Without an override, Firefly still uses openai:gpt-4o.
Model controls are abstracted too:
| Firefly option | Purpose |
|---|---|
max_tokens, request_timeout |
Output limit and per-request timeout |
temperature, top_p, top_k, seed, stop_sequences |
Sampling and generation controls where supported |
reasoning or reasoning_budget_tokens |
Reasoning effort or an explicit thinking budget |
parallel_tool_calls |
Parallel tool-call policy |
store_responses, output_verbosity |
Provider response storage and output detail where supported |
Pass ModelOptions to an agent, decorator, ModelSpec, or individual run.
Per-run options inherit agent settings and override explicitly supplied fields;
explicit None clears an inherited option. Firefly translates these controls and
raises ModelOptionsError for unsupported combinations. Portability does not make
every provider support every feature: choose a model with the tools, structured
output, and reasoning capabilities your application needs. Native model_settings
remain available for advanced integrations.
See the model and options reference,
migration guide, and
Responses example. For GPT-6 reasoning with tools,
use Responses and leave FIREFLY_AGENTIC_DEFAULT_TEMPERATURE unset; the
model compatibility guide lists endpoint
constraints.
-
Explicit extension contracts — Components expose Python protocols or abstract base classes; protocols marked
@runtime_checkablealso support runtime checks. Contracts span every layer —AgentLike,ToolProtocol,GuardProtocol,AgentMiddleware,DelegationStrategy,ReasoningPattern,ValidationRule,StepExecutor,Checkpointer,Chunker,MemoryStore,EmbeddingProtocol,VectorStoreProtocol, plus the new workflow ports (AgentRunner,JournalBackend,ModelSelectionStrategy) and more — so you can swap or extend any component without modifying framework internals. -
Convention over configuration — Sensible defaults everywhere.
FireflyAgenticConfigis a Pydantic Settings singleton that reads from environment variables prefixed withFIREFLY_AGENTIC_and.envfiles. One config object governs model defaults, retry counts, token limits, telemetry emission (observability_enabled), strict-cost mode (cost_strict), memory backends, and validation thresholds — override only what you need. -
Layered composition — Core, agent, intelligence, experimentation, and orchestration group related responsibilities. Applications compose these modules through explicit contracts; the architecture diagram is a conceptual map, not an enforced Python import hierarchy.
-
Optional dependencies — Storage and numerical libraries (
chromadb,pinecone,asyncpg,numpy) are declared as pip extras ([vectorstores-chroma],[postgres],[embeddings],[all]). The framework imports them lazily so that you install only what your deployment requires. Agent provider integrations are declared explicitly with the Pydantic AI dependency.
Follow the visual guide for a reading path through the diagrams and their implementation guides. Open any diagram at full size to inspect its labels.
Implement a module's protocol or abstract base class to create your own components. The diagram distinguishes structural protocols from inheritance-based extension points.
-
Models and options —
ModelOptionstranslates typed request controls for the selected model and API.ModelSpec,Credential, andModelFactorysupport application-managed model catalogues and credentials. Explicit Chat Completions and Responses routing preserves existing agents while enabling newer APIs; unsupported option combinations raiseModelOptionsError. -
Agents —
FireflyAgentwrapspydantic_ai.Agentwith metadata, lifecycle hooks, and automatic registration.AgentRegistryprovides singleton name-based discovery.DelegationRouterroutes prompts across agent pools via seven strategies (RoundRobinStrategy,CapabilityStrategy,ContentBasedStrategy,CostAwareStrategy,ChainStrategy,FallbackStrategy,WeightedStrategy). A composable middleware stack (MiddlewareChainoverAgentMiddleware) wraps every run — withdefault_middleware=True,LoggingMiddlewareis installed andObservabilityMiddlewareis added whenobservability_enabled, withPromptGuardMiddleware,OutputGuardMiddleware,CostGuardMiddleware,CacheMiddleware,PromptCacheMiddleware,ExplainabilityMiddleware,ValidationMiddleware,RetryMiddleware, andCircuitBreakerMiddlewareavailable to add.FallbackModelWrapper/run_with_fallbackprovide automatic model failover, andResultCache/CacheStatisticsback response caching. The@firefly_agentdecorator defines an agent in one statement. Five template factories (create_summarizer_agent,create_classifier_agent,create_extractor_agent,create_conversational_agent,create_router_agent) cover common use cases out of the box.
-
Tools —
ToolProtocol(duck-typed) andBaseTool(inheritance) let you choose your extensibility style.ToolBuilderprovides a fluent API for building tools without subclassing. Four guard types (ValidationGuard,RateLimitGuard,SandboxGuard,CompositeGuard) intercept calls before execution (a rejected guard raisesToolGuardError). For human-in-the-loop, mark a toolrequires_approval=True: the agent run pauses before executing it and returns aDeferredToolRequests(detected viais_deferred(result)), which you resume withdeferred_tool_results=— approving (ToolApproved), denying (ToolDenied), or auto-deciding inline via anapproval_handler=. The native deferred-tools types are re-exported fromfireflyframework_agentic.tools. Three composition patterns (SequentialComposer,FallbackComposer,ConditionalComposer) build higher-order tools.ToolKitgroups tools for bulk registration. Nine built-in tools (calculator, datetime, filesystem, HTTP, JSON, search, shell, text, database) are ready to attach to any agent. -
Prompts —
PromptTemplaterenders Jinja2 templates with variable validation and token estimation.PromptRegistrymaps names to versioned templates. Three composers (SequentialComposer,ConditionalComposer,MergeComposer) combine templates at render time.PromptValidatorenforces token limits and required sections.PromptLoaderloads templates from strings, files, or entire directories. -
Reasoning — Six pluggable patterns share the
ReasoningPatterninterface and structured results. The base loop is_reason→_act→_observe→_should_continue; Tree of Thoughts and Goal Decomposition override execution: ReAct (observe-think-act), Chain of Thought (step-by-step), Plan-and-Execute (goal → plan → steps with optional replanning), Reflexion (execute → critique → retry), Tree of Thoughts (branch → evaluate → select), and Goal Decomposition (goal → phases → tasks). All produce structuredReasoningResultwithReasoningTrace. Prompts are slot-overridable. Each pattern's structured output is wrapped in a pydantic-ai output mode — selected per-pattern viaoutput_mode=or framework-wide via thereasoning_output_modeconfig —"tool"(ToolOutput),"native"(provider structured output), or"prompted"(PromptedOutput, portable to any model).OutputReviewercan validate final outputs.ReasoningPipelinechains patterns sequentially.
-
Content —
TextChunkersplits by tokens, sentences, or paragraphs with configurable overlap;MarkdownChunkerchunks structure-aware on Markdown headings.DocumentSplitterdetects page breaks and section separators.ImageTilercomputes tile coordinates for VLM processing.BatchProcessorruns chunks through an agent concurrently with a semaphore.ContextCompressordelegates to pluggable strategies (TruncationStrategy,SummarizationStrategy,MapReduceStrategy) —ContextCompressor.compressis async.SlidingWindowManagermaintains a rolling token-budgeted context window. The[binary]-gatedcontent.binarysubmodule normalises uploaded files into consumer-ready artifacts:BinaryNormalizer(withBinaryConfig) producesBinaryArtifacts,sniff_media_typedetects formats,build_office_converterselects anOfficeConverter(GotenbergConverter,LibreOfficeConverter,NoOpOfficeConverter), andPdfGuard,ImageNormalizer,ArchiveUnpacker, andEmailUnpackerhandle PDFs, images, archives, and emails. -
Memory —
ConversationMemorystores per-conversation turn history with token-budget enforcement (selecting the newest turns that fit).WorkingMemoryprovides a scoped key-value scratchpad backed byMemoryStore(InMemoryStore,FileStore,SQLiteStore,PostgreSQLStore, orMongoDBStore).MemoryManagercomposes both behind a unified API and supportsfork()for isolating working memory in delegated agents or pipeline branches while sharing conversation context.create_llm_summarizerbuilds an LLM-backed history summarizer for long conversations. Completed streaming turns retain provider metadata and reasoning state. Close database-backed managers withclose()oraclose()when their owning application shuts down. -
Validation — Five composable rules (
RegexRule,FormatRule,RangeRule,EnumRule,CustomRule) feed intoFieldValidatorandOutputValidator.OutputReviewerwraps agent calls with parse-then-validate retry logic: on failure it builds a feedback prompt and retries up to N times.RubricRevieweradds LLM-as-judge grading against a rubric (RubricReviewer.from_rubric_file). QoS guards (ConfidenceScorer,ConsistencyChecker,GroundingChecker, plus theQoSGuardaggregator returning aQoSResult) detect hallucinations and low-quality extractions before they propagate downstream. -
Pipeline —
DAGholdsDAGNodeandDAGEdgeobjects with cycle detection and topological sort.PipelineEngineschedules nodes as their dependencies complete, with per-node condition gates, retries, and timeouts.PipelineBuilderoffers a fluent API (add_node/add_edge/chain). Step types adapt agents, patterns, and functions to DAG nodes:AgentStep,ReasoningStep,CallableStep,FanInStep,BatchLLMStep,EmbeddingStep, andRetrievalStep. State-based routing uses.branch()andSend; legacyBranchStepandFanOutStepare deprecated. State reducers (append,extend,merge_dict,replace) merge fan-out results, and control signals (Pause,Send) drive branching and human-in-the-loop pauses.Checkpointer/FileCheckpointer(withCheckpointRecord) persist and resume long runs, and a pluggable audit-log family (AuditLog,FileAuditLog,LoggingAuditLog,OtelAuditLog,QueryableAuditLogoverAuditEntry) records execution traces.
- Workflows —
@workflowdefines a code-defined orchestration DSL;await subworkflow(...)composes child workflows over your agents — a complement to the declarativepipelineDAG, both living in the Orchestration layer. Compose async primitives —agent(),parallel(),pipeline(),stream(),phase(),human()(human-in-the-loop),map_agents()andlog()— inside aWorkflowContextthat carries aWorkflowBudget(concurrency, agent-count and token/cost ceilings), aJournal(JournalBackend/FileJournalBackend) for replaying completed operations, and a pluggableAgentRunner.FireflyAgentRunner(the default) runs every sub-agent call through a fullFireflyAgent(middleware, guards, budget, model fallback);DefaultAgentRunneris the lightweight path.SmartRoutingRunnerselects a model through its configuredModelSelectionStrategy(ComplexityHeuristicStrategy,CostFloorStrategy), and verification helpers (cascade,adversarial_verify,judge_panel,loop_until_dry) add quality checks. Replay requires stable control flow and call order; application side effects need idempotency. See docs/workflows.md.
-
Observability —
FireflyTracercreates OpenTelemetry spans scoped to agents, tools, and reasoning steps.FireflyMetricsrecords tokens (total, prompt, completion), latency, cost, errors, and reasoning depth via the OTel metrics API.FireflyEventsemits structured log records.@tracedand@metereddecorators instrument any function with one line. Native pydantic-ai instrumentation is on by default — rich GenAI-convention spans (and metrics) per model request and tool call, nested under the framework's agent span, with prompt/response content stripped by default for privacy (toggle vianative_instrumentation_enabled). The framework emits model and agent telemetry purely through the OpenTelemetry API; the host application owns OTel SDK and exporter configuration.UsageTrackerautomatically records token usage, cost estimates, and latency for every agent run, reasoning step, and pipeline execution. Cost is computed through a resolver chain (resolve_cost,genai_prices_cost,provider_reported_cost,DEFAULT_RESOLVERS); setcost_strictto raiseUnknownModelCostErrorwhen no price is found.BudgetGateenforces token/cost budgets per scope (BudgetRule,BudgetMode,BudgetWindow), and a pluggable sink family (LoggingSink,JSONLFileSink,OTelMetricsSink,EventBusSink,CostSink) routes usage records wherever you need them. -
Explainability —
TraceRecordercaptures every LLM call, tool invocation, and reasoning step asDecisionRecordobjects.ExplanationGeneratorturns records into human-readable narratives.AuditTrailprovides an append-only, immutable log with JSON export for compliance.ReportBuilderproduces Markdown and JSON reports with statistics. -
Security —
PromptGuardscans inbound prompts for injection and jailbreak patterns;OutputGuardredacts secrets and PII from model output (default_prompt_guard/default_output_guardprovide ready-to-use instances). At-rest protection comes fromAESEncryptionProvider(behind theEncryptionProviderprotocol) andEncryptedMemoryStore, which encryptsMemoryEntry.contentwhile leaving keys, metadata, and timestamps in plaintext. Inbound request authentication and authorization are a hosting concern, not the framework's. -
Resilience —
CircuitBreakertrips after a configurable failure threshold and rejects calls withCircuitBreakerOpenErrorwhile open, transitioning throughCircuitState(closed → open → half-open).CircuitBreakerMiddlewareplugs it into the agent middleware chain so a failing model is short-circuited before it drains your budget. -
Storage —
StorageBackendabstracts blob/object storage withLocalBackendout of the box;DatabaseStorepersists artifacts with leasing (WriteSession,LockToken), a configurableRetryPolicy, andStorageMetadata. Typed errors (StorageUploadError,StorageDownloadError,StorageLeaseError,StorageTransientError,StoreUnavailableError) make failure handling explicit. -
Experiments —
Experimentdefines variants with model, temperature, and prompt overrides.ExperimentRunnerexecutes all variants against a dataset via anagent_factorycallable.ExperimentTrackerpersists results with optional JSON export.VariantComparatorcomputes latency, output length, and comparison summaries. -
Lab —
LabSessionmanages interactive agent sessions with history.Benchmarkruns agents against standardised inputs and reports p95 latency.EvalOrchestratorscores agent outputs with pluggableScorerfunctions.EvalDatasetloads/saves test cases from JSON.ModelComparisonruns the same prompts across multiple agents for side-by-side analysis. -
Evaluation — LLM-as-judge metrics (faithfulness, relevancy, answer correctness, RAGAS, …) and deterministic retrieval metrics (recall@k, MRR, MAP, nDCG, …) for assessing LLM and pipeline outputs. Each metric is a plain function you call directly. Install with
pip install "fireflyframework-agentic[evaluation]". See docs/evaluation.md for the full guide.Optional developer tooling.
fireflyframework_agentic.experiments(A/B experiments) andfireflyframework_agentic.lab(offline evaluation / benchmarking) are leaf modules — nothing in the core imports them and they add no third-party dependencies. Import them only if you run experiments or evaluations; agent-building consumers can ignore them. -
Embeddings —
EmbeddingProtocol(duck-typed) andBaseEmbedder(inheritance with auto-batching) provide provider-agnostic text embedding. Eight providers ship out of the box: OpenAI, Azure OpenAI, Cohere, Google, Mistral, Voyage AI, AWS Bedrock, and Ollama (local).EmbedderRegistrymanages named instances. Built-in similarity utilities (cosine_similarity,euclidean_distance,dot_product) compare vectors without external dependencies. Configuration viaembedding_batch_size,embedding_max_retries, anddefault_embedding_model. -
Vector Stores —
VectorStoreProtocolandBaseVectorStoreprovide pluggable storage and retrieval with six backends: InMemoryVectorStore (zero-dependency, brute-force cosine), ChromaVectorStore, PineconeVectorStore, QdrantVectorStore, PgVectorVectorStore (Postgres + pgvector), and SqliteVecVectorStore (embedded sqlite-vec). A multi-tenant isolation layer (ScopedVectorStore,TenantScopedVectorStore, plusscope_namespace/parse_scope_namespacehelpers) namespaces documents per tenant or scope. Auto-embedding upserts documents without pre-computed vectors.search_textembeds a query string and searches in one call, andSearchFilternarrows results by metadata. Namespace scoping isolates document collections.VectorStoreRegistrymanages named instances.EmbeddingStepandRetrievalStepintegrate directly into DAG pipelines for retrieval-augmented workflows.
- Studio — moved to its own repository:
fireflyframework-agentic-studio.
A browser-based visual IDE for building agent pipelines (drag-and-drop
canvas, code generation, AI assistant, time-travel debugging). Install
with
pip install fireflyframework-agentic-studioand launch withfirefly studio.
Firefly Agentic is the agentic member of the Firefly Framework — a polyglot platform that brings one cohesive programming model to many runtimes. Each member shares the same firefly-in-the-dark identity, recolored per language.
- PyFly — the Python implementation (Spring-Boot DX, async-native).
- Firefly for Rust — reactive,
tokio+axummicroservices. - Firefly Studio — a browser-based visual IDE for building agent pipelines.
Runtime:
Core dependencies (installed automatically):
- pydantic-ai
>=2.51.0,<3— Agent engine (model calls, tool dispatch, streaming), with provider extras declared explicitly - pydantic
>=2.13,<3— Data validation and settings - pydantic-settings
>=2.14.2,<3— Environment-based configuration - Jinja2
>=3.1.0— Prompt template engine - httpx
>=0.28.0— Async HTTP client (built-in HTTP tool, Gotenberg converter) - OpenTelemetry API
>=1.29.0— Tracing and metrics - OpenTelemetry SDK
>=1.29.0— Telemetry primitives - genai-prices
>=0.1.9,<0.2— LLM pricing data for cost resolution - markdown-it-py
>=3.0— Structure-aware Markdown chunking - python-dotenv
>=1.0.0—.envloading for example scripts - PyYAML
>=6.0— YAML prompt and skill metadata
Optional dependencies (installed via extras):
[embeddings]— numpy for fast in-memory vector math[openai-embeddings]— openai>=1.0.0for OpenAI/Azure embeddings[vectorstores-chroma]— chromadb>=0.5.0[vectorstores-pinecone]— pinecone>=5.0.0[vectorstores-qdrant]— qdrant-client>=1.12.0[vectorstores-pgvector]— asyncpg>=0.30.0for Postgres + pgvector[vectorstores-sqlite-vec]— sqlite-vec>=0.1.6for embedded vector search[binary]— pypdf, Pillow, pillow-heif, cairosvg, py7zr, extract-msg forcontent.binary[script-execution]— Pydantic Monty for sandboxed Python execution[evaluation]— Ragas and LangChain adapters; see the dependency caveat[all]— Runtime integration bundle (memory backends, security, embedding providers, vector stores, watch, binary, script execution); excludes evaluation, reasoning-eval, and dev tools
Hosted LLM credentials (for the provider you select):
OPENAI_API_KEYfor OpenAI modelsANTHROPIC_API_KEYfor Anthropic modelsGEMINI_API_KEYfor Google Gemini modelsGROQ_API_KEYfor Groq models- Or any Pydantic AI-supported provider
Local models and test models may not require a provider API key.
Install the wheel published with v26.09.0 in a Python 3.13+ virtual environment:
python -m pip install "https://github.com/fireflyframework/fireflyframework-agentic/releases/download/v26.09.0/fireflyframework_agentic-26.9.0-py3-none-any.whl"To include an optional integration, use a direct-reference requirement:
python -m pip install "fireflyframework-agentic[postgres] @ https://github.com/fireflyframework/fireflyframework-agentic/releases/download/v26.09.0/fireflyframework_agentic-26.9.0-py3-none-any.whl"Releases use YY.MM.Patch, starting at patch 0 for each new month. Python
normalizes 26.09.0 to 26.9.0 in package metadata and filenames. GitHub Releases
provides both the wheel and source archive; the release workflow does not publish
to PyPI. The wheel above pins the framework version for CI; lock your application's
dependencies as well when you need a reproducible environment.
The interactive installer detects your platform, checks Python and UV, lets you
choose extras, and installs from the repository's current main branch.
macOS / Linux:
curl -fsSL https://raw.githubusercontent.com/fireflyframework/fireflyframework-agentic/main/install.sh | bashWindows (PowerShell):
irm https://raw.githubusercontent.com/fireflyframework/fireflyframework-agentic/main/install.ps1 | iexThe PowerShell installer also accepts explicit non-interactive options after you
download install.ps1:
# Windows — install with all extras, no prompts
.\install.ps1 -NonInteractive -Extras allgit clone https://github.com/fireflyframework/fireflyframework-agentic.git
cd fireflyframework-agentic
uv sync --extra all # or: pip install -e ".[all]"| Extra | What it adds | When you need it |
|---|---|---|
postgres |
asyncpg, SQLAlchemy | PostgreSQL memory / storage persistence |
mongodb |
motor, pymongo | MongoDB memory persistence |
security |
cryptography | At-rest encryption (EncryptedMemoryStore, AESEncryptionProvider) |
embeddings |
numpy | Fast in-memory vector math |
openai-embeddings |
openai | OpenAI / Azure text embeddings |
cohere-embeddings |
cohere | Cohere text embeddings |
google-embeddings |
google-generativeai | Google text embeddings |
mistral-embeddings |
mistralai | Mistral text embeddings |
voyage-embeddings |
voyageai | Voyage AI text embeddings |
azure-embeddings |
openai | Azure OpenAI text embeddings |
bedrock-embeddings |
boto3 | AWS Bedrock text embeddings |
ollama-embeddings |
httpx | Ollama local text embeddings |
vectorstores-chroma |
chromadb | ChromaDB vector store backend |
vectorstores-pinecone |
pinecone | Pinecone vector store backend |
vectorstores-qdrant |
qdrant-client | Qdrant vector store backend |
vectorstores-pgvector |
asyncpg | Postgres + pgvector vector store backend |
vectorstores-sqlite-vec |
sqlite-vec | Embedded sqlite-vec vector store backend |
binary |
pypdf, Pillow, pillow-heif, cairosvg, py7zr, extract-msg | content.binary file normalisation |
watch |
watchfiles | File-watching for content sources |
script-execution |
pydantic-monty | Sandboxed Python execution and Firefly Code Mode |
reasoning-eval |
numpy, pandas | Reasoning quality comparisons |
evaluation |
Ragas, LangChain provider adapters | LLM-as-judge and retrieval evaluation; compatibility constraints apply |
dev |
pytest, Ruff, Pyright, pre-commit, Testcontainers | Framework development and checks |
all |
Runtime integrations above | Excludes reasoning-eval, evaluation, and dev; uv sync --all-extras includes these too |
python -c "import fireflyframework_agentic; print(fireflyframework_agentic.__version__)"macOS / Linux:
curl -fsSL https://raw.githubusercontent.com/fireflyframework/fireflyframework-agentic/main/uninstall.sh | bashWindows (PowerShell):
irm https://raw.githubusercontent.com/fireflyframework/fireflyframework-agentic/main/uninstall.ps1 | iexOr manually remove the cloned directory and its virtual environment.
Create a .env file beside your script:
OPENAI_API_KEY=your-api-key
FIREFLY_AGENTIC_DEFAULT_MODEL=openai-responses:gpt-6-lunaKeep .env out of version control. To use another provider, change the model
selector and supply that provider's credentials. Choose a model that supports tool
calling. The script below loads .env; existing environment variables take
precedence.
Save this complete script as app.py. It uses only Firefly's agent, tool, memory,
and model-options APIs:
import asyncio
from dotenv import load_dotenv
load_dotenv()
from fireflyframework_agentic.agents import FireflyAgent
from fireflyframework_agentic.memory import MemoryManager
from fireflyframework_agentic.models import ModelOptions
from fireflyframework_agentic.tools import firefly_tool
GLOSSARY = {
"bounded retries": "Retry a failed operation up to a configured attempt limit.",
"backoff": "Increase the delay between retry attempts to reduce load.",
}
@firefly_tool("lookup", auto_register=False)
async def lookup(term: str) -> str:
"""Look up an engineering term in the local glossary."""
return GLOSSARY.get(term.strip().lower(), "No glossary entry exists for that term.")
async def main() -> None:
memory = MemoryManager()
agent = FireflyAgent(
name="glossary",
instructions="Use the lookup tool for glossary definitions. Keep answers brief.",
tools=[lookup],
memory=memory,
model_options=ModelOptions(max_tokens=4096),
auto_register=False,
)
conversation_id = memory.new_conversation()
first = await agent.run(
"Use the glossary to explain bounded retries.",
conversation_id=conversation_id,
)
print(first.output)
follow_up = await agent.run(
"Which term did I ask you to explain?",
conversation_id=conversation_id,
model_options=ModelOptions(max_tokens=2048),
)
print(follow_up.output)
if __name__ == "__main__":
asyncio.run(main())python app.pyThis makes real provider requests. tools=[lookup] attaches the tool to the agent;
reusing conversation_id preserves the conversation. Changing the configured model
or supported API does not require rewriting the tool or memory code.
For the configured GPT-6 Luna Responses model, add reasoning="low" to
ModelOptions when reasoning is useful. Other models and APIs have their own
reasoning/tool constraints; consult the model compatibility guide.
Responses-specific storage is controlled with store_responses=False; these
options are validated against the selected model and API. See the
Responses example for reasoning, structured output,
and streaming with those controls.
| Add | Guide or runnable example |
|---|---|
| Plain Pydantic output schemas and streaming | Model-agnostic agent |
| Decorator-defined agents | Agent decorators |
| Tool approval and deferred runs | Human-in-the-loop tools |
| Persistent working memory and conversation export/import | Memory |
| Reasoning patterns | Reasoning |
| Output validation and review | Validation |
| Multi-agent pipelines | Pipeline |
| Embeddings and retrieval | Vector stores |
Browse the example catalogue for setup requirements, credentials, and offline or live execution modes.
firefly-agentic works seamlessly in Jupyter notebooks and JupyterLab.
Since the framework is async-first, use await directly in notebook cells
(Jupyter provides a running event loop automatically).
# From your clone directory
cd fireflyframework-agentic
source .venv/bin/activate # activate the venv created by the installer
pip install ipykernel # install Jupyter kernel support
python -m ipykernel install --user --name fireflyagentic --display-name "Firefly Agentic"
jupyter lab # or: jupyter notebookThen select the Firefly Agentic kernel when creating a new notebook.
# Cell 1 — configure
import os
from getpass import getpass
from dotenv import load_dotenv
load_dotenv()
if not os.environ.get("OPENAI_API_KEY"):
os.environ["OPENAI_API_KEY"] = getpass("OpenAI API key: ")
os.environ["FIREFLY_AGENTIC_DEFAULT_MODEL"] = "openai-responses:gpt-6-luna"# Cell 2 — create an agent
from fireflyframework_agentic.agents import FireflyAgent
from fireflyframework_agentic.models import ModelOptions
agent = FireflyAgent(name="notebook-bot", model_options=ModelOptions(max_tokens=4096))
result = await agent.run("Explain quantum entanglement in two sentences.")
print(result.output)# Cell 3 — use memory for multi-turn conversations
from fireflyframework_agentic.memory import MemoryManager
memory = MemoryManager(max_conversation_tokens=32_000)
agent_with_mem = FireflyAgent(name="chat", memory=memory)
cid = memory.new_conversation()
result = await agent_with_mem.run("My name is Alice.", conversation_id=cid)
print(result.output)
result = await agent_with_mem.run("What is my name?", conversation_id=cid)
print(result.output) # Alice# Cell 4 — reasoning patterns
from fireflyframework_agentic.reasoning import ReActPattern
react = ReActPattern(max_steps=5)
result = await react.execute(agent, "What are the top 3 uses of Python in 2026?")
print(result.output)# Cell 5 — structured output with validation
from pydantic import BaseModel
from fireflyframework_agentic.validation import OutputReviewer
class Summary(BaseModel):
title: str
bullet_points: list[str]
confidence: float
summary_agent = FireflyAgent(name="summary", output_type=Summary)
reviewer = OutputReviewer(output_type=Summary, max_retries=2)
result = await reviewer.review(summary_agent, "Summarize the benefits of async Python.")
result.output # displays the structured Summary object in the notebookTip: You do not need
asyncio.run()ornest_asyncioin Jupyter —awaitworks at the top level of any cell because Jupyter runs its own event loop.
docs/tutorial.md is an 18-chapter, hands-on guide that teaches every concept from zero to expert through a real-world Intelligent Document Processing pipeline. Start here if you want to learn the framework thoroughly.
docs/use-case-idp.md is a focused walkthrough of building a 7-phase IDP pipeline that ingests, splits, classifies, extracts, validates, assembles, and explains data from corporate documents — using agents, reasoning, document splitting, content processing, validation, explainability, and pipelines.
Detailed guides for each module:
- Architecture — Design principles and layer diagram
- Agents — Lifecycle, registry, delegation, decorators, human-in-the-loop approval
- Models — Model factory, provider settings, Chat Completions/Responses selection, model capabilities
- Migration — Pydantic AI 2.x, OpenAI API selection, tool and workflow changes
- Template Agents — Summarizer, classifier, extractor, conversational, router
- Tools — Protocol, builder, guards, composition, built-ins, native HITL approval (
requires_approval, deferred resume) - Prompts — Templates, versioning, composition, validation
- Reasoning Patterns — 6 patterns, structured outputs, output modes (
output_mode/reasoning_output_mode), custom patterns - Content — Chunking, compression, batch processing
- Memory — Conversation history, working memory, storage backends
- Validation — Rules, QoS guards, output reviewer
- Embeddings — 8 providers, auto-batching, similarity, registry
- Vector Stores — 6 backends, tenant scoping, auto-embedding, search_text, namespaces
- Pipeline — DAG orchestrator, parallel execution, checkpointing, audit log, retries
- Dynamic Workflows — Code-defined orchestration DSL over agents:
@workflow,agent/parallel/pipeline/stream, budgets, journal resume, smart routing, sub-workflows, HITL,FireflyAgentRunner - Observability — Tracing, native pydantic-ai instrumentation, metrics, events, provider-agnostic cost resolvers, budget gates
- Resilience — Circuit breaker (state machine + middleware), fast-fail on cascading failures
- Storage — Managed-SQLite durable layer: atomic writes, cross-process leasing
- Explainability — Decision recording, audit trails, reports
- Security — Prompt/output guards, at-rest encryption
- Secure Script Execution — Deny-by-default Monty sandbox, static safety pre-screen,
SecureScriptRunner, Firefly Code Mode - Experiments — A/B testing, variant comparison
- Lab — Benchmarks, datasets, evaluators
- Evaluation — LLM-as-judge metrics, RAGAS, retrieval metrics
- Studio — moved to fireflyframework-agentic-studio
git clone https://github.com/fireflyframework/fireflyframework-agentic.git
cd fireflyframework-agentic
uv sync --all-extrasuv run pytest -m "not nightly"
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv tool run --with-requirements docs/requirements.txt mkdocs build --strictThe PostgreSQL integration tests start a disposable Docker container. Without
Docker, use uv run pytest -m "not nightly and not integration" for a smaller
suite. Live provider tests require credentials and are opt-in nightly tests.
See the test guide for coverage and execution boundaries.
Dev dependencies (the [dev] extra, included by uv sync --all-extras):
pytest >=8.3.0,
pytest-asyncio >=0.24.0,
pytest-cov >=6.0.0,
ruff >=0.9.0,
pyright >=1.1.0,
httpx >=0.28.0.
See CONTRIBUTING.md for guidelines.
See CHANGELOG.md for notable changes.
Apache License 2.0. See LICENSE for the full text.