DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running.
-
Updated
Sep 26, 2026 - Go
DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running.
See what Claude Code and Codex actually send to the API — and what each part costs.
Context compression for AI coding agents: compresses tool output before it enters the model, dedups repeats to 13-token refs. Claude Code, Cursor, Codex, Kiro, Zed, any MCP client. Rust, zero LLM calls.
Streamline your workflow with Lynkr, a CLI tool that acts as an HTTP proxy for efficient code interactions using Claude Code CLI.
Companionship, chat, coding, and work share one memory and context framework — the kind of AI you see in science fiction: it keeps you company, and it gets things done with you.(这是一个基于上下文和注意力机制做的一个多元化的agent项目)
Your agent pays twice for output it has already seen. OMNI returns a handle instead: 97.2% off a file read twice. Nothing deleted, nothing invented.
Local-first MCP tools and dashboard for investigating Codex token usage, credits, costs, caching, and thread patterns.
The Multi-Agent Reasoning framework creates an interactive chatbot where AI agents collaborate via structured reasoning and Swarm Integration for optimal answers. Simulating a team that discusses, debates, and refines responses, it enables complex problem-solving and precise results. Now with Prompt Caching to reduce latency and costs.
A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.
Context intelligence layer for LLM agents: persistent memory with write-time dedup, sensitivity tagging, conflict detection, and hierarchical decay. ~12ms. No LLM calls. MIT.
Anthropic Claude API wrapper for Go
🚀 Autocache - Intelligent Anthropic API Cache Proxy Automatically inject cache-control fields into Claude API requests to reduce costs by up to 90% and latency by up to 85%. Works as a transparent drop-in replacement for popular AI platforms like n8n, Flowise, Make.com, LangChain, and LlamaIndex—no code changes required
Independent research on Claude Code internals, Claude Agent SDK, and related tooling.
AI coding agent with one Python core and three front-ends — headless CLI, Textual TUI, and an Electron desktop. Works with any OpenAI-compatible API, with risk-tiered permissions, event-sourced replayable sessions, and a fail-closed OS-level sandbox.
Multi-Agent Code Review Cost Benchmark 2026: Librarian vs Prompt Cache
Drop-in prompt-caching fixes for the LLM agent harness you use. Point your AI coding agent at this repo and it ships the patches.
Claude Opus 5 API (claude-opus-5): model id, per-million-token pricing, cache write and cache read rates, streaming and Anthropic-style requests through an OpenAI-compatible gateway.
A curated list of strategies, tools, papers, and resources for reducing LLM token costs and improving efficiency in production.
A Complete faster runtime for coding agents. Optimize the journey, not the hop. Make coding agents 25% faster and 30% cheaper on average while keeping the quality same or more. Same Task, Same Quality, Faster and Cheaper.
A tiny Claude Code skill that keeps your prompt cache warm during idle sessions, so your next message reads from cache instead of paying full price.
To associate your repository with the prompt-caching topic, visit your repo's landing page and select "manage topics."