Skip to content
Use this GitHub action with your project
Add this Action to an existing workflow or create a new one
View on Marketplace

Repository files navigation

Codelens

High performance code analysis tool written in Rust — stats, health scores, hotspots, change coupling, diffs, trends, cost estimation, and CI/AI quality gates.

Features

  • Fast: Parallel file traversal, 30-50x faster than Python alternatives
  • 65+ Languages: Built-in support for popular programming languages
  • Smart Filtering: Respects .gitignore, .gitattributes linguist attributes, auto-excludes build directories
  • Multiple Outputs: Console, JSON, CSV, Markdown, HTML with charts, OpenMetrics, badge JSON, SARIF
  • Complexity Analysis: Function count, cyclomatic + cognitive complexity, nesting depth
  • Health Score: Project/directory/file-level health grading (A-F) across six dimensions including line duplication (ULOC/DRYness), with pluggable scoring models
  • CI Quality Gates: --fail-under absolute gate, --baseline regression gate, and codelens diff for ref-to-ref health deltas ("clean as you code"), plus an official GitHub Action and pre-commit hooks
  • Hotspot Detection: Risky files via churn × complexity, with code age, author knowledge risk, and function-level breakdown
  • Change Coupling: Files that keep changing together — hidden dependencies the module structure doesn't show
  • Trend Tracking: Save snapshots, compare evolution, chart the full history
  • Cost Estimation: Multi-model development cost estimation (COCOMO Basic/II, Putnam, LOCOMO)
  • LLM Token Estimation: How many tokens a repo is, and whether it fits a model's context window
  • AI Agent Integration: Built-in MCP server (codelens mcp) plus an Agent Skill teaching agents the analysis workflows — see docs/ai-integration.md
  • Extensible: Add custom languages via TOML configuration

Installation

Homebrew (macOS/Linux)

brew install DropFan/tap/codelens

Cargo

cargo install codelens

Build from source

git clone https://github.com/DropFan/codelens
cd codelens
cargo build --release

Usage

# Analyze current directory
codelens

# Analyze specific directories
codelens src tests

# Only count specific languages
codelens -l rust,go,python

# Output JSON
codelens -f json -O stats.json

# Output HTML report
codelens -f html -O report.html

# Show top 20 languages by code lines
codelens --top 20 --sort code

# Exclude directories
codelens --exclude vendor,dist,node_modules

# List supported languages
codelens --list-languages

Health Score

Score code health across six dimensions (complexity, function size, comment ratio, file size, nesting depth, line duplication) with grades from A to F.

The historical v1 pipeline remains the default so existing dashboards and CI gates keep their established scores. It aggregates raw metrics across all analyzed files with the original weights, curves, function matcher, and bracket-depth behavior.

The opt-in v2 pipeline scores each production file first. A file score is 85% weighted dimensions plus 15% from its weakest reliable core dimension. Language and directory scores combine an sqrt(code lines) weighted center with the worst 10% of their files, using an 85/15 split. The project score then weights language scores by production code lines. Test code is reported separately; documentation and data formats do not affect the main score. Unavailable measurements are omitted and the remaining weights are renormalized, with confidence showing how much of the configured model was actually measured. New analyses retain both v1 and v2 complexity inputs; snapshots created before this feature fall back to their original metrics. v1 and v2 scores are intentionally not comparable, so use the same version on both sides of a baseline and rebuild saved baselines when migrating.

codelens health .               # Project, directory, and file-level report
codelens health . --top 20      # Show top 20 worst files
codelens health . -f json       # Output as JSON
codelens health . --health-model v2   # Use the multilingual algorithm
codelens health . --fail-under B   # CI gate: exit 1 if health is below B
codelens health . --baseline main --fail-on-regression   # regression gate

--health-model v1|v2 is a global option and applies to the default report, health, diff, and MCP server. It can also be set in .codelens.toml:

health_model = "v1"

Command-line selection overrides configuration; configuration overrides the built-in default (v1). Machine-readable reports retain the stable internal names default (v1) and default-v2 (v2); HTML reports display the clearer user-facing names v1 and v2.

--fail-under accepts a grade (A/B/C/D) or a numeric score (75), turning the health report into a CI quality gate — fail a PR when project health drops below your threshold.

--baseline compares against a trend snapshot (latest, latest~1, a date) or any git ref (main, HEAD~1, a tag — analyzed via a temporary worktree). With --fail-on-regression, the gate fails when the project or an existing file drops a grade, when either drops at least 5 points within the same grade, or when a new file receives an F. Legacy debt stays visible without blocking unrelated changes ("clean as you code"). The delta ("B 87.9 → C 77.2") renders in console, markdown (great for PR comments), HTML, and JSON. Comparisons use only measurements available on both sides and include per-language score changes. A tests-only ↔ production scope change is called out explicitly; its project-level delta is informational and does not fail the regression gate.

On very large codebases, --no-dup-scan skips line-duplication collection to save memory; the health score then omits the duplication dimension (it is not scored as clean — the remaining weights are renormalized), and machine-readable output reports the scoring model as default-no-dup for v1 or default-v2-no-dup for v2.

Hotspot Detection

Find the riskiest files by combining git change frequency (churn) with code complexity — files that change often AND are complex are the most likely sources of bugs.

codelens hotspot .              # Last 90 days (default)
codelens hotspot . --since 30d  # Last 30 days
codelens hotspot . --since 6m --top 5  # Last 6 months, top 5
codelens hotspot . --functions  # Which functions inside absorb the churn

Each hotspot shows its age (days since first commit, rename-aware) and its author concentration: a risky file owned ≥75% by a single author is flagged as a ★ knowledge island — frequently changed, complex, and effectively known by one person. --functions intersects diff hunks with function spans to show which functions inside the top files actually change (approximate, no AST).

Diff Two Refs

Compare two git refs — or a ref against your working tree — with health movement as the headline (raw line counts are already covered by git diff):

codelens diff main                   # main vs working tree
codelens diff main..HEAD             # two refs
codelens diff v1.0 v2.0 --fail-on-regression   # CI gate

Output: project health before → after, the files whose grade dropped, and complexity deltas. This is the data source behind the PR story "this change takes health from B to C — blocked".

Change Coupling

Find file pairs that keep changing in the same commits — hidden dependencies the module structure does not express, and prime refactoring targets.

codelens coupling .                   # Last 90 days, noise-filtered
codelens coupling . --for src/api.rs  # What changes together with this file
codelens coupling . --min-shared 3 --min-coupling 20  # Lower thresholds
codelens coupling . --include-tests   # Keep test files in the pairing

Bulk commits (more than --max-changeset files, default 30) are excluded from pairing and reported, so formatting sweeps don't fake coupling.

Test files (foo_test.go, test_*.py, *.test.ts, tests/, ...) are excluded by default: a test changing together with the code it tests is expected behavior, not a hidden dependency, and those pairs would drown the real signal. Use --include-tests to keep them; --for pointed at a test file implies it.

Trend Tracking

Save snapshots and compare codebase evolution over time. Snapshots are stored in .codelens/snapshots/. Use latest, latest~N, or a date prefix like 2025-01-01 as references.

codelens trend --save --label v1.0     # Save a labeled snapshot
codelens trend                         # Compare latest two snapshots
codelens trend --list                  # List all snapshots
codelens trend --compare latest~2 latest  # Compare specific snapshots

Cost Estimation

Estimate development cost, schedule, and team size using four pluggable models. Default mode runs all models and shows a comparison table.

codelens estimate .                       # All models comparison (default)
codelens estimate . --model cocomo-basic   # Single model with per-language breakdown
codelens estimate . --model cocomo2        # COCOMO II Post-Architecture
codelens estimate . --model putnam --ck 11000  # Putnam with custom productivity
codelens estimate . --model locomo         # LLM generation cost
codelens estimate . --avg-wage 120000      # Custom salary across all models
Model Description Reference Typical Use
COCOMO Basic Classic Boehm 1981 regression Software Engineering Economics, Boehm 1981 (ISBN 0-13-822122-7) Quick estimates, scc comparison
COCOMO II Modern 2000 calibration with scale factors Software Cost Estimation with COCOMO II, Boehm et al. 2000 (ISBN 0-13-026692-2) Organization-level planning
Putnam/SLIM Rayleigh-curve conservative model A General Empirical Solution to the Macro Software Sizing and Estimating Problem, IEEE TSE 1978 Risk assessment, worst case
LOCOMO LLM token cost model scc LOCOMO model, Boyter 2026 AI-assisted development cost

Output Formats

Format Flag Description
Console -f console Colored terminal output (default)
JSON -f json Structured data for processing
CSV -f csv Spreadsheet compatible
Markdown -f markdown Documentation friendly
HTML -f html Interactive report with charts
OpenMetrics -f openmetrics Prometheus text format for scraping
Badge -f badge shields.io endpoint JSON (codelens health -f badge → live code-health badge)
SARIF -f sarif SARIF 2.1.0 for GitHub code scanning (upload-sarif) or reviewdog -f=sarif

CI Integration

GitHub Action — health gate + sticky PR comment + step summary in one step (see docs/github-action.md):

- uses: DropFan/codelens@rust
  with:
    fail-under: 'C'
    baseline: 'origin/${{ github.base_ref }}'
    fail-on-regression: 'true'

pre-commit — gate commits locally with the bundled .pre-commit-hooks.yaml. Install codelens first (the hooks run the binary on your PATH), and pin rev to v0.1.6-rust or newer (earlier tags do not ship the hook manifest):

repos:
  - repo: https://github.com/DropFan/codelens
    rev: v0.2.0-rust
    hooks:
      - id: codelens-health
        args: ['--fail-under', 'C']

AI Agent Integration

codelens mcp runs a built-in MCP server so coding agents can query repository stats, health, hotspots, and coupling before editing code (see docs/ai-integration.md):

claude mcp add codelens -- codelens mcp

codelens . --tokens estimates the repository's LLM token count and whether it fits common context windows (byte-based estimate).

Filtering & Detection

codelens --by-file --top 20         # Per-file statistics (respects --sort/--top)
codelens --by-dir --dir-depth 2     # Directory tree rollups (files/code/complexity)
codelens --count-as jsp:html        # Count .jsp files as HTML
codelens --no-duplicates            # Skip files with identical content
codelens --no-min-gen               # Skip minified/generated files
codelens --no-dup-scan              # Skip line-duplication analysis (saves memory on huge repos)
  • Extensionless scripts are detected via shebang (#!/usr/bin/env python).
  • Drop a .codelensignore file (gitignore syntax) anywhere in the tree to exclude paths, like scc's .sccignore / tokei's .tokeignore.
  • .gitattributes linguist attributes are honored by default so numbers match GitHub: linguist-language=X overrides detection, linguist-vendored / linguist-generated exclude files (--no-linguist opts out).
  • Files matching test conventions (tests/, *_test.go, *.spec.ts, FooTest.java, ...) are reported separately with a test/code ratio.

Configuration

Create .codelens.toml in your project root. CLI flags override config file values, which override built-in defaults:

# Exclude patterns
excludes = "*test*,*mock*"

# Target languages
lang = "rust,go,python"

# Extension remapping
count_as = "jsp:html,tpl:php"

# Output format
output = "json"

# Per-file statistics
by_file = true

# Threading
threads = 8

# Depth limit
depth = 10

# Show git info
git_info = true

# Skip line-duplication analysis (health score omits the duplication dimension)
no_dup_scan = true

# Health scoring pipeline: "v1" (default) or "v2" (multilingual, opt-in)
health_model = "v2"

# Custom language definitions (relative paths resolve against this file's directory)
languages_file = "my-langs.toml"

Custom Languages

Teach codelens new languages with --languages-file: a TOML file with one table per language id. A table whose id matches a built-in language replaces that definition entirely — fields are not merged — so when overriding a built-in language, restate every field you still need (extensions, line_comments, ...); an override that omits extensions stops those files from being detected at all. The built-in languages.toml shows all fields in use, including string_delimiters for accurate string/comment parsing. function_pattern and complexity_keywords are validated when the file loads; an invalid regex is a hard error rather than silently disabling function and complexity analysis.

# my-langs.toml

# A new language:
[mylang]
name = "MyLang"
extensions = [".myl"]
filenames = ["Mylfile"]           # exact filename matches (like Makefile)
line_comments = ["#"]
block_comments = [["/*", "*/"]]   # [open, close] pairs
nested_comments = false           # true if /* /* */ */ nests
function_pattern = "(?m)^\\s*def\\s+\\w+"
complexity_keywords = ["if", "elif", "else", "for", "while"]

# Overriding a built-in language: the whole definition is replaced, so
# copy the fields you want to keep from the built-in languages.toml.
[rust]
name = "Rust"
extensions = [".rs"]              # omitting this would stop .rs detection
line_comments = ["//"]            # omitting this would count // as code
block_comments = [["/*", "*/"]]
nested_comments = true
codelens --languages-file my-langs.toml                    # analyze with the extra definitions
codelens --languages-file my-langs.toml --list-languages   # verify: MyLang appears in the list
codelens health . --languages-file my-langs.toml           # works on subcommands too

--count-as can map further extensions onto a custom language (--count-as myx:mylang). Analysis commands also pick the path up from .codelens.toml (languages_file = "my-langs.toml"); a relative path there resolves relative to the directory containing the config file. --list-languages only honors the explicit flag.

One built-in ambiguity to know about: .m belongs to both Objective-C and MATLAB. codelens assigns it to Objective-C, which matches what scc reports on real Objective-C sources. MATLAB projects can reclaim the extension with --count-as m:matlab — prefer that over overriding [matlab] in a languages file, since an override replaces the whole definition.

License

MIT

About

High performance code analysis tool — stats, health scores, hotspots, change coupling, diffs, trends, cost estimation, and CI/AI quality gates

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages