Skip to content

Analyse a gene list typed into the chat - #298

Merged
adamjohnwright merged 4 commits into
mainfrom
gene-list-analysis
Sep 25, 2026
Merged

adamjohnwright merged 4 commits into
mainfrom
gene-list-analysis

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Asked "can we do a gsa analysis in the chat. I want to do it with genes TP53, ERBB2 and RUNX2", the chat replied with upload instructions for an expression matrix. Three genes are not a matrix. What Reactome runs on a gene list is over-representation analysis, and the Analysis Service runs it in about a second, so the chat now offers to do it.

What the reader sees

  1. An offer, not a run. The reply explains that a gene list means over-representation rather than GSA, and lists exactly the identifiers it read. Then it shows Run it and No, answer my question. The reader can click either, type yes, or keep typing.
  2. Run shows "Matched N of M", the top 10 pathways by FDR (names escaped, stable IDs shown), and a link to the result in the Pathway Browser of the service that holds the token (beta by default). Unmatched identifiers are listed. Small lists get a caution.
  3. Follow-ups work. The exchange is seeded as the thread's previous turn. Checked in a browser: "which has the lowest FDR?" is answered from the result.
  4. No:
    • a message that asked about GSA gets the matrix instructions, without the gene-list line (so the declined list is not offered again);
    • anything else goes to the model, as before.

Why it asks first

Recognising a request is a heuristic. The first version ran immediately, and review found it answered 14 of 44 ordinary questions ("Could you explain why IFNG and TNF are enriched…") with a results table. With the offer, a misfire costs one click, and the reader sees what will be submitted before anything is sent.

Three rounds of adversarial review

  • Round 1 — the recogniser. False positives and misses, plus shouting detection firing on gene symbols themselves.
  • Round 2 — speed and flow.
    • Quadratic parsing: a 4,000-gene paste took 38 s, and 40,000 newlines took 24 s, on the shared event loop.
    • Trailing sentences were read as genes.
    • AskActionMessage locked the message box until a click.
    • No looped back to the offer.
  • Round 3 — the button flow.
    • Resume crashed: offers lived in user_session, which is saved as JSON, so cl.Action came back as null.
    • Button callbacks ran with no task (no Stop, errors swallowed).
    • A stale typed yes could run an old offer.
    • More trailing-word cases.

Everything found is fixed, and each fix has a test that fails when the fix is removed (sabotaged and checked).

Measured rates

The phrase sets are in tests/analysis/gene_list_phrases.py, with where each came from. The honest numbers come from sets written after each rewrite and never tuned against:

Set Requests read exactly Questions that trigger an offer
1 15/15 0/20
2 11/12 1/15
4 (latest) 9/10 0/12

Round 3's own fresh set had 4/20 questions trigger an offer before the last fix. The remaining wrong cases are pinned in KNOWN_LIMITS. Pasted tables are one of them: attach the file instead.

Performance

Parsing is a single pass. Gap patterns are possessive, and messages are capped at 60K characters. Timing tests cover each pathological input the reviews found; the worst of them now runs in tens of milliseconds.

Verified

  • ./checks.sh (the same as CI) passes.
  • 23 browser checks pass against a local instance on the beta network, including:
    • the exact message from the report;
    • Run, typed yes, and No (with both routes);
    • an older offer's Run;
    • typing on without answering the offer;
    • Stop shown and sending blocked while a declined question is answered;
    • three question-shaped controls;
    • the Pathway Browser link opening the analysis.

Data

The identifiers go to Reactome's Analysis Service (beta by default, per ANALYSIS_BASE_URL). The seeded turn goes to the model provider on the next message. The reader typed the list into a chat whose every message already goes there.

🤖 Generated with Claude Code

adamjohnwright and others added 4 commits September 25, 2026 19:59
Asked "can we do a gsa analysis in the chat. I want to do it with genes
TP53, ERBB2 and RUNX2", the chat replied with upload instructions for an
expression matrix. A gene list is not a matrix; what Reactome runs on one
is over-representation, and the Analysis Service runs it in a second.

The chat now recognises a request to analyse two or more identifiers,
submits them to the Analysis Service, and replies with the matched count,
the top pathways by FDR, and a Pathway Browser link on the service that
holds the token. The exchange is seeded as the thread's previous turn so
follow-ups are answered from the result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…them

An adversarial review of the first version found the recogniser answered
14 of 44 ordinary questions ("Could you explain why IFNG and TNF are
enriched...") with a results table, and missed or trimmed 20 of 25 real
requests: shouting detection fired on the gene symbols themselves, and
lower-case symbols were never read.

- The chat now shows the identifiers it read and asks: Run it, or No,
  answer my question -- which sends the message to the model as before.
- Lists are read as lists: after a colon, question mark, newline or
  preposition, across commas and "and", in any case.
- Questions about genes (explain, why, compare, ...) are not requests.
- Capitalised words and accession-like strings (NOT, HUMAN, GSE12345,
  chr17) are not submitted; beta matched NOT as a real identifier.
- 'and N more' counts from the service's total, an impossible match count
  is not shown, odd result shapes do not raise, and a newline in a name
  cannot end a table row.
- Seeding uses the same thread id as the next turn, and logs when it
  cannot seed.

Measured on 35 phrasings written after the rewrite and never tuned
against: 15/15 requests read exactly, 0/20 questions fire.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second adversarial review found:
- quadratic parsing: a pasted 4,000-gene column took 38 s, 40,000 newlines
  24 s, on the event loop all sessions share;
- 12 of 40 fresh questions still proposed an analysis;
- trailing sentences read as genes ("... using default settings");
- AskActionMessage locked the message box until a click, and a timeout
  dropped the question;
- No on a message saying "gsa" offered the declined list again.

Now:
- One pass over the tokens, with possessive gap patterns and a 60K cap.
  Timing tests cover each pathological input the review found.
- A list keeps one separator throughout; a blank line after a list ends
  reading; "X and Y" in a question is prose, not a list.
- The offer is a message with actions: the reader can click Run, click
  No, type "yes", or type on. An older offer's buttons still work.
- No sends a GSA question to the matrix-only how-to, anything else to
  the model.

Measured on a third set written after this and never tuned against:
11/12 requests read exactly, 1/15 questions offered an analysis. Both
failures are pinned as known limits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Round three of adversarial review:
- Resuming a thread and typing yes raised AttributeError: offers lived in
  user_session, which Chainlit saves as JSON, and cl.Action became null.
  Offers are now plain data in a bounded in-process store.
- Button callbacks ran with no task: no Stop, sending not blocked, errors
  swallowed, and the work held the action HTTP request open. They now run
  as a stoppable task that reports failure, as a message does.
- No after a restart said nothing; it now says the offer lapsed. Evicted
  offers lose their buttons.
- The typed-yes state is read and cleared before any early return.
- A lower-case word after a list in capitals is the sentence going on
  ("TP53, MDM2, then show me..."); a lone hyphen is a separator; "in"
  opens a list; two genes in a question are its subject, not a list.
- The column timing test now reaches the parser (it was over the cap).

Measured on a fourth set written after these fixes and never tuned
against: 9/10 requests exact, 0/12 questions offered an analysis.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 87af16f into main Sep 25, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the gene-list-analysis branch September 25, 2026 21:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant