Analyse a gene list typed into the chat - #298
Merged
Merged
Conversation
Asked "can we do a gsa analysis in the chat. I want to do it with genes TP53, ERBB2 and RUNX2", the chat replied with upload instructions for an expression matrix. A gene list is not a matrix; what Reactome runs on one is over-representation, and the Analysis Service runs it in a second. The chat now recognises a request to analyse two or more identifiers, submits them to the Analysis Service, and replies with the matched count, the top pathways by FDR, and a Pathway Browser link on the service that holds the token. The exchange is seeded as the thread's previous turn so follow-ups are answered from the result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…them
An adversarial review of the first version found the recogniser answered
14 of 44 ordinary questions ("Could you explain why IFNG and TNF are
enriched...") with a results table, and missed or trimmed 20 of 25 real
requests: shouting detection fired on the gene symbols themselves, and
lower-case symbols were never read.
- The chat now shows the identifiers it read and asks: Run it, or No,
answer my question -- which sends the message to the model as before.
- Lists are read as lists: after a colon, question mark, newline or
preposition, across commas and "and", in any case.
- Questions about genes (explain, why, compare, ...) are not requests.
- Capitalised words and accession-like strings (NOT, HUMAN, GSE12345,
chr17) are not submitted; beta matched NOT as a real identifier.
- 'and N more' counts from the service's total, an impossible match count
is not shown, odd result shapes do not raise, and a newline in a name
cannot end a table row.
- Seeding uses the same thread id as the next turn, and logs when it
cannot seed.
Measured on 35 phrasings written after the rewrite and never tuned
against: 15/15 requests read exactly, 0/20 questions fire.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second adversarial review found:
- quadratic parsing: a pasted 4,000-gene column took 38 s, 40,000 newlines
24 s, on the event loop all sessions share;
- 12 of 40 fresh questions still proposed an analysis;
- trailing sentences read as genes ("... using default settings");
- AskActionMessage locked the message box until a click, and a timeout
dropped the question;
- No on a message saying "gsa" offered the declined list again.
Now:
- One pass over the tokens, with possessive gap patterns and a 60K cap.
Timing tests cover each pathological input the review found.
- A list keeps one separator throughout; a blank line after a list ends
reading; "X and Y" in a question is prose, not a list.
- The offer is a message with actions: the reader can click Run, click
No, type "yes", or type on. An older offer's buttons still work.
- No sends a GSA question to the matrix-only how-to, anything else to
the model.
Measured on a third set written after this and never tuned against:
11/12 requests read exactly, 1/15 questions offered an analysis. Both
failures are pinned as known limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Round three of adversarial review:
- Resuming a thread and typing yes raised AttributeError: offers lived in
user_session, which Chainlit saves as JSON, and cl.Action became null.
Offers are now plain data in a bounded in-process store.
- Button callbacks ran with no task: no Stop, sending not blocked, errors
swallowed, and the work held the action HTTP request open. They now run
as a stoppable task that reports failure, as a message does.
- No after a restart said nothing; it now says the offer lapsed. Evicted
offers lose their buttons.
- The typed-yes state is read and cleared before any early return.
- A lower-case word after a list in capitals is the sentence going on
("TP53, MDM2, then show me..."); a lone hyphen is a separator; "in"
opens a list; two genes in a question are its subject, not a list.
- The column timing test now reaches the parser (it was over the cap).
Measured on a fourth set written after these fixes and never tuned
against: 9/10 requests exact, 0/12 questions offered an analysis.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Asked "can we do a gsa analysis in the chat. I want to do it with genes TP53, ERBB2 and RUNX2", the chat replied with upload instructions for an expression matrix. Three genes are not a matrix. What Reactome runs on a gene list is over-representation analysis, and the Analysis Service runs it in about a second, so the chat now offers to do it.
What the reader sees
Why it asks first
Recognising a request is a heuristic. The first version ran immediately, and review found it answered 14 of 44 ordinary questions ("Could you explain why IFNG and TNF are enriched…") with a results table. With the offer, a misfire costs one click, and the reader sees what will be submitted before anything is sent.
Three rounds of adversarial review
AskActionMessagelocked the message box until a click.user_session, which is saved as JSON, socl.Actioncame back as null.Everything found is fixed, and each fix has a test that fails when the fix is removed (sabotaged and checked).
Measured rates
The phrase sets are in
tests/analysis/gene_list_phrases.py, with where each came from. The honest numbers come from sets written after each rewrite and never tuned against:Round 3's own fresh set had 4/20 questions trigger an offer before the last fix. The remaining wrong cases are pinned in
KNOWN_LIMITS. Pasted tables are one of them: attach the file instead.Performance
Parsing is a single pass. Gap patterns are possessive, and messages are capped at 60K characters. Timing tests cover each pathological input the reviews found; the worst of them now runs in tens of milliseconds.
Verified
./checks.sh(the same as CI) passes.Data
The identifiers go to Reactome's Analysis Service (beta by default, per
ANALYSIS_BASE_URL). The seeded turn goes to the model provider on the next message. The reader typed the list into a chat whose every message already goes there.🤖 Generated with Claude Code