Skip to content

Site search: find symbols by name, rank by reader intent, add hand-marked index entries - #210

Merged
KubaO merged 61 commits into
twinbasic:mainfrom
KubaO:claude/paintpicture-docs-runtime-f3250d
Sep 27, 2026
Merged

KubaO merged 61 commits into
twinbasic:mainfrom
KubaO:claude/paintpicture-docs-runtime-f3250d

Conversation

@KubaO

@KubaO KubaO commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Site search couldn't find most reference members. PaintPicture wasn't in the top 10. On the first eval's symbol queries, hit@10 was 20.5% and mean reciprocal rank (MRR) 0.182. The eval at the tip has 10,327 queries: 98.2% find the right page at rank 1, and 99.1% within the top 10. WIP.Search.md has the design, every measurement and the known misses.

  • builder: the search index splits at ### headings, so a member such as ### PaintPicture gets its own entry. Generic sub-sections (See Also, Example, Remarks and the like) fold back into the member they belong to. Both are set in docs/_config.yml under search:.
  • builder: each entry carries the symbols documented on it, taken from tB/symbols.json, as bare and qualified names. Printer.Fonts and Printer Fonts both find the member.
  • search: a bare name ranks the way a reader who types it expects: type names and language elements first, then members, then enum constants, then prose. The order is a priority, not a filter. Exact names and whole page titles rank above pages that only share words with the query. A kind word, as in MaxHeight property, narrows the results only while some entry names the thing.
  • search: stop words are kept, and runs of dots split. Operators such as <> and &= are found by name. HTML entities are decoded per token, so &H80004005 is found. A query of only * no longer crashes.
  • search: the index is built on the first keystroke, with a loading message, not on page load.
  • lunr 2.3.9: two patches, in both the site client and the offline build:
    • Its token set could invent words and lose real ones; a query containing a threw.
    • Its set union was quadratic for short wildcard words; a page went from 809 ms to 87 ms.
  • index entries: a page's front matter (index:, index_also:) or a heading ({: index="..." }) names the terms a reader looks up to find it, as a book's index does. Each term has one main entry across the site, and the build refuses a second. A mark on anything but a heading fails the build. The first entries cover jargon that no page's own text finds, such as late binding, immediate window, standard exe and registry access. Every target was approved by hand.
  • docs: new content where no page had it: passing arguments ByRef and ByVal, and optional arguments with default values (Sub); returning more than one value (Function); named arguments (Call). Every sample compiles, and its printed output was checked. There are Glossary entries for namespace and default member, and smaller additions to the IDE pages: the manifest, adding a library reference, and the Variables pane.
  • eval: eval/search_quality.mjs measures eval/site_search.mjs, a replica of the client's search, against ground truth. That covers every symbol, bare, qualified and with its kind; every page title; every page title with a section name; and 108 hand-picked prose queries. --compare against the saved baseline lists every query that got worse. The research scripts and their logs are in eval/search-experiments/.
  • test.bat: test/search.test.mjs unit-tests the search-data generator. It fails if the site client, the offline build's search and the eval's replica drift apart.
  • Authoring.md, Builder.md and Pipeline-Stages.md describe index entries and the new search data. build.bat, check.bat, test.bat and lint are clean at the tip.

Why members documented under ### headings, such as PaintPicture, don't
come up in site search; what a replay of 8,012 queries against the built
index measured; the design that follows from it; the rollout in five
steps; and operator search as future work.
A search for `*` or `**` reached lunr's trailing-wildcard term as a bare
wildcard, which throws inside lunr's query engine and leaves the page's
search broken until it reloads. update() now drops tokens made only of
asterisks, and shows "No results found" when none remain. The offline
build replaces only initSearch() and navLink(), so it inherits the fix.

eval/site_search.mjs now sets the client's tokenizer separator,
/[\s\-\/]+/, which it never did: it tokenised with lunr's default and so
could not reproduce the crash. It mirrors the asterisk filter too.

Step 1 of WIP.Search.md's rollout.
Replays 8,012 queries against a built search-data.json: bare and
Container.Name symbol queries derived from the build's tB/symbols.json,
and 20 hand-picked prose queries in search_prose_queries.json. It reports
hit@1/5/10, MRR, per-category and per-kind breakdowns, and index cost,
and --compare lists the queries that changed rank against a saved run.

It takes its index setup and query logic from eval/site_search.mjs, which
now exports them, so the two cannot drift apart. search_baseline.json holds
the numbers for today's index.

Step 2 of WIP.Search.md's rollout.
search-data.json now has an entry per heading down to h3
(search.heading_level: 3), so a member documented as `### PaintPicture`
is a search result of its own, linking to its anchor, instead of part
of a long "Methods" entry. Sections titled in search.fold_headings (See
Also, Example, Remarks, ...) are appended to the section before them
rather than becoming entries titled only "See Also".

Measured with eval/search_quality.mjs against the previous baseline:
entries 3,781 -> 6,713, gzip 1,039 -> 1,111 KB, hit@10 20.5% -> 35.7%,
MRR .182 -> .312, bare symbol names hit@10 55.8% -> 84.6%, prose
unchanged at 80%. The baseline file now holds these numbers.

test/search.test.mjs is the first unit test of search.mjs: the split,
folding, the prefix entry and determinism, run by test.bat.

Step 3 of WIP.Search.md's rollout.
Every search entry that documents a symbol now carries two fields from
the symbol index: `names` (bare names, boost 100) and `qualified`
(Container.Name forms, boost 50). searchData waits for symbolIndex and
joins them by URL (joinSymbolsToEntries). There are two fields rather
than one because BM25 discounts a match inside a long field. This finds
enum values that only appear in table rows, such as vbKey0.

The client also splits a query token at a dot between identifiers, and
keeps the whole token: Form.PaintPicture finds Form's PaintPicture,
Debug.Print still ranks its own entry first, and 1.0, e.g. and i.e. are
left alone. The split is done without a lookbehind, which Safari before
16.4 cannot parse. offline.mjs's initSearch copy has the same fields,
eval/site_search.mjs mirrors both changes, and test/search.test.mjs
fails if the three field lists drift apart.

Measured with eval/search_quality.mjs against step 3: hit@10 35.7% ->
97.8%, MRR .312 -> .933, qualified hit@10 7.9% -> 97.7%. Prose drops
from 80% to 75%: "symbol index" falls from 1 to 32, because Index is a
property on many controls. 25 queries rank worse than before the whole
rollout, mostly short keywords, which step 5 addresses.

Step 4 of WIP.Search.md's rollout.
lunr's index pipeline drops English stop words, but its search pipeline
keeps them. Do, For, If, Is, On, With and Each are stop words, so a
search for those keywords could never match a page. The index now keeps
them. A tokenizer wrapper also turns runs of two or more dots into
spaces, so Do...Loop and For Each...Next index their keywords as words.
With both, those six keywords find their pages at rank 1.

Keeping stop words makes the in-browser index cost 1,294 ms and 243 MB
(from 1,184 ms and 204 MB; 964 ms and 157 MB before this rollout), and
the browser built it on every page load. It is now built on the first
keystroke instead. "Loading search index..." shows in the results panel
and the a11y live region until results replace it, and the search runs
on whatever the box holds when the build finishes. A failed load says
so, and the next keystroke retries. offline.mjs's initSearch copy does
the same, and test/search.test.mjs checks that all three copies (client,
offline, eval replica) keep stop words and split dot runs.

Measured with eval/search_quality.mjs: hit@10 98.0%, MRR .933, prose
back to 80%, and 18 queries worse than before the rollout, down from 25.
Checked in a browser, online and in the offline tree: no request for the
index on page load, the loading message stays up with no empty panel in
between, and arrow keys and Esc work on the results.

Step 5 of WIP.Search.md's rollout.
Records the measurements after the rollout. Results are now judged by
what a reader wants, not against the pre-rollout index. There's a tiered
ground truth for bare names (type and language element, then member,
then enum constant, then prose), the current rank-1 failure list, and
three client-side fixes measured at hit@1 89.0% -> 92.4% with no query
worse. The tiered exact fields were rejected, with the real cause of the
CheckBox miss: an empty top entry on the class page, not the explanation
the measuring agent gave. A resumption section at the top holds the
state, the criteria, the tools and the lunr pitfalls needed to continue
in a new session.

eval/search-experiments/ keeps the scripts and records of every
measurement round, with a README on regenerating the data they need.
A bare-name query now counts as correct only on a URL of the name's best
tier: a type or language element, then a class member, then an enum
constant. The looser "any page documenting the name" rate is still
reported. New outputs: hit@1 per category and per tier, tier-order recall
and violations for names on 2+ URLs, and --failures N, the queries missing
rank 1. The baseline records which ground truth it was saved with, and
--compare warns when they differ.

Baseline at the current build: hit@1 89.0% (bare 93.3%, qualified 86.6%,
prose 10 of 20), hit@10 97.9%; 526 names on 2+ URLs, recall@10 92.0%,
45 out of tier order.
… first

Three client changes, in both copies of just-the-docs.js and in
eval/site_search.mjs, judged by reader intent:

- An `exact` field (boost 50) holds each bare name lowercased, without a
  trailing `$`, with `_` appended. A one-word query matches it whole, so
  `Node` ranks the Node class above `Nodes`. The other clauses skip it.
- A `page` field (boost 5) holds the page title, so a page titled with the
  query outranks a section elsewhere that mentions it.
- With two or more words, entries containing all of them come first: each
  word is required as its stem plus a trailing wildcard. If none contains
  them all, the ordinary query runs.

Query tokens are also trimmed as the index trims them, so `Date$` finds
`date`; an all-asterisk token trims to nothing, which keeps the crash
guard. Both fields are derived in the browser; search-data.json is
unchanged.

The experiment's X1/X2/X3 made 14 queries worse as measured: a lone
clause's boost cancels out in lunr's scoring, wildcards also searched the
exact field, and the all-words query never applied to capitalised or
hyphenated phrases. With those corrected, against the intent baseline:
hit@1 89.0% -> 93.5% (bare 97.0%, qualified 91.5%, prose 16 of 20), hit@10
98.5%, names out of tier order 45 -> 19, 418 queries better and none
worse. Index build 1,538 -> 1,625 ms, heap 251 -> 280 MB, paid only on the
first search. Both clients checked in a browser against the replica.
…me names

The symbol join now also emits `primary`: the names at each URL that are
types or language elements (a statement, keyword, operator, attribute or
directive outside any package, or a procedure of a module). The clients
index it as exact names at field boost 1000, so `Left` finds the Strings
function before 40 controls' Left properties, and `BorderStyle` the enum
before the same-named properties.

exactName() now spells non-word characters as `_` and their hex code
instead of dropping `$`: lunr's trimmer had turned `#If` into `if` and the
operators into nothing. `#If`, `Time$`, `Error$` and all 24 operators now
reach rank 1. The exact-name clause runs for a query that names one thing,
counting `With statement` as naming `With` (KIND_WORDS, the symbol index's
kinds), and the all-words-first REQUIRED clauses no longer search the name
fields, a leak that had pushed `error handling` down.

Against the previous baseline: hit@1 93.5% -> 94.1%, bare 97.0% -> 98.8%
(hit@10 100%), names out of tier order 19 -> 0; 53 queries better, none
worse. search-data.json +23 KB raw, +4 KB gzip; index build +5-6%, heap
280 -> 288 MB. Both clients checked in a browser against the replica.
…ions changed

A prose query may now list `behind` pages, which a reader should find
right behind the expected one (within the top three). Reported on its
own line; it never changes a query's rank.

`conditional compilation` now expects /tB/Core/Topic-Preprocessor, the
#If/#Const page, with /Reference/Compiler-Constants right behind: rank 6
and 3 today. `symbol index` also accepts Permanent-Links' definition of
the symbol index, which ranks first: 3 -> 1. Ground truth is now
intent-2; the baseline is saved under it.

hit@1 94.15%, prose 17 of 20 at rank 1.
… five

A page could not be found by a term its text never uses: the #If/#Const
page never says "conditional compilation". Authors now name such terms,
as a book's index does: `index:` / `index_also:` in a page's frontmatter,
or `{: index="a; b" }` / `{: index_also="..." }` on a heading.
render.mjs takes them off the heading before the HTML is written (and
fails on either attribute anywhere else); search.mjs attaches them to the
entry holding that heading, or to the page's own entry; the build fails
when two places claim one term as the main entry.

In all three client copies, one lunr field `index` (boost 1000) holds
each term as a single stemmed token, a secondary term with one more `_`.
Every run of up to four query words is looked up whole, at clause boost
5 for a main term and 1 for a secondary one, so a term matches only a
query that names all of it. The terms are also appended to the content,
so the all-words pass still finds a marked entry, and the field's BM25
average length is pinned at one term: empty on nearly every entry, its
real average was near zero and a marked match counted for almost nothing.
One field, not two: every lunr field costs a slot on every term, and two
fields plus a words field took 24 MB of heap where one takes 5 MB.

Entries: conditional compilation (Topic-Preprocessor, then Compiler
Constants), late binding (Data-Types#object, then CreateObject), 64-bit
compilation (Features/64bit).

Against intent-2: hit@1 94.15% -> 94.18%, prose 17 -> 20 of 20 at rank 1,
the `behind` expectation met, 0 queries worse, 3 better. Heap 287 -> 292
MB, index build +3-8%. Both clients checked in a browser against the
replica. Authoring.md documents the keys, and its note on which headings
get a search entry now says h3, as it has been since step 3.
`FileListBox.Name` ranked the FileListBox page first, and `Slider.KeyDown`,
documented under `KeyDown, KeyPress, KeyUp`, ranked 20th, behind every
other control's KeyDown. `qualified` held the right token, but at boost 50
it counted for less than a title naming the container.

In all three copies, `qualified` is now boost 500, and only a qualified
name reaches it:
- a plain word's trailing wildcard skips `qualified`, since at that weight
  `vbfile*` completing to every `vbfileattribute.*` pulled VbFileAttribute
  above the vbFile constant;
- two adjacent words, joined with a dot, are a term on `qualified` at
  clause boost 10 (`FileListBox Name`);
- in the all-words pass each word is REQUIRED over every text field at
  boost 0, deciding presence only (an entry may name its container
  nowhere but `qualified`), and scored by a second clause off `qualified`.
  Scored there, `form*` pushed Form's members above Form#events for
  `Form events`.

Against intent-2: hit@1 94.18% -> 99.46%, MRR .9627 -> .9972, qualified
hit@1 91.5% -> 99.8%, 423 queries better, 0 worse. Every qualified symbol
written as two words: hit@1 97.75% -> 99.82%, 106 better, 0 worse. No new
field or term, so search-data.json and heap are unchanged. Both clients
checked in a browser against the replica; test/search.test.mjs's
qualified-name guard fails on each of five mutations of the replica.

The vendor README also gets back a paragraph that had drifted below the
index-term section from the reader-intent one.
…essons

The baseline's ground truth is intent-2, not intent-1. The eval can't see
multi-word queries, so the notes say how the qualified-name round checked
them. Two lunr facts from that round: a plain word's wildcard reaches
whole qualified names, and a REQUIRED clause at boost 0 requires without
scoring. The order of the next steps is marked as this session's, not
the user's, where it is.
extractSections took any heading that read the same as the page's title
for the title, giving it the page's URL. On Shape and Timer, whose h1 is
"Shape class" / "Timer class", that was `### Shape` / `### Timer`, the
Shape and Timer properties, so neither had a #shape / #timer entry and
`Shape.Shape` and `Timer.Timer` found nothing.

Four pages' entries change: Shape and Timer, and New-Functions and
Challenges/1, whose second heading repeats the title. Those two now have
an empty page entry beside the section, as the 272 other pages whose h1
differs from their title already do.

Against the baseline: hit@1 99.46% -> 99.49%, hit@10 99.98% -> 100%,
2 queries better, 0 worse. test/search.test.mjs covers the rule and
fails on the old code.
`Printer.Font` and `Printer.Fonts` both stem to `printer.font`, so every
clause scored them alike and `Printer.Fonts` came second; so did
`Collection.Item`, `Global.Printers`, `OLE.Update`, `Report.Page` and both
`GetHeaders`, typed with a dot or as two words.

In all three copies, stemTwins() finds the qualified names whose stem
another shares (104, mostly a function and its `$` form), and
qualifiedField() appends each twin an entry holds to `qualified` whole,
as exactName() writes it. The query adds, on `qualified` at clause boost
10, every word with a dot and every two adjacent words joined with one,
written the same way.

Only the twins: every qualified name whole cost 19 MB of heap. And in
`qualified`, not `exact`: there the longer field pushed `InStrB` behind
its own section.

Against the baseline: hit@1 99.49% -> 99.58%, qualified hit@1 100%,
7 better, 0 worse. Every qualified name as two words: 99.86% -> 100%,
0 worse. Heap 292.8 -> 293.1 MB; search-data.json unchanged. Both
clients checked in a browser against the replica; test/search.test.mjs
fails on each of eight mutations.
Decided with the user. Where a symbol's URL is a page, any section of
that page now counts as well: `DefInt` is documented by the Deftype page,
and its section headed `DefBool, DefByte, DefInt, ...` lands the reader
on the same definition. Ground truth intent-3.

Exactly 24 queries move, all from rank 2 to 1 (the 14 Def* statements
and the B/W string functions); no other rank changes. hit@1 99.58% ->
99.88%, MRR .9978 -> .9993. `MidB$` stays a miss: its first result is
the `MidB =` statement, a different page.

WIP.Search.md also lists two probes to look at next: `New Functions`
and `Form events`.
The six rollout hashes were already stale from an earlier rewrite; all
sixteen now name commits on this branch. eval/search-experiments/README.md
named one of them too.
…ched

Any query with the word `a` threw inside lunr ("Cannot read properties of
undefined (reading '_index')"), in both clients and the replica: 103 of
3,658 titles did. lunr 2.3.9 minimises the index's token set by merging
nodes whose TokenSet#toString() match, and that key runs each edge's
label into its child's numeric id, so `{1 -> 656}` and `{1 -> 6, 5 -> 6}`
both key as `01656`. Merged, the token set held `amp;h80004001010`, which
no entry has, and lost `amp;h80004001` and `amp;h80004005`; a wildcard
clause reaching the invented word found no postings and threw. Whether
keys collide depends on the whole term set: f8e630e alone doesn't, and
this branch's fields shifted the ids until it did.

separateTokenSetKeys() replaces TokenSet#toString() with a key that ends
each id with `,`, installed once in all three copies beside the
tokenizer wrapper. The token set now equals the index's terms, no title
throws, and the eval is unchanged. Both clients checked in a browser
against the replica with no console errors. test/search.test.mjs takes
lunr's own toString() from lunr.min.js, checks a small fixture still
collides with it, and that both patched keys keep it exact; each of five
mutations fails it.

WIP.Search.md records the diagnosis of the `New Functions` and `Form
events` probes, and a new item: `&H80004005` can't be found, because the
search content keeps `&amp;`.
…d measured

Both are general. A page typed by its own multi-word title reaches rank 1
84.6% of the time (a one-word symbol entry wins, more so beside a kind
word), and `<page> <section>` (`DTPicker Properties`) only 22.7%, since
the page's `X class` heading has X in its title and the section only in
`page`. Measured with knobs, not shipped: a score x3 for a result whose
whole title, or page title plus title, is the multi-word query, with
plural kind words no longer making a name, takes those to 94.0% and
96.0% with nothing worse in the eval, the sets, or every qualified name
as two words.

Also recorded: multi-word queries with a short word take ~800 ms in the
replica (`a page`), unseen while `a` threw.

eval/search-experiments/probes/ keeps the scripts and the knob diff.
Ship both the whole-title re-rank and the plural kind-word rule, and add
the page-title and page-plus-section sets to the eval as ground truth.
The symbol queries are one word each, so the eval never saw a multi-word
query. Two sets derived from the build now judge them, as ground truth
intent-4:

- every page's own title of two or more words, typed as the reader sees
  it; any entry of that page counts (140 queries: 87.9% at rank 1);
- `<page title> <section title>` for the one-word section titles 20+
  pages share, such as `DTPicker Properties`; only that section counts
  (300 queries: 22.7% at rank 1).

The operator pages (`&, &=`) are left out: a reader types one operator,
which the bare-name set already measures. No ranking change: all 8,012
earlier queries keep their ranks. Baseline re-saved.
The search data keeps the page's HTML entities, and must: the results
panel inserts titles and content with innerHTML and highlights by
character position in that text. So decoding them in search-data.json
would let `&lt;Object&gt;` become a tag and shift every highlight after
an entity. But the index took them as words: `&amp;H80004005` was the
term `amp;h80004005`, and `&H80004005` found none of the five pages that
mention it (only an unrelated fuzzy match).

The tokenizer wrapper now decodes each token after the split
(decodeTokenEntities(): named and numeric references, lowercased), so the
token keeps its position in the escaped text. In all three copies: the
online client, offline.mjs's initSearch() and the replica.

  &H80004005        all five pages, nothing else (before: none)
  eval              unchanged (8,452 queries)
  spaced names      unchanged, 100% (5,108)
  index terms       26,136 -> 25,897; none holds an entity (507 did);
                    the token set still equals the terms

Both clients checked in a browser: the same five pages as the replica,
highlighted at `(&H80004005)`.
…choice

Records item 4 (e6237fe) with its numbers and why it is not a content
fix, and leaves items 5 and 6 open for the user to choose.
83c4dbb wrote the MidB$ sentence through a replacement string, and
$` pasted everything before it in its place.
lunr 2.3.9's Index#query gathers a REQUIRED clause's entries as a
running total, c = c.union(S), once per expanded term and field, and
Set#union copied both sets each time. The all-words pass requires every
word with a trailing wildcard, and `a*` reaches thousands of terms:
82% of the time went to union and the Set constructor.

accumulateSetUnions(), in all three copies, lets a set that union()
made take the next set in place, with lunr's own length.

Replica, per search: `a page` 809 -> 87 ms, `a p` 1334 -> 278 ms,
`Form events` 104 -> 52 ms. Chrome: `a page` 501 -> 120 ms, `a p`
1102 -> 235 ms. Every page and section title as a query (10,292):
718 s -> 86 s, 716 -> 7 over 200 ms, identical refs and scores. Eval
unchanged (8,452 queries).
…hing

A query naming one thing and its kind (`MaxHeight property`,
`Continue statement`) required the kind word in the all-words pass, and
a member's section rarely says "property" or "event": only 61.6% of
the 1,836 "<name> <kind>" queries found their symbol at rank 1, and
662 not at all. Now, if no entry found has the name in its title or as
its exact name, the pass runs again with the kind words optional (they
still score). A blanket "optional" made 22 worse (`Mid function`,
`Line statement`); an exact-name-only test made 4 worse.

"<name> <kind>" (eval/search-experiments/probes/kinds.mjs): hit@1
61.6% -> 91.1%, top 10 63.9% -> 94.9%, not found 662 -> 93; 577
better, none worse. Eval unchanged (8,452), qualified names as two
words unchanged (5,108). Checked in both clients against the replica.
Every symbol typed as its name and kind (`MaxHeight property`), for
the client's kind words less `sub` and `member`, which readers don't
say: 1,785 queries, 91.1% at rank 1 (about 62.6% before df5a34b).
Any symbol of that name and kind counts, and any section of a page
documenting one. KIND_WORDS is exported from the replica so the eval
and probes/kinds.mjs use its list. Baseline re-saved; the 8,452
earlier queries keep their ranks.
33 queries the user approved, with the shortened forms and variants
asked for (`register COM dll`, `register dll`, `inline initialization`,
`field initialization`, `typedecl char[acter]`, `type char[acter]`).
Before any entry: prose hit@1 47.2% (25 of 53), `behind` within 3 for
1 of 10. Baseline re-saved with these ranks; every other query unchanged.
The user approved all 16 recommended rows, with shortened forms. Terms
the page's own words already find first get no entry (`watches`,
`call stack`, `COM registration`, `inline initialization`). Left out,
and why:
- `declaration`, `comment`: the entry's stem is a bare name's (`Declare`,
  `Comments`), which it pushed to rank 2. The Glossary's definition is
  2nd / 5th without one, which the user ruled counts within the top 5;
  the two queries now accept it.
- `pointers`: the Pointers page is 3rd; the entry cost `Pointer` 6 -> 7
  and `Pointer field` 3 -> 4.
`access key` is a main entry on Label too, since CommandButton's
secondary entry alone put it first; `type char` has its own term, as
`type character` doesn't match it.

Prose hit@1 47.2% -> 94.3% (50 of 53), `behind` within 3 for 10 of 10;
27 queries better, none worse (eval 10,270), qualified names as two
words unchanged. Capitals and hyphens make no difference.
The user's choices from the rows left to them: 11 queries, with
`user types` added. Before any entry: 1 of 11 at rank 1. Baseline
re-saved; every other query unchanged.
…space

The rows left to the user, as decided: `user defined types` / `user
types` -> Type, then UDTs; `event handlers` -> Handlers, then the Forms
tutorial's handler step; `twinpack` -> Creating-TWINPACK, then the
Packages page, File-Format and Importing-TWINPACK; `class module`,
`enumerated constant` (then Enum), `module variable`, `inherited
property`, `z-order` (Form, then MDIForm and Report), `windows api
tutorial`. `namespaces`: a new Glossary definition, linking the
package, project, static-linking and AppObject senses, takes the main
entry, with the Packages page second.

Importing-TWINPACK takes `twinpack` as a secondary entry too: the
Packages page's alone put it above Importing's own title typed whole.
Type takes `user types` as a main entry, which it already ranked first
for, so UDTs' secondary entry doesn't overtake it.

Prose 10 of 11 -> all 11 at rank 1; `behind` within 3 for 17 of 17.
10 better, none worse (eval 10,281); qualified names as two words
unchanged. Checked in the browser against the replica.
The user chose the New Project dialog first, with Build Type and
Project-Types matched behind it. Before any entry: 1 of 3 at rank 1.
As the user chose: `standard exe`, `create an ActiveX DLL` and `create
ActiveX DLL` are main entries on the New Project dialog's options, and
secondary on Project Settings' Build Type and the Project-Types page.
Project-Types says it covers types beyond the traditional EXE and
ActiveX DLL, so it doesn't lead.

`pointers` is a secondary entry on the Pointers page, as the user asked
("at a lower priority"): the page goes 3 -> 1, and bare `Pointer`, which
a main entry pushed 6 -> 7, is untouched. It still costs `Pointer field`
3 -> 4; kept at the user's request.

Prose 62 of 67 -> 65 of 67 at rank 1, `behind` within 3 for 20 of 20;
3 better, 1 worse (`Pointer field`); qualified names as two words
unchanged. All gates pass.
Where it stands, the eval by category, the known misses, the four
candidates for next, and the rules for index entries; the criteria,
tools and lunr notes gain this session's rulings and traps.
…r fixes

Two surveys: the old content-gap list re-checked, and about 120 reader
terms run through the replica. 33 of the new queries miss rank 1 today;
8 are guards for spellings that already land (by reference, type
library, named arguments, manifest, DPI awareness, ...).

Prose hit@1: 65 of 67 -> 73 of 108. Nothing else moves (10,284
queries compared: 0 worse, 0 better).
New sections, each with a sample that compiles (check_build) and whose
printed values were confirmed with tbrun:
- Sub: "Passing arguments ByRef and ByVal" and "Optional arguments and
  default values".
- Function: "Returning more than one value" (ByRef parameters, a UDT,
  an array).
- Call: "Named arguments". A positional argument after a named one is
  TB5103, checked. The old example declared MessageBeep in the 16-bit
  "User" library and was marked inert; it is now a module that
  compiles.
Smaller polish: the Glossary's by reference, by value and named
argument link to these; the Project Explorer's manifest section says
what the manifest does; Library References says how to add one; the
Variables pane names VB6's Locals window.

Call's Example stays before Named arguments: the Example folds into
the page's top entry, which is where "Call statement" finds the word
"statement" (with Named arguments first, it fell 1 -> 5).

Prose hit@1: 73 -> 80 of 108 (ByRef, ByVal vs ByRef, application
manifest, variables window, optional arguments, default parameter and
argument values). 10,284 other queries: 0 worse. Link check: 0 broken.

    examples.bat --only "^Reference/Core/(Sub|Function|Call)\.md"
    check_examples: 9 sample(s), 9 compile, 0 finding(s), 10.4s -- clean
Main entries (index):
- Sub: ByRef, ByVal; optional parameters. Function: return multiple
  values, multiple return values. Call: named parameters.
- Glossary, type library: typelib. Project Explorer's manifest section:
  visual styles.
- Attributes: packing alignment. Categories, State Management: registry
  access.
- Project Settings: add reference (Library References), high DPI
  (Force DPI Awareness At Startup).
- IDE-Features, Theme System: dark mode.
- IDE pages: new project dialog; find and replace, search and replace;
  memory window/panel; variables panel, locals window; diagnostics
  window; history panel/window; outline view/window/panel.
Secondary (index_also): add reference on the Project menu's
References; dark mode on the Window menu's Theme.

Not shipped: `default property` on DefaultMember. As a main or a
secondary entry, it puts DefaultMember above CommandButton's Default
property for the name-and-kind query "Default property" (1 -> 2). Which
reading comes first is left for the user.

Prose hit@1: 80 -> 105 of 108 (left: declaration and comment, accepted
by the Glossary ruling, and default property); behind within 3: 22 of
22. 10,284 other queries: 0 worse. Every title as a query: no throws,
the same 9 empty and 2 misses as before. Link check: 0 broken; test.bat
and check.bat gates pass.
The resume section names the one open decision (default property), the
work left for later, and the tools this item used (tbrun for what a
sample prints, the every-title check). Baseline: 10,325 queries.
…s.mjs

Staging moved every eval tool onto lib/repo-paths.mjs; this one was
written on the branch before that.
Its target is Attributes' [DefaultMember]; a Glossary definition of the
term counts too, as for `by reference`. Measured before the definition
exists: unchanged, the query at 6.
Headed *default member* so it doesn't match the whole query `default
property`: headed *default property*, it took rank 1 (a whole title
scores x3) and pushed CommandButton's Default property to 2 for the
name-and-kind query `Default property`, the same trade the user
declined for an Attributes entry. As *default member* it lands at 5,
where a Glossary definition counts.

Eval: prose `default property` 6 -> 5; worse 0, better 1.
… set aside

Every branch hash remapped to its rebased commit (eval/search-experiments'
README too). The baseline saved again: staging's quoted `#` titles
turned the title query `Line Input` into three, 10,327 queries in all;
`Input #` at 2 is a new known miss.
@KubaO
KubaO force-pushed the claude/paintpicture-docs-runtime-f3250d branch from d383b7e to 086e467 Compare September 27, 2026 18:59
@KubaO
KubaO merged commit 77cc40a into twinbasic:main Sep 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant