Subtext is an open-source Python library for exploratory text analysis. It accepts ordinary Python data, extracts inspectable features, and keeps results linked to their source texts.
import subtext as stTypical uses include comparing editions or translations, studying fictional dialogue, following a publication over time, and examining human- and machine-written samples.
Subtext 3 is alpha software. The core package has no runtime dependencies; Pandas output is optional. The API and saved-file formats may still change.
From a checkout of this repository:
python -m pip install -e .To use .to_pandas():
python -m pip install -e '.[pandas]'Subtext requires Python 3.10 or later.
This example compares two editions of the same passages. passage identifies
matched texts. chapter is used to check the direction of a result in separate
parts of the work.
import subtext as st
texts = [
{
"id": "old-1",
"text": "It was quite impossible; nevertheless, she continued.",
"edition": "old",
"passage": 1,
"chapter": "opening",
},
{
"id": "new-1",
"text": "It was impossible, but she continued.",
"edition": "new",
"passage": 1,
"chapter": "opening",
},
{
"id": "old-2",
"text": "He was quite certain; nevertheless, he waited.",
"edition": "old",
"passage": 2,
"chapter": "ending",
},
{
"id": "new-2",
"text": "He was certain, but he waited.",
"edition": "new",
"passage": 2,
"chapter": "ending",
},
]
feature_set = st.features.style() + st.features.terms(ngram_range=1)
report = st.compare(
texts,
by="edition",
groups=("old", "new"),
paired_by="passage",
across="chapter",
features=feature_set,
min_document_frequency=1,
)
print(report)print(report) gives a compact table. This is the output from the example
above; the report warns that two pairs are not enough for a serious analysis.
ComparisonReport('old' vs 'new', by='edition', n_features=41)
feature method estimate a estimate b effect size p value q value context consistency
------------------------------ ----------------------- ---------- ---------- ----------- ------- ------- -------------------
token_1gram_count:but paired_permutation_rate 0 166.7 — 0.5 1 1
token_1gram_count:nevertheless paired_permutation_rate 142.9 0 — 0.5 1 1
token_1gram_count:quite paired_permutation_rate 142.9 0 — 0.5 1 1
token_1gram_count:he paired_permutation_rate 142.9 166.7 0 1 1 1
token_1gram_count:was paired_permutation_rate 142.9 166.7 — 0.5 1 1
length:characters paired_permutation 49.5 33.5 — 0.5 1 1
sentences:mean_characters paired_permutation 49.5 33.5 — 0.5 1 1
token_1gram_count:certain paired_permutation_rate 71.43 83.33 0 1 1 1
… 33 more rows
Warnings:
- A group has fewer than five texts; estimates are unstable.
- Only 2 matched 'passage' units were available.
Feature names use family:name. In sentences:mean_tokens, sentences is the
family and mean_tokens is the measurement. It is not a regular expression.
The full name avoids collisions between measurements from different families.
Subtext stores the name as a FeatureKey and looks up its registered label. A
short name such as mean_tokens also works when it is unique.
feature(), documents(), outliers(), influence(), and by_context() all
resolve a feature name:
name = "sentences:mean_tokens"
row = report.feature(name) # the comparison row
sources = report.documents(name) # source texts and measured values
unusual = report.outliers(name) # texts far from their group mean
influence = report.influence(name) # leave-one-out sensitivity
chapters = report.by_context(name) # the result split by chapterFor example, print(chapters) produces:
ContextReport(feature='sentences:mean_tokens', across='chapter')
context n a n b estimate a estimate b difference
------- --- --- ---------- ---------- ----------
ending 1 1 7 6 1
opening 1 1 7 6 1
concordance() is separate. It searches the source text, literally by default
or as a regular expression when regex=True:
hits = report.concordance("nevertheless")
print(hits)ConcordanceResult(query='nevertheless', n_hits=2)
text id left match right
------- ------------------------ ------------ ----------------
old-1 It was quite impossible; nevertheless , she continued.
old-2 He was quite certain; nevertheless , he waited.
Analysis functions accept data directly:
st.profile("One text.")
st.profile(["First text.", "Second text."])
st.profile([{"text": "One record.", "speaker": "A"}], by="speaker")
st.profile(dataframe, by="section")
st.profile(huggingface_dataset, by="label")
st.profile(pathlib.Path("article.txt"))Supported inputs are:
| Input | Interpretation |
|---|---|
str |
One literal text. A string is never treated as a filename. |
| iterable of strings | One text per item. Stable IDs are generated. |
| iterable of mappings | text, optional id and tokens, plus metadata. |
| dataframe-like table | Rows obtained through to_dict(orient="records"). |
| Polars-like table | Rows obtained through to_dicts(). |
| Arrow-like table | Rows obtained through to_pylist(). |
| dataset or row iterator | Mapping records consumed without a hard dependency. |
pathlib.Path |
One UTF-8 text file. |
TextRecord / TextCollection |
An already prepared Subtext input. |
Use st.collect() when column names differ or when the same texts will be used
more than once:
texts = st.collect(dataframe, text="body", id="article_id")
subset = texts.where(publication="Example Daily")
print(texts)
print(texts.metadata_fields)
print(texts.missing_count)TextCollection stores immutable TextRecord objects with text boundaries,
IDs, metadata, missing values, and optional tokens. It does not imply that the
data is balanced or representative.
Printing it shows the collection and a short preview:
TextCollection(n_texts=4, missing=0, metadata=[chapter, edition, passage])
old-1: 'It was quite impossible; nevertheless, she continued.'
new-1: 'It was impossible, but she continued.'
old-2: 'He was quite certain; nevertheless, he waited.'
new-2: 'He was certain, but he waited.'
A collection can be converted back to ordinary records or to a Pandas DataFrame:
records = texts.to_records()
records_with_tokens = texts.to_records(include_tokens=True)
frame = texts.to_pandas()to_records() returns a list[dict], which is also the common route into
other table libraries. Since TextCollection is a sequence, list(texts),
tuple(texts), indexing, and slicing also work; those forms contain
TextRecord objects rather than dictionaries.
polars_frame = polars.DataFrame(texts.to_records())
arrow_table = pyarrow.Table.from_pylist(texts.to_records())
huggingface_data = datasets.Dataset.from_list(texts.to_records())Subtext does not currently provide separate to_polars(), to_arrow(), or
to_dataset() methods. These constructors keep those libraries optional.
Only token-based features require a tokenizer. Raw character counts, punctuation, regular expressions, and character n-grams do not tokenize the input. When token features are combined, they share one tokenization per text.
The default tokenizer is intentionally simple. Supply a callable when word boundaries need domain- or language-specific treatment:
def whitespace_tokens(text):
return text.split()
table = st.extract(
texts,
st.features.lexical(),
tokenizer=whitespace_tokens,
)Precomputed tokens can be strings or mappings with offsets:
prepared = [
{
"id": "x1",
"text": "New York changed.",
"tokens": [
{"text": "New York", "start": 0, "end": 8},
{"text": "changed", "start": 9, "end": 16},
],
}
]
st.extract(prepared, st.features.lexical())Feature bundles are FeatureSet objects and can be added together:
| Bundle | Included measurements |
|---|---|
st.features.surface() |
characters, lines, paragraphs, punctuation |
st.features.lexical() |
token counts, TTR, MATTR, MTLD, HD-D, repetition |
st.features.structure() |
sentence count and sentence-length distributions |
st.features.style() |
surface + lexical + structure |
st.features.terms() |
token n-gram counts and their opportunities |
st.features.function_words() |
counts for a small built-in English list |
style() does not add thousands of n-gram columns unless requested:
features = st.features.style(
character_ngrams=True,
include_function_words=True,
) + st.features.terms(ngram_range=(1, 2))Individual extractors are also public:
features = st.FeatureSet([
st.TextLength(),
st.Punctuation(normalize="count"),
st.LexicalDiversity(mattr_window=50),
st.RegexMatches("hedges", r"\b(?:perhaps|possibly|may)\b"),
st.DictionaryMatches(
"certainty",
("clearly", "certainly", "undoubtedly"),
separate=True,
),
st.CharacterNgrams(ngram_range=4),
])Feature keys include a family, name, and semantic kind. Counts also carry an exposure—the number of opportunities for that count—so analysis code does not have to infer a denominator from a column name.
A custom extractor declares its input, returns feature values, and records its configuration:
class UppercaseShare:
requires = frozenset({"text"})
def extract(self, context):
letters = [char for char in context.text if char.isalpha()]
share = (
sum(char.isupper() for char in letters) / len(letters)
if letters else 0.0
)
return {st.FeatureKey("case", "uppercase_share", "rate"): share}
def config(self):
return {"feature": "UppercaseShare"}
table = st.extract(texts, UppercaseShare())Custom count features must return st.Measurement(value, exposure). Subtext
checks that both numbers are finite and that the exposure is non-negative.
| Function | Question it addresses | Result |
|---|---|---|
st.measure() |
What does one text measure? | one sparse feature row |
st.extract() |
What does every text measure? | FeatureTable |
st.profile() |
What is typical within each group? | ProfileReport |
st.compare() |
How do two groups differ? | ComparisonReport |
st.trend() |
How does a feature move over ordered periods? | TrendReport |
st.collocations() |
Which adjacent words occur together unusually often? | CollocationReport |
st.concordance() |
Where does a literal or regular expression occur? | ConcordanceResult |
Examples:
st.profile(texts, by="speaker", features=st.features.style())
st.compare(
texts,
by="translation",
paired_by="passage",
across="chapter",
features=st.features.style(),
)
st.trend(
texts,
over="year",
by="publication",
features=st.features.terms(ngram_range=1),
)
st.collocations(
texts,
min_frequency=5,
min_document_frequency=3,
order_by="log_dice",
)
st.concordance(texts, r"\bhowever\b", regex=True, where={"speaker": "A"})An existing FeatureTable can be reused by profile(), compare(), and
trend(); do not pass features again in that case.
Subtext chooses a method from the feature semantics and the study design:
| Data | Default method | Reported effect |
|---|---|---|
| unpaired counts | smoothed binomial log-odds with a normal approximation | log-odds difference |
| paired counts | paired mean-difference permutation test on per-opportunity rates | paired Hedges' g |
| unpaired scalar/rate | label-permutation test of the mean difference | Hedges' g |
| paired scalar/rate | sign-permutation test of paired differences | paired Hedges' g |
Small permutation spaces are enumerated exactly; larger ones are sampled.
Setting n_resamples=0 uses paired or Welch t tests instead. Confidence
intervals are percentile bootstrap intervals. P-values are adjusted across the
report with the Benjamini–Hochberg procedure by default.
Each comparison row includes group estimates, the raw counts and exposures when available, effect size, test statistic, p- and q-values, interval, document coverage, normalized dispersion, cross-context direction consistency, and the largest leave-one-out shift.
A large effect can still be uncertain. A small p-value can describe a trivial effect. A feature found mainly in one text is fragile. The report exposes the source records needed to check these cases.
These publications describe methods used or adapted by Subtext. They do not imply that the authors reviewed or endorse this package.
- MATTR uses the moving-window type-token ratio described by Covington and McFall (2010), “Cutting the Gordian Knot: The Moving-Average Type–Token Ratio”. Subtext uses a default window of 50 tokens and shortens the window when the whole text is shorter.
- MTLD and HD-D follow the definitions studied by McCarthy and Jarvis (2010), “MTLD, vocd-D, and HD-D: A Validation Study of Sophisticated Approaches to Lexical Diversity Assessment”. MTLD averages forward and reverse passes with a default threshold of 0.72. HD-D uses the direct hypergeometric probability with a default sample size of 42.
- Unpaired count comparison is informed by the regularized log-odds approach in Monroe, Colaresi, and Quinn (2008), “Fightin' Words: Lexical Feature Selection and Evaluation for Identifying the Content of Political Conflict”. It is not an exact implementation of that paper: Subtext treats each feature as a binomial rate and adds a Jeffreys baseline plus a pooled empirical Beta prior. Its interval and p-value use a normal approximation to that smoothed log-odds estimate.
- Permutation tests use label permutations for independent groups and sign permutations for matched differences, in the tradition of Pitman (1937), “Significance Tests Which May Be Applied to Samples from Any Populations”. Monte Carlo paths add one to both the numerator and denominator following Phipson and Smyth (2010), “Permutation P-values Should Never Be Zero”.
- Bootstrap intervals use the basic percentile procedure, not BCa or a studentized interval. See Efron (1981), “Nonparametric Standard Errors and Confidence Intervals”.
- Hedges' g uses the small-sample bias correction from Hedges (1981), “Distribution Theory for Glass's Estimator of Effect Size and Related Estimators”. In paired analyses, Subtext standardizes the mean of the within-pair differences by their sample standard deviation before applying the correction.
- Benjamini–Hochberg adjustment uses the step-up false-discovery-rate procedure from Benjamini and Hochberg (1995), “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing”. Text features are often dependent, so the original paper's assumptions still matter.
- Welch's t test, used only when resampling is disabled for an unpaired scalar or rate, follows Welch (1947), “The Generalization of ‘Student's’ Problem When Several Different Population Variances Are Involved”.
- PMI for collocations follows the word-association treatment of Church and Hanks (1990), “Word Association Norms, Mutual Information, and Lexicography”. Subtext counts strictly adjacent tokens. For n-grams longer than two, it uses the product of the unigram marginals as the independence baseline.
- NPMI uses the normalization described by Bouma (2009), “Normalized (Pointwise) Mutual Information in Collocation Extraction”.
- logDice follows Rychlý (2008), “A Lexicographer-Friendly Association
Score”, for bigrams. The
package's
n > 2form is a direct extension that replaces the factor 2 and two marginal counts withnand the sum of all component marginals. - Collocation t-score uses the observed-minus-expected count divided by the square root of the observed count, as presented by Church et al. (1991), “Using Statistics in Lexical Analysis”.
- Dispersion is Shannon entropy divided by the maximum entropy for the observed number of texts. See Shannon (1948), “A Mathematical Theory of Communication”.
The lexical-diversity scores depend on what counts as a token and a type. They should not be compared across different tokenization, case-folding, or sample construction rules. None of the cited measures turns an observational text collection into a controlled experiment.
Feature tables and reports export plain Python records. Pandas conversion is available through the optional extra.
table = st.extract(texts, st.features.style())
records = table.to_records()
frame = table.to_pandas()
ids, columns, matrix = table.to_dense()
ids, columns, sparse_rows = table.to_sparse_rows()Extraction records the collection fingerprint, source adapter, tokenizer, sentence splitter, feature configuration, document-frequency threshold, and whether supplied tokens were used:
print(table.processing)Tables and reports can be saved as JSON envelopes:
table.save("features.stx")
loaded_table = st.FeatureTable.load("features.stx")
assert loaded_table.verify(texts)
report.save("comparison.stx")
snapshot = st.load("comparison.stx")
assert snapshot.valid()
assert snapshot.verify(texts)The checksum detects accidental changes or edits made without recomputing it. It is not a digital signature and does not establish who created a file. Loading a snapshot does not execute stored code.
- Subtext offers measurements and diagnostics, not causal conclusions, authorship proof, plagiarism findings, or model attribution.
- The default tokenizer and sentence splitter are lightweight heuristics. Use preprocessing suited to the language and domain.
- Statistical inference assumes that the analysed units match the design. Repeated passages, authors, conversations, or publications may require pairing or a model that Subtext does not provide.
- Multiple-testing correction does not repair biased sampling, confounding, inappropriate features, or dependent observations.
- Collocation scores are unstable for rare terms. Use both frequency and document-frequency thresholds, then read the concordance lines.
- Trend slopes describe the supplied periods; they do not model nonlinear change, autocorrelation, or uncertainty in the slope.
The original Signetar/Subtext history,
tags, and remote are retained. Subtext 2.0.3's main.py is preserved verbatim
under subtext.legacy, including its old behaviour and quirks:
from subtext import legacy
legacy.__version__The legacy module is an archive and compatibility layer, not the recommended API for new analysis.
Subtext is licensed under the MIT License.