Skip to content

Latest commit

 

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Subtext

Subtext is an open-source Python library for exploratory text analysis. It accepts ordinary Python data, extracts inspectable features, and keeps results linked to their source texts.

import subtext as st

Typical uses include comparing editions or translations, studying fictional dialogue, following a publication over time, and examining human- and machine-written samples.

Subtext 3 is alpha software. The core package has no runtime dependencies; Pandas output is optional. The API and saved-file formats may still change.

Install

From a checkout of this repository:

python -m pip install -e .

To use .to_pandas():

python -m pip install -e '.[pandas]'

Subtext requires Python 3.10 or later.

Quick start

This example compares two editions of the same passages. passage identifies matched texts. chapter is used to check the direction of a result in separate parts of the work.

import subtext as st

texts = [
    {
        "id": "old-1",
        "text": "It was quite impossible; nevertheless, she continued.",
        "edition": "old",
        "passage": 1,
        "chapter": "opening",
    },
    {
        "id": "new-1",
        "text": "It was impossible, but she continued.",
        "edition": "new",
        "passage": 1,
        "chapter": "opening",
    },
    {
        "id": "old-2",
        "text": "He was quite certain; nevertheless, he waited.",
        "edition": "old",
        "passage": 2,
        "chapter": "ending",
    },
    {
        "id": "new-2",
        "text": "He was certain, but he waited.",
        "edition": "new",
        "passage": 2,
        "chapter": "ending",
    },
]

feature_set = st.features.style() + st.features.terms(ngram_range=1)

report = st.compare(
    texts,
    by="edition",
    groups=("old", "new"),
    paired_by="passage",
    across="chapter",
    features=feature_set,
    min_document_frequency=1,
)

print(report)

print(report) gives a compact table. This is the output from the example above; the report warns that two pairs are not enough for a serious analysis.

ComparisonReport('old' vs 'new', by='edition', n_features=41)
feature                         method                   estimate a  estimate b  effect size  p value  q value  context consistency
------------------------------  -----------------------  ----------  ----------  -----------  -------  -------  -------------------
token_1gram_count:but           paired_permutation_rate  0           166.7       —            0.5      1        1
token_1gram_count:nevertheless  paired_permutation_rate  142.9       0           —            0.5      1        1
token_1gram_count:quite         paired_permutation_rate  142.9       0           —            0.5      1        1
token_1gram_count:he            paired_permutation_rate  142.9       166.7       0            1        1        1
token_1gram_count:was           paired_permutation_rate  142.9       166.7       —            0.5      1        1
length:characters               paired_permutation       49.5        33.5        —            0.5      1        1
sentences:mean_characters       paired_permutation       49.5        33.5        —            0.5      1        1
token_1gram_count:certain       paired_permutation_rate  71.43       83.33       0            1        1        1
… 33 more rows
Warnings:
  - A group has fewer than five texts; estimates are unstable.
  - Only 2 matched 'passage' units were available.

Feature names

Feature names use family:name. In sentences:mean_tokens, sentences is the family and mean_tokens is the measurement. It is not a regular expression. The full name avoids collisions between measurements from different families. Subtext stores the name as a FeatureKey and looks up its registered label. A short name such as mean_tokens also works when it is unique.

feature(), documents(), outliers(), influence(), and by_context() all resolve a feature name:

name = "sentences:mean_tokens"

row = report.feature(name)       # the comparison row
sources = report.documents(name) # source texts and measured values
unusual = report.outliers(name)  # texts far from their group mean
influence = report.influence(name)  # leave-one-out sensitivity
chapters = report.by_context(name)  # the result split by chapter

For example, print(chapters) produces:

ContextReport(feature='sentences:mean_tokens', across='chapter')
context  n a  n b  estimate a  estimate b  difference
-------  ---  ---  ----------  ----------  ----------
ending   1    1    7           6           1
opening  1    1    7           6           1

concordance() is separate. It searches the source text, literally by default or as a regular expression when regex=True:

hits = report.concordance("nevertheless")
print(hits)
ConcordanceResult(query='nevertheless', n_hits=2)
text id  left                      match         right
-------  ------------------------  ------------  ----------------
old-1    It was quite impossible;  nevertheless  , she continued.
old-2    He was quite certain;     nevertheless  , he waited.

Input

Analysis functions accept data directly:

st.profile("One text.")
st.profile(["First text.", "Second text."])
st.profile([{"text": "One record.", "speaker": "A"}], by="speaker")
st.profile(dataframe, by="section")
st.profile(huggingface_dataset, by="label")
st.profile(pathlib.Path("article.txt"))

Supported inputs are:

Input Interpretation
str One literal text. A string is never treated as a filename.
iterable of strings One text per item. Stable IDs are generated.
iterable of mappings text, optional id and tokens, plus metadata.
dataframe-like table Rows obtained through to_dict(orient="records").
Polars-like table Rows obtained through to_dicts().
Arrow-like table Rows obtained through to_pylist().
dataset or row iterator Mapping records consumed without a hard dependency.
pathlib.Path One UTF-8 text file.
TextRecord / TextCollection An already prepared Subtext input.

Use st.collect() when column names differ or when the same texts will be used more than once:

texts = st.collect(dataframe, text="body", id="article_id")
subset = texts.where(publication="Example Daily")

print(texts)
print(texts.metadata_fields)
print(texts.missing_count)

TextCollection stores immutable TextRecord objects with text boundaries, IDs, metadata, missing values, and optional tokens. It does not imply that the data is balanced or representative.

Printing it shows the collection and a short preview:

TextCollection(n_texts=4, missing=0, metadata=[chapter, edition, passage])
  old-1: 'It was quite impossible; nevertheless, she continued.'
  new-1: 'It was impossible, but she continued.'
  old-2: 'He was quite certain; nevertheless, he waited.'
  new-2: 'He was certain, but he waited.'

Converting a collection

A collection can be converted back to ordinary records or to a Pandas DataFrame:

records = texts.to_records()
records_with_tokens = texts.to_records(include_tokens=True)
frame = texts.to_pandas()

to_records() returns a list[dict], which is also the common route into other table libraries. Since TextCollection is a sequence, list(texts), tuple(texts), indexing, and slicing also work; those forms contain TextRecord objects rather than dictionaries.

polars_frame = polars.DataFrame(texts.to_records())
arrow_table = pyarrow.Table.from_pylist(texts.to_records())
huggingface_data = datasets.Dataset.from_list(texts.to_records())

Subtext does not currently provide separate to_polars(), to_arrow(), or to_dataset() methods. These constructors keep those libraries optional.

Tokenization

Only token-based features require a tokenizer. Raw character counts, punctuation, regular expressions, and character n-grams do not tokenize the input. When token features are combined, they share one tokenization per text.

The default tokenizer is intentionally simple. Supply a callable when word boundaries need domain- or language-specific treatment:

def whitespace_tokens(text):
    return text.split()

table = st.extract(
    texts,
    st.features.lexical(),
    tokenizer=whitespace_tokens,
)

Precomputed tokens can be strings or mappings with offsets:

prepared = [
    {
        "id": "x1",
        "text": "New York changed.",
        "tokens": [
            {"text": "New York", "start": 0, "end": 8},
            {"text": "changed", "start": 9, "end": 16},
        ],
    }
]

st.extract(prepared, st.features.lexical())

Features

Feature bundles are FeatureSet objects and can be added together:

Bundle Included measurements
st.features.surface() characters, lines, paragraphs, punctuation
st.features.lexical() token counts, TTR, MATTR, MTLD, HD-D, repetition
st.features.structure() sentence count and sentence-length distributions
st.features.style() surface + lexical + structure
st.features.terms() token n-gram counts and their opportunities
st.features.function_words() counts for a small built-in English list

style() does not add thousands of n-gram columns unless requested:

features = st.features.style(
    character_ngrams=True,
    include_function_words=True,
) + st.features.terms(ngram_range=(1, 2))

Individual extractors are also public:

features = st.FeatureSet([
    st.TextLength(),
    st.Punctuation(normalize="count"),
    st.LexicalDiversity(mattr_window=50),
    st.RegexMatches("hedges", r"\b(?:perhaps|possibly|may)\b"),
    st.DictionaryMatches(
        "certainty",
        ("clearly", "certainly", "undoubtedly"),
        separate=True,
    ),
    st.CharacterNgrams(ngram_range=4),
])

Feature keys include a family, name, and semantic kind. Counts also carry an exposure—the number of opportunities for that count—so analysis code does not have to infer a denominator from a column name.

Custom features

A custom extractor declares its input, returns feature values, and records its configuration:

class UppercaseShare:
    requires = frozenset({"text"})

    def extract(self, context):
        letters = [char for char in context.text if char.isalpha()]
        share = (
            sum(char.isupper() for char in letters) / len(letters)
            if letters else 0.0
        )
        return {st.FeatureKey("case", "uppercase_share", "rate"): share}

    def config(self):
        return {"feature": "UppercaseShare"}

table = st.extract(texts, UppercaseShare())

Custom count features must return st.Measurement(value, exposure). Subtext checks that both numbers are finite and that the exposure is non-negative.

Analysis

Function Question it addresses Result
st.measure() What does one text measure? one sparse feature row
st.extract() What does every text measure? FeatureTable
st.profile() What is typical within each group? ProfileReport
st.compare() How do two groups differ? ComparisonReport
st.trend() How does a feature move over ordered periods? TrendReport
st.collocations() Which adjacent words occur together unusually often? CollocationReport
st.concordance() Where does a literal or regular expression occur? ConcordanceResult

Examples:

st.profile(texts, by="speaker", features=st.features.style())

st.compare(
    texts,
    by="translation",
    paired_by="passage",
    across="chapter",
    features=st.features.style(),
)

st.trend(
    texts,
    over="year",
    by="publication",
    features=st.features.terms(ngram_range=1),
)

st.collocations(
    texts,
    min_frequency=5,
    min_document_frequency=3,
    order_by="log_dice",
)

st.concordance(texts, r"\bhowever\b", regex=True, where={"speaker": "A"})

An existing FeatureTable can be reused by profile(), compare(), and trend(); do not pass features again in that case.

What compare() calculates

Subtext chooses a method from the feature semantics and the study design:

Data Default method Reported effect
unpaired counts smoothed binomial log-odds with a normal approximation log-odds difference
paired counts paired mean-difference permutation test on per-opportunity rates paired Hedges' g
unpaired scalar/rate label-permutation test of the mean difference Hedges' g
paired scalar/rate sign-permutation test of paired differences paired Hedges' g

Small permutation spaces are enumerated exactly; larger ones are sampled. Setting n_resamples=0 uses paired or Welch t tests instead. Confidence intervals are percentile bootstrap intervals. P-values are adjusted across the report with the Benjamini–Hochberg procedure by default.

Each comparison row includes group estimates, the raw counts and exposures when available, effect size, test statistic, p- and q-values, interval, document coverage, normalized dispersion, cross-context direction consistency, and the largest leave-one-out shift.

A large effect can still be uncertain. A small p-value can describe a trivial effect. A feature found mainly in one text is fragile. The report exposes the source records needed to check these cases.

Research basis and implementation notes

These publications describe methods used or adapted by Subtext. They do not imply that the authors reviewed or endorse this package.

The lexical-diversity scores depend on what counts as a token and a type. They should not be compared across different tokenization, case-folding, or sample construction rules. None of the cited measures turns an observational text collection into a controlled experiment.

Output and provenance

Feature tables and reports export plain Python records. Pandas conversion is available through the optional extra.

table = st.extract(texts, st.features.style())

records = table.to_records()
frame = table.to_pandas()
ids, columns, matrix = table.to_dense()
ids, columns, sparse_rows = table.to_sparse_rows()

Extraction records the collection fingerprint, source adapter, tokenizer, sentence splitter, feature configuration, document-frequency threshold, and whether supplied tokens were used:

print(table.processing)

Tables and reports can be saved as JSON envelopes:

table.save("features.stx")
loaded_table = st.FeatureTable.load("features.stx")
assert loaded_table.verify(texts)

report.save("comparison.stx")
snapshot = st.load("comparison.stx")
assert snapshot.valid()
assert snapshot.verify(texts)

The checksum detects accidental changes or edits made without recomputing it. It is not a digital signature and does not establish who created a file. Loading a snapshot does not execute stored code.

Limits

  • Subtext offers measurements and diagnostics, not causal conclusions, authorship proof, plagiarism findings, or model attribution.
  • The default tokenizer and sentence splitter are lightweight heuristics. Use preprocessing suited to the language and domain.
  • Statistical inference assumes that the analysed units match the design. Repeated passages, authors, conversations, or publications may require pairing or a model that Subtext does not provide.
  • Multiple-testing correction does not repair biased sampling, confounding, inappropriate features, or dependent observations.
  • Collocation scores are unstable for rare terms. Use both frequency and document-frequency thresholds, then read the concordance lines.
  • Trend slopes describe the supplied periods; they do not model nonlinear change, autocorrelation, or uncertainty in the slope.

Legacy

The original Signetar/Subtext history, tags, and remote are retained. Subtext 2.0.3's main.py is preserved verbatim under subtext.legacy, including its old behaviour and quirks:

from subtext import legacy

legacy.__version__

The legacy module is an archive and compatibility layer, not the recommended API for new analysis.

License

Subtext is licensed under the MIT License.

About

Subtext is a framework used to assist natural language processing. It provides useful tools and advanced algorithms for processing human language.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Contributors

Languages