gavinbowden.me — home

Lexical Substitution Pipeline

Given a sentence and a target word, this pipeline generates candidate single-word substitutes from four sources, filters them through spaCy syntax and WordNet checks, scores them with Sentence-BERT, BERT fluency, and a learned ranker, optionally reranks with a cross-encoder, and re-inflects the winner to match the target's tense and number. Then I evaluated its ranking on the SWORDS benchmark, and the evaluation turned out to be the interesting part: my headline number was partly scored on training data, one model had quietly collapsed into a constant, and the most expensive stage in the pipeline makes it worse.

Not a thesaurus lookup

Lexical substitution sounds like a thesaurus lookup right up until you make it context-aware. Take "falling" in "His soul swooned slowly as he heard the snow falling faintly." You want "drifting" or "settling". You do not want "failing" or "decreasing", and you want it conjugated to match: "drifting", not "drift".

So why not hand the whole job to one good model? Because no single signal can be trusted on its own. A masked language model will happily suggest a word meaning the exact opposite. A thesaurus will hand you a synonym that's the wrong part of speech for this sentence. A sentence embedding will call two sentences similar even when the swapped word doesn't inflect correctly.

So each stage exists to catch what the stage before it can't see (kinda like the Magi in Neon Genesis Evangelion!).

Some real runs, straight from the repo's example outputs. On that Joyce line, the pipeline ranks settling, descending, dropping, melting, drifting. On "The horror! The horror!" it offers fear, fright, scare, nightmare. And on "Call me Ishmael." its one surviving answer is "Leave".

"Leave me Ishmael" is grammatical, confident, and a completely different novel.

Every candidate carries six separate scores (lexical-resource match, target-word similarity, sentence-level similarity, MLM fluency, a learned validity score, and a cross-encoder rerank) that get combined into the final order.

Six signals, because one can't be trusted

  • Candidate generation pulls from four independent sources: WordNet (synonyms plus similar-to, verb-group, hypernym, hyponym, and derivationally related forms, from only the top three senses as picked by SBERT similarity between the sentence and each sense's gloss), SWORDS used purely as a lexical resource (never as ground truth), FLAN-T5 prompted up to eight different ways, and the top 150 BERT masked-language-model completions at the target's position. Each source gets its own prior, from 1.0 down to 0.25, because a WordNet hit deserves more trust than a T5 guess. A mode flag (full/balanced/fast/resources/mlm/t5) picks which sources run, since the full T5 prompt sweep is by far the slowest part. Balanced mode uses three prompts instead of eight.
  • Linguistic filtering runs every candidate back through spaCy in context. It gets dropped if it's already in the sentence, a stopword, a named entity, shares a lemma or substring with the target, is a WordNet antonym (including one hop out through similar-to synsets, not just direct antonym pairs), or comes back the wrong part of speech once it's re-inflected and re-parsed. For verbs specifically, the dependency-parsed argument frame (subject, object, indirect object, preposition, particle) has to be a subset of the candidate's own frame, with a WordNet double-object check for ditransitive verbs, so a verb needing a direct object doesn't get swapped for one that can't take one.
  • Semantic filtering scores what survives: Sentence-BERT cosine similarity between the original and substituted sentence, an optional bidirectional NLI contradiction check, and a supervised substitute-validity model (RoBERTa with a regression head, fine-tuned on SWORDS' soft human-agreement labels). That last one has a story, below. There's also a WiC-style sense classifier in the code that never got wired into the pipeline, which I'm counting as a future feature and not a lie.
  • Ranking combines those signals as a weighted sum with multiplicative penalties for candidates under the semantic/target/MLM/validity thresholds, rather than hard-cutting them (a candidate can survive a bad score, just discounted). The MLM fluency score is 1/log2(rank + 1) of the candidate's rank across BERT's whole vocabulary, not a raw probability. The alternative is a gradient-boosted ranker (HistGradientBoostingRegressor) trained on SWORDS labels over the five raw scores plus six engineered interaction and gap features. The CLI defaults to the weighted sum; the evaluation runs use the learned ranker.
  • A cross-encoder reranks the top 24 candidates and gets blended back into the final score. In theory it's the most powerful single signal, since it attends over the whole sentence pair instead of pooling to embeddings first. In practice, see below.
  • Morphology re-inflects the winning candidate to match the target's tense, number, and degree via lemminflect, with hand-written fallback rules for plurals, -ing forms, past tense, third-person singular, and comparative/superlative, plus a regex guard against the classic double-suffix bug where an inflector hands back "greaterer" or "classeses".

How do you grade a synonym?

SWORDS is a lexical-substitution benchmark of real sentences where crowdworkers scored a large candidate pool per target word, so evaluation isn't "did it guess the one right answer" but "how well does its ranking agree with a distribution of human judgments." To be precise about what I measured: the evaluation hands the pipeline SWORDS' own candidate lists and scores how well it ranks them. It measures ranking, not end-to-end generation. I scored all 370 dev-split targets (22,978 candidates) with NDCG@k, MAP@k, precision@k, and pairwise accuracy (the fraction of gold-scored candidate pairs ordered correctly, independent of k), all implemented by hand with per-POS and per-k breakdowns.

Two numbers from the dataset shaped decisions before any modeling happened. 82.4% of SWORDS' candidates are single words, which justified scoping this to single-word substitution instead of chasing multi-word paraphrases. And the mean gold score across all candidates is 0.111: most candidates are mediocre-to-bad even by crowdworker judgment, so a ranker has to actually find the few good ones.

Line chart of NDCG, MAP, and precision at k from 1 to 10, all decreasing as k grows
Precision@k drops from 0.58 at k=1 to 0.26 at k=10 mechanically. SWORDS gold sets are small, so requiring 10 returns caps precision once you run out of genuinely good substitutes. NDCG stays near 0.53 across k because it's rank-weighted rather than a raw hit count.
Bar chart comparing NDCG, MAP, and precision at k=10 across VERB, NOUN, ADJ, and ADV target words
Verbs are the hardest part of speech for this pipeline (NDCG 0.505, precision 0.244), plausibly the cost of the strict subcategorization-frame filtering, on top of verbs just carrying more sense ambiguity than nouns or adjectives.

Results, honestly (I leaked my own test set)

The first number I reported was NDCG@10 of 0.531. Then, writing this up, I went back to check which split the learned ranker had actually been trained on.

SWORDS dev. The same 370 targets I was evaluating it on.

With an 80/20 split grouped by target, that means 296 of the 370 targets behind my headline number were training data. On the 74 held-out targets the ranker never saw, the honest numbers are NDCG@10 0.488, MAP@10 0.346, precision@10 0.243, and pairwise accuracy 0.608. That's the number on the receipt now.

The substitute-validity model had a quieter problem. Across 840 example outputs, its scores range from 0.110907 to 0.110915.

That is not a signal. That is a constant, and specifically it is the dataset's mean gold score.

On labels that sit mostly near zero, the regression head worked out that predicting the average keeps the loss low, and then stopped there. Its correlation with gold is 0.03. Since filtered-out candidates score 0, the "feature" was really a flag for "survived filtering", which the ranker could have had for free, minus the 500 MB of weights.

I also ran a feature-ablation sweep: drop one signal, re-score the whole dev set, see what breaks. (These runs use the same contaminated set, so read them as comparisons between configurations, not as absolute scores.) Removing the MLM fluency score hurts the most, taking NDCG from 0.531 to 0.469, which is exactly what I expected.

Here is what I did not expect. Removing the cross-encoder reranker entirely beats the full pipeline on every metric: NDCG 0.541 vs 0.531, MAP 0.365 vs 0.356, precision 0.275 vs 0.263, pairwise 0.639 vs 0.624.

The most expensive stage in the pipeline, an extra transformer forward pass per candidate in the rerank pool, is net negative.

Bar chart of NDCG, MAP, and precision at k=10 for the full pipeline versus six single-feature ablations, showing "No rerank" scoring highest on all three metrics
"No rerank" is the tallest bar on all three metrics. The pipeline ranks better with the cross-encoder stage removed entirely than with it included.

My best guess at why: the cross-encoder is a general sentence-pair similarity model (cross-encoder/stsb-roberta-base, trained for semantic textual similarity), not anything tuned for whether one specific word swap is correct. It plausibly rewards paraphrase-level closeness, two sentences that "mean about the same thing", over the surgical judgment this task needs, and at a blend weight of 0.75 that reward dominates the final score for anything in the rerank pool. It's easy to miss if you only eyeball a handful of examples where the top result looks fine. The pipeline shipped with the reranker on by default before I ran this ablation.

What I would do differently

  • Split my data the way I'd tell anyone else to: train, dev, and test, each doing exactly one job. Every problem above traces back to the dev split doing three jobs at once (training the ranker, tuning the weights, and grading the result).
  • Turn the reranker off by default, or replace the general STS cross-encoder with one fine-tuned on SWORDS pairs, instead of trusting a similarity model to do a job it was never trained for.
  • Catch a collapsed model at training time. A one-line check on the variance of its predictions would have flagged the RoBERTa model before it shipped.
  • Add real unit tests. The SWORDS evaluation is the only thing exercising this code, and it never touches the heuristic corners like the double-suffix guard or the verb-frame matching. That's exactly the kind of code that breaks quietly and only shows up as a slightly worse aggregate score.
  • Measure what single-word-only scoping actually costs, instead of assuming the ~18% of multi-word candidates weren't worth the complexity.