The short version
- Amazon indexes description text. For distinctive phrases whose words appear nowhere else on the listing, the book surfaced in organic results 1.6% of the time, against 0.0064% for matched control books. Roughly 260 times the odds.
- Distinctive phrasing carries the effect. Phrases rare across the corpus ranked at 1.6%, moderately common ones at 0.9%, and boilerplate at 0.24%. Stock phrases like "New York Times bestseller" contribute close to nothing.
- The description is the weakest of the three search surfaces. Phrases in the title and other visible metadata surfaced the book about 35% of the time. Publisher-submitted backend keywords in our earlier baseline, about 8%. Description-only phrases, 1.6%. The keyword baseline reflects how publishers commonly choose keywords, not the ceiling of the field.
- Practical read: description copy is doing double duty, and specific language gets picked up where generic language does not. It adds incremental visibility. It does not do the keyword field's job.
Publishers treat the description as sales copy, written for a reader who has already landed on the page. That is its main job. The open question has been whether Amazon also reads it as retrieval text, and if so, whether the effect is large enough to influence how the copy gets written.
Kadaxis ran the study across 25,199 books with usable descriptions. Amazon does index description text. The rate is low overall, and it climbs several times over for distinctive language.
The hierarchy of surfaces
The most useful result is the hierarchy. In this study, phrases in visible metadata and distinctive description-only phrases surfaced books at 34.96% and 1.60%. Earlier Kadaxis research at shallower search depth placed publisher-submitted backend keywords between them, at about 8%.
| Where the phrase appears | Book surfaced in organic results |
|---|---|
| Title, subtitle, or other visible metadata | 34.96% |
| Publisher-submitted backend keywords (earlier Kadaxis baseline)* | about 8% |
| Description only, distinctive phrase | 1.60% |
| Description only, all phrases | 0.49% |
*The 8% figure measures publisher-submitted keyword quality, not the ceiling of the keyword field. Kadaxis keywords, measured the same way, rank about five times more often.
A phrase in the title is roughly twenty times more likely to surface the book than the same class of phrase sitting only in the description. Backend keywords land between the two.
That ordering gives publishers a clean priority. Title and subtitle carry the most retrieval weight across the fewest slots. The keyword field sits second, and the eight percent figure above describes the publisher-submitted baseline, not what the field can do. The linked research traces the low rate to how those keywords are chosen. For most publishers the keyword field is the most underused surface on the listing. Descriptions carry a small amount of weight across a much larger body of text, at no extra cost, because the copy is being written regardless.
Distinctive phrases carry the effect
Within description text, the effect concentrates in phrases that are rare across the corpus.
| Phrase frequency across 25,199 descriptions | Surfaced in organic results |
|---|---|
| Distinctive (appears in 3 or fewer descriptions) | 1.60% |
| Moderate (4 to 20 descriptions) | 0.90% |
| Boilerplate (more than 20 descriptions) | 0.24% |
The gradient runs in the direction lexical indexing predicts. A distinctive phrase has few competitors, so a book carrying it has a realistic path to the top of that result set. A boilerplate phrase like "New York Times bestseller" or "a gripping page-turner" is carried by thousands of listings, and no single book surfaces for it.
This gradient also settles an obvious objection. If the ranking came from books being popular rather than from their description text, popular books would rank across all three strata equally. The strongest results would sit with common phrases, where the most competitive titles cluster. The opposite pattern showed up.
The practical version: description copy written in specific, concrete language is retrievable in a way that stock jacket-copy language is not. The phrases that worked in this study name particular characters, particular techniques, particular concepts. The phrases that did nothing are the ones any book in the category could carry.
What the study did
A book ranking for a phrase in its description proves nothing on its own. The phrase may also sit in its title, its categories, or its Amazon-displayed bullets, or Amazon may understand the book semantically with no lexical match anywhere. Isolating the description means stripping out every other explanation first.
Start from real searches. Take every 2-to-6-word phrase in a book's ONIX description, and keep only phrases that exactly match a query someone had already searched and that had a successful result-page fetch on record. This restricts the analysis to phrases that function as plausible reader queries.
Remove the book's own keywords. Exclude any phrase that is an assigned or delivered keyword for that book, from either the ONIX keyword field or the monitored delivery set. A hit from a book's own keyword field proves nothing about descriptions.
Remove phrases with any other lexical home. Two levels here. The first removes phrases whose words appear anywhere in the book's other ONIX fields. The second, which produces the headline number, goes further: it removes any phrase whose words appear anywhere on the Amazon-visible surface, including the page title, subtitle, authors, bullets, browse categories, and format labels, across every edition in the book's family. What survives is a phrase whose words exist on the listing only inside the description.
Confirm the phrase is on the page Amazon sees. ONIX descriptions and Amazon-displayed descriptions can differ. The headline cell requires the phrase to appear in the Amazon-displayed description of at least one edition, proving the text was on the page Amazon indexes. 220 of the 222 headline hits cleared this check.
Verify identity. A raw match on an edition family can produce false positives when a shared ASIN or a classics reprint sits in the family. A hit counts only when the matched edition has a high-confidence relation to the book or when the result title matches the book title. Raw and verified rates were both computed; every number here is verified.
That produced 276,748 candidate pairs across the corpus, with 158,803 in the strictest exclusivity tier and 13,863 in the headline cell of distinctive, description-only phrases.
Controls and effect size
A 1.6% rate means nothing without knowing what the same query does for a comparable book that lacks the phrase.
For every treated pair, the study built a control pool: books in the same top-level browse category, in the same ratings-count quintile, whose descriptions do not contain the phrase, to whom the query was never assigned or delivered, evaluated against the same result-page fetch. Matching on category and ratings volume removes the two most obvious confounders, since a book that ranks for anything tends to be more popular and more topical than the corpus average.
Control books in the headline cell surfaced at 0.0064%. Treated books surfaced at 1.60%.
The stratified effect size, computed as a Mantel-Haenszel odds ratio across 12,782 query-by-category-by-ratings strata, is 259.7, with a 95% confidence interval of 224.1 to 300.9. Confidence intervals on the rates come from a 1,000-iteration bootstrap clustered on books, so a handful of unusual titles cannot drive the result.
Two further checks. The rate holds steady across fetch months, at 1.56% in May and 1.88% in July, so this is not an artifact of one crawl window. And no headline hit came from a sponsored placement; every one is organic.
The clearest cases
The aggregate statistics bound the alternative explanation without eliminating it. For a topical phrase like "negotiation skills," Amazon may be retrieving the book semantically, from its category and subject signals, rather than from a literal string match in the description.
Some hits do not have that ambiguity. The study surfaced phrases at character-name level, strings that exist nowhere on the listing except inside the description text:
- "bear demon mikki" and "grasslander wizard ivah" surfaced their novel at position 1
- "superstitious schoolmaster" surfaced The Legend of Sleepy Hollow at position 2
- "resourceful spaceman" surfaced The Green Odyssey at position 1
- "fermi are awake" surfaced Beyond the Reach of Earth at position 1
No category signal, subject code, or behavioral pattern explains a book appearing first for the name of a minor character. The description text is the only place that string lives, which makes lexical indexing of description text the parsimonious explanation for this class of hit.
Commercial examples run the same way. "Tactical empathy" and "fbi hostage negotiator" surfaced Never Split the Difference at positions 1 and 2, on phrases that describe the book's actual content rather than its shelf.
What the result does and does not establish
Description phrases are retrievable in organic Amazon Books search, at rates far above matched-control background, with the effect confirmed against the Amazon-displayed page and stable over time.
For topical phrases, semantic retrieval remains a live alternative at the level of any individual pair. The matched-control design bounds how much of the effect that could account for; it does not remove it. Character-name cases make lexical indexing the only reasonable account for part of the effect, and the honest position is that both mechanisms are likely contributing.
Review text is a second alternative worth naming. Amazon indexes customer reviews, and reviewers often echo the language of the description, including character names. This study excluded the title, subtitle, authors, bullets, categories, and format from the phrase's possible lexical homes, and did not exclude review text. The matched-control comparison is unaffected, since control books carry reviews too, and the character-name cases carry a little less weight than they would if reviews were ruled out. Separating the two surfaces is its own study.
Absence is also not evidence. Result pages were checked to a depth of roughly 32 to 112 positions. A book indexed for a description phrase but sitting at position 200 counts as a miss in this data while still being indexed.
What publishers should do with this
Three things follow.
Write descriptions with distinctive, concrete language. Name the specific technique, the specific place, the specific premise, the character. Those phrases are retrievable. Category-generic jacket copy is carried by thousands of other listings and returns nothing.
Keep the strongest signals in the strongest surfaces. A phrase worth ranking for belongs in the title, subtitle, or keyword field before it belongs in the description. The description picks up what the higher-weight surfaces have no room for.
Treat the description as a second pass at coverage. Most books have far more true, specific, searchable attributes than fit in a title and twenty keyword slots. The description is long, it is being written anyway, and the words chosen for it produce a small amount of additional organic visibility. That visibility is available for the price of choosing better words.
A 1.6% rate is modest against a 35% rate for title placement, and it is roughly 260 times what a comparable book gets without the phrase. Across a backlist of thousands of titles and descriptions running several hundred words each, a weak surface applied at that scale still produces real discovery.