← The Threshold Archive

Methodology

How 9,764 accounts became the numbers on the previous page — what was measured, what was validated, and what is still provisional.

01 — Sources and extraction

Three archives, three architectures.

The corpus is the complete public archives of three sister sites run by the Experience Research Foundation. Each was scraped with a purpose-built extractor, because each stores its accounts differently.

ArchiveContentAccountsWordsDiscovery method
nderf.orgNear-death experiences5,7508,192,29952 static archive index pages
adcrf.orgAfter-death communication1,7661,850,928A–Z sitemap ∪ 8 era archives
oberf.orgOut-of-body, prebirth, deathbed and related2,2482,817,541A–Z sitemap ∪ 20 category pages
Discovery

Working from the published indexes

Each archive publishes its own index of accounts, and those indexes are the basis for everything here. Walking them yields 5,811 story pages for the near-death archive from plain HTML, which can be checked against the sites themselves. Where a demographic field — country, age, editorial classification — is not present on the page, it is taken from the archive's own published record for that account and attached afterwards. Nothing is inferred that the archives do not already state.

Coverage

Union, because every index is incomplete

On both sister sites the A–Z sitemap and the category pages each omit accounts the other contains. oberf.org's sitemap has a corrupted stretch through section D — 19 accounts appear only in the category listings, while 56 older ones appear only in the sitemap. Discovery therefore unions every index and de-duplicates. Final coverage: 9,764 of 9,793 discovered URLs. All 29 misses are genuine dead links published by the sites themselves.

Politeness

Rate limits and shared hosting

Collection was deliberately slow and light: two concurrent requests per server at most, a minimum spacing between them, backoff on any error, and a descriptive user agent so the traffic is identifiable. Two of the archives share a server, so they share one budget rather than getting one each. Pages are cached locally, so re-running an analysis costs the archives nothing.

02 — What the data is

A survey, not a pile of prose.

This is the fact that makes quantitative analysis possible. Every contributor writes a free-text narrative and answers a standardised questionnaire — roughly 50 items on nderf.org, answered by 5,300–5,700 people each.

So each account yields two independent records of the same experience: what the person chose to write, and how they responded when asked directly. The site's published research draws on the questionnaire. Comparing the two halves is where most of the findings on the previous page come from.

Enrichment

Demographics

Gender, religion before and after, and the date of the experience are all answers the experiencer gave in the questionnaire. Country, age bracket and editorial classification come from the archive's own record rather than from the page text, and are matched back to each account. Final coverage: 93.2% for country and 88.5% for age. The remaining gap skews toward recent submissions, so any trend across time slightly under-weights the newest accounts — stated wherever those figures appear.

03 — Response coding

The questionnaire has no single answer format.

This is the least glamorous part of the work and the one most capable of destroying every number downstream. Responses arrive in at least three shapes, and no generic rule handles all of them.

ShapeExample itemWhat a naive parser does
Yes / No / UncertainDid you pass through a tunnel?Works
Coded value plus free text“Yes” followed by a paragraph of elaborationDiscards the answer entirely
Descriptive alternativesDid time seem to speed up or slow down?Reports 0% — “yes” never appears
Non-response“No response”, “No comment”Counted as a negative

Every item's option vocabulary was therefore enumerated from the data and coded by hand as affirmative, partial, negative, uncertain or missing. Denominators count only valid responses; “partial” is reported separately rather than folded into the headline figure. The coding table is published in full.

04 — Element detection

Two rates, never merged.

Prompted rates come from the questionnaire across the full corpus. Spontaneous rates come from reading narratives and recording what the experiencer described unbidden. They answer different questions and are never summed or averaged.

Sampling

1,500 narratives, 500 per archive

Equal allocation rather than proportional, so each archive carries the same precision for cross-corpus comparison; sampling weights are recorded so overall rates can be reweighted to the true corpus mix. Narratives under 60 words were excluded. Batches were sized by word count rather than narrative count, so a single 15,000-word account could not crowd out its batch.

Coding

21 elements, read individually

Each narrative was read and coded for 21 elements — presence only, negations excluded, borderline calls recorded separately as uncertain rather than forced. A further 600 narratives were coded in more depth for event order, beings encountered, settings, corroboration claims and style. Coverage: 1,497 of 1,500 and 598 of 600.

05 — Validation

The first method failed its own test.

Element detection was originally done with hand-written text patterns. Before trusting it, 80 narratives were coded blind by two independent readers and the patterns scored against them.

MeasureResultReading
Inter-reader agreement (κ)0.96The task itself is reliable — readers agree on 98.4% of judgements
Pattern precision0.84What the patterns found was usually right
Pattern recall0.40They missed three fifths of everything
Recall range across elements0.00 – 0.93Fatal: relative frequencies became artefacts of pattern quality

The κ of 0.96 is what makes the verdict unambiguous. Because two independent readers agreed almost perfectly, the disagreement with the patterns was pattern failure, not ambiguity in the task. There was no defence available.

The uneven recall mattered more than the low average. “Peace” scored 0.04 and “religious figure” 0.93, so any ranking between elements would have measured how well each pattern happened to be written. The pattern lexicon was retired, flagged invalid in the published codebook, and replaced by read-through coding. Scoring code and the labelled sample are in the repository.

06 — Five failures

Every one produced a confident wrong number.

None of these threw an error. Each was caught because a figure was implausible, not because anything broke.

Elaborated answers were discarded

Many people answer “Yes” and then keep writing. Matching only the bare token dropped every one of them from both numerator and denominator.

Saw a light: 9.7% → 62.6%

Six items were never yes/no

They offer descriptive alternatives, so no valid answer contains the word “yes”.

Time distortion: 0.0% → 76.6% · Border: 0.1% → 65.4%

Pattern matching missed most of its targets

Recall of 0.40, varying from 0.00 to 0.93 by element, making cross-element comparison meaningless.

Retired; replaced by 1,500 read-through judgements

An age effect was a recall effect

Reports of psychic ability appeared to vary sharply with age at the time of the experience. Cross-tabulating against years elapsed showed most of the gradient was simply time to notice.

Apparent 34-point age gradient → largely recall interval

Pooling three archives invented a correlation

Combining collections with very different base rates made encounters with the dead appear to repel every other element. Within a single archive they are among the strongest pairings — Simpson's paradox.

Caught by re-running every correlation within corpus

07 — Limits and confounds

What these numbers cannot tell you.

This is not a population sample

Contributors chose to submit to a research site about near-death experience. Nothing here estimates how common any of this is in the general population — only what this corpus contains.

Prompting is not neutral

Asking “did you pass through a tunnel?” invites a judgement about a specific element. That is precisely why prompted and spontaneous rates are reported separately and never merged: the gap between them is a measurement, not a nuisance.

The corpus is 61% United States

Of accounts with a recorded country, 3,284 of 5,358 are American. Canada, France, Australia and the United Kingdom follow far behind. Any “cross-cultural” comparison here is really the United States against a thin and uneven tail.

Recall intervals are long and uneven

Many accounts were written decades after the event. Every content finding was tested against years-elapsed; one apparent age effect dissolved entirely under that test, and it is the first check any new finding gets.

The questionnaire changed

Earlier submissions answered roughly 23 questions; later ones answer 50–57. Any trend across time is confounded with instrument version unless segmented, so decade comparisons are reported within recall window.

Coverage is not quite complete

Country and age reach 93.2% and 88.5% of accounts. The unmatched remainder skews toward newer submissions, so it is not a random slice — trends across time will slightly under-weight recent accounts.

08 — Still provisional

Flagged rather than fixed.

Three things on the previous page are weaker than the rest, and are marked here so nobody cites them as settled.

Claim taxonomy

Categories for corroboration claims were defined in advance rather than derived from the data, and 314 of 880 claims fell into “other”. The specificity breakdown is unaffected — it is coded independently — but the kinds of claim should be rebuilt from the corpus before anyone relies on them.

Message themes

Recurring themes in what people report being told are keyword matches over paraphrases, never checked against blind reading. Unlike the element labels, they carry no validation. They are the weakest numbers produced and are not shown on the main page.

The life review measure

Its two affirmative options behave very differently from one another, which suggests merging them into a single rate was probably a mistake. The decline across decades holds in both recall windows, but the absolute level should be treated as soft.

Everything here is descriptive. It reports what tens of thousands of pages say and how often they say it. It makes no claim about what caused the experiences, and none of the analysis is capable of addressing that question.