Paper 5 of Five
Abstract
This is the guardrail paper and the one that protects the series. Every claim in Papers 1 through 4 is an absence claim, and absence claims are the easiest kind to make and the easiest to make badly, because the evidence is by construction not there. The load-bearing move in all four is that some text, question, obligation, or standard failed to appear where it bore — and that move is worthless without a base rate. Homiletic genre produces innocent silences. So do occasional preaching, lectionary constraint, and the plain economics of printing. This paper supplies the procedure for building the control corpus against which the other four are measured, the thresholds below which a silence claim must be abandoned, and a set of innocent-silence generators that any absence claim must clear before it counts as a finding. It is written explicitly as the paper an adversary would need in order to attack Papers 1 through 4, and it is published alongside them rather than after them for that reason. It also carries three results from the series’ own diagnostic runs in which the guardrail caught something: a register test that could not be run, an amendment that excluded the contested, and a test that passed while blind. Falsification constraint: if the control corpus, built to this specification, shows the same asymmetries as the demonstration corpus, Papers 1 through 3 have no referent and should be withdrawn.
1. Why This Paper Exists
Consider the shape of the argument the series makes. A verse is missing where it bore (Paper 1). A question was not disputed where it should have been (Paper 2). One half of a reciprocal command was cited and the other was not (Paper 3). A standard was out of scope where the decision was made (Paper 4).
Each of these is a claim that something did not happen. Each has the same vulnerability, and it is not a subtle one: things fail to happen constantly, for reasons having nothing to do with anyone’s interests. A preacher does not cite every apt verse. A controversialist does not answer every argument. A register does not consider every standard. Most of what did not happen did not happen for no reason at all.
The question every absence claim must answer is therefore: compared to what? An observed non-citation rate of 90% means nothing until we know what rate obtains for comparable texts on comparable questions in comparable documents. If the comparable rate is 88%, there is no finding. If it is 15%, there may be one.
This is the comparison class problem, and it is the whole methodological difficulty of the series compressed into a sentence. Papers 1 through 4 supply the categories. This paper supplies the denominator without which the categories cannot be applied.
A note on where this paper sits. It is not a limitations section. A limitations section concedes weakness after the argument is made and asks the reader to discount accordingly. This paper is the argument’s condition of possibility, and the other four are not established until it has been executed. Every one of them says so in its own notes.
2. The Innocent-Silence Generators
Before a control corpus can be built, one must know what it is controlling for. Eight generators of innocent silence, each of which can produce the observed pattern with no directional mechanism whatever.
G1 — Genre convention. A homily is not a brief. Its conventions govern what may be raised, in what order, at what length, and with what degree of contention. A form that opens on a lectionary text and moves to application will not canvass the canon, and its failure to do so is a fact about the form.
G2 — Occasion. Occasional preaching is preaching to a moment: a fast day, an installation, a funeral, a statute just enacted. The occasion determines the question, and questions the occasion did not raise go unaddressed for that reason.
G3 — Lectionary constraint. Where a preaching tradition assigns texts by calendar, the distribution of cited verses is partly a function of the calendar. Verses outside the cycle appear less often, and this has nothing to do with what they say.
G4 — Printing economics. Not everything preached was printed, not everything printed survives, and length cost money. Publication is a filter, and the filter’s criteria — salability, sponsorship, controversy value — are not the historian’s criteria.
G5 — Memory and frequency effects. Citation reflects what is in the citer’s working memory, and what is in working memory reflects prior frequency of exposure, which reflects prior citation. This is self-reinforcing and indifferent to content. The bibliometric literature has documented the accumulation dynamic at length, and there is no reason the pulpit is exempt.
G6 — Position within a passage. Opening verses are more citable than closing ones for reasons of structure alone. Paper 3 §5.4 is built around this and treats it as a falsification condition rather than a nuisance.
G7 — Polemical need. Writers cite what is contested. An uncontested proposition generates no citations because there is nothing to establish. Paper 3 §6.4 identifies this as the objection it cannot fully answer, and this paper’s §5 supplies the test.
G8 — Survival and digitization. What is countable is what survives and has been scanned. Both filters are non-random and both correlate with institutional prominence, which correlates with the variables of interest.
Any absence claim must clear all eight, and the floor principle applies: the claim is as strong as its weakest cleared generator, not as strong as the number cleared.
3. Building the Control Corpus
3.1 The matching specification
The control corpus consists of documents matched to the demonstration corpus on the variables that drive G1 through G8, addressing questions where no absence is alleged.
Match on: denomination; region; decade; publication venue and imprint; document form (published sermon, pamphlet, address, quarterly article, denominational proceeding); occasion type; and, where determinable, author prominence.
Do not match on: the direction of the argument, the author’s position on the demonstration question, or anything correlated with the hypothesis. Matching on these would build the finding into the control.
Control questions are power-relation questions handled by the same men in the same venues where no quarantine or subtraction is alleged: employer and hireling, creditor and debtor, magistrate and subject, rich and poor within the congregation, parent and child, and — for Paper 2’s title work — questions of church property, denominational schism, and competing land claims where origin was in fact litigated.
3.2 Sampling
Frame: the full set of imprints meeting the match criteria in the accessible bibliographies, enumerated before selection.
Selection: random within stratum, strata defined by decade × denomination × form. Convenience sampling is barred; where a stratum cannot be filled, the shortfall is reported and the stratum is dropped rather than backfilled with whatever is at hand.
Size: determined by the precision needed on the smallest contrast in Papers 1 through 3, computed and fixed in advance. Not by what proves convenient to collect.
3.3 The within-document and within-passage controls
The two strongest designs in the series do not need a separate corpus at all, and this is worth stating plainly because it is the series’ best evidentiary asset.
Within-passage (Paper 3 §2.3): comparing citation of Ephesians 6:5 against 6:9 holds constant G1 through G5 and G8 by construction, since both halves share book, author, context, familiarity, printing cost, lectionary position, and survival. Only G6 and G7 remain live, and both have dedicated tests (§5 below and Paper 3 §5.4).
Within-author, between-register (Paper 4 §4.2): comparing one man’s treatment of a question across two venues holds constant belief, competence, era, and position.
Where these designs are available they are preferred, and the external control corpus serves to establish the general citation environment rather than to carry the primary contrast.
3.4 The register corpus
Paper 3’s A5 condition and Paper 4’s S5 signature both require a corpus of non-polemical registers: family instruction manuals, devotional commentaries, catechetical material for households, pastoral works addressed to householders.
One exclusion, and it is not optional. Sequential commentaries on Ephesians and Colossians are excluded. A commentary covers every verse by genre obligation, so its treatment of a verse carries no information about selection. Including them would manufacture the appearance of register presence.
4. Thresholds and Abandonment
Stated in advance and binding.
T-A — Differential threshold. A silence claim requires that the demonstration rate exceed the control rate by a margin exceeding the control’s own between-stratum variance. A difference smaller than the variation among control strata is not a difference.
T-B — Reliability floor. Krippendorff’s alpha ≥ 0.67 on the primary codes, computed separately per code. Below the floor, the coding is reported as failed. No recoding to reach threshold; a second attempt uses fresh material and is reported as a second attempt.
T-C — Indeterminate ceiling. 35%. Above it, the coding is reported as failed.
T-D — Generator clearance. All eight generators cleared, floor principle applied.
T-E — Population threshold. For Paper 4’s region-based work, a region must contain at least 5% of the identifiable stream, with sensitivity at 1% and 10%.
T-F — Preregistration integrity. Region definitions, position-reversed passage sets, register classifications, and control-question lists are fixed in writing before any counting. A set assembled after the primary result is known is not a control.
5. The Test That Decides Paper 3
G7 — polemical need — is the generator that survives the within-passage design, and it is therefore the one that can defeat the series’ strongest measure. It deserves its own procedure.
The argument to be defeated: the servant’s half was cited because it was contested; the master’s half was not cited because nobody denied it. Uncontroversial propositions generate no citations. On this account the asymmetry is a fact about the shape of the dispute and implies nothing about the direction of obligation.
The test: if the master’s obligation went uncited because it was uncontested common ground, it should appear frequently in non-polemical registers, where uncontested truths are the ordinary furniture. Devotional writing, family instruction, and pastoral works addressed to householders are full of things nobody disputes; that is what they are for.
- If the self-binding half is thin there too, G7 fails and the asymmetry is not explained by contestation.
- If it is well represented there and thin only in controversy, G7 largely succeeds, and the finding converts from a subtraction (Paper 1) into a register effect (Paper 4).
This test runs before the primary analysis, per Paper 3’s fixed analysis order. A design that tests its own thesis last has arranged to know the answer before checking whether the answer means anything.
Status: attempted, not completed. The attempt is reported at §6.1, and its single scoping observation was mixed in a way that forced a refinement of the hypothesis itself.
6. Three Cases Where the Guardrail Caught Something
A guardrail that has never stopped anything is decoration. The series’ diagnostic runs produced three instances where this paper’s discipline changed a result, and they are reported here because they are the evidence that the discipline is operative rather than professed.
6.1 The register test that could not be run
Paper 3’s A5 test was attempted. The sources proved fully public, out of copyright, and machine-readable; the obstacle was retrieval infrastructure rather than evidence, and the distinction matters for whether the constraint is real or decorative. A single work was partially read and produced one observation with no inferential weight: inward-binding pastoral material present in the register, and half-verse warrant also present in the same work.
The mixed result forced a split of the hypothesis into H4a (self-binding halves are not well represented in non-polemical registers) and H4b (where they appear, they are not brought to bear as obligations on the holder). The refinement — present versus brought to bear — proved more useful than the test would have been, and became the operational core of Paper 4.
What the guardrail did: it prevented a mixed one-book observation from being reported as a result, and it converted an incomplete test into a sharper hypothesis rather than a hedge.
6.2 The amendment that excluded the contested
Paper 4’s region-based work acquired an amendment (A4) requiring that the governing standard be acknowledged by the parties to bind conduct in a candidate region, on the reasoning that a standard’s own scope exclusions are not gaps.
Run on two domains it was not derived from, A4 produced determinate rulings — and produced them by excluding every region whose boundary was disputed. Contestation itself became the ground of exclusion. Since a disputed boundary is precisely the condition under which a gap is most likely to exist, an amendment introduced to correct an over-finding bias had introduced an under-finding one. Four instances were logged across two domains.
What the guardrail did: it caught a defect that was invisible on the case the amendment was derived from, and it did so because the amendment was tested off-case. This is the strongest single vindication of the series’ preregistration commitments and it should be read as such.
6.3 The test that passed while blind
A control domain was run to check whether the region method over-finds. It returned the predicted result — no populated non-closing region — and the prediction was confirmed.
The domain nonetheless contained a well-documented failure of the professional standard to reach conduct, produced not by any boundary but by a register holding jurisdiction and declining to exercise it on discretionary grounds. The method reads published boundaries and was structurally incapable of seeing it.
What the guardrail did: it required asking what a confirmed prediction was evidence for, and the answer was: less than it appeared. “Returns no region” is satisfied both when the method works and when it is blind, and only the second was true. The lesson recorded — that predictions on apparatus must name a direction of error and a condition under which confirmation would be uninformative — is a general one and belongs here rather than in the run that produced it.
7. Steelmanning
7.1 The control corpus cannot be built
The objection. The specification at §3 is a fantasy of resources. Stratified random sampling from an enumerated frame of nineteenth-century imprints, matched on seven variables, sized by advance power calculation, coded by three independent raters — this describes a funded multi-year project, not work an independent scholar can execute. A methodology whose guardrail is unbuildable is a methodology with no guardrail, and the honest description of Papers 1 through 4 is that they are conjectures presented in the vocabulary of measurement.
Response. The objection is largely correct about the full specification and largely wrong about the consequence.
Correct: the external control corpus at §3.1–3.2 is expensive. I have not built it and may not be able to.
Wrong about the consequence, for two reasons. First, the strongest designs do not need it. §3.3 shows that the within-passage and within-author controls neutralize six of the eight generators by construction, and those designs are cheap — they require the demonstration corpus and nothing more. Second, an unbuilt control is a stated debt rather than a hidden one. Every paper in the series says in its own notes that nothing is established until this corpus exists. A reader who takes them as conjecture has read them as written.
What I cannot claim is that the debt will be paid. That is a real limit and the series should be judged with it in view.
7.2 The generator list is arbitrary and incomplete
The objection. Eight generators, chosen by the investigator, with no argument that they exhaust the space. Any absence claim survives by clearing the generators the investigator thought of. The ninth generator — the one not on the list — defeats the finding, and there is no procedure for finding it.
Response. Accepted without qualification; the list is not closed and cannot be. Two mitigations, neither adequate.
The generators are stated in advance and in public, which makes additions to the list a legitimate move for an adversary and makes the list’s contents auditable. And the within-passage design defends against unknown generators in a way that generator-by-generator clearance cannot, because it holds constant everything shared by two halves of one passage, including generators nobody has named. That is the argument for preferring it, and it is the strongest structural reply the series has.
7.3 The three cases at §6 are self-reported
The objection. A researcher reporting that his own guardrail caught his own errors is offering the least verifiable form of evidence. Each of the three cases was identified by me, characterized by me, and its significance assessed by me. §6.2 in particular reports an amendment failing — but I wrote the amendment, I chose the test domains, and I graded the result.
The strongest form. Worse: reporting caught errors is a well-known credibility strategy. A paper that displays its own self-correction purchases trust for the claims it did not catch, and the reader has no way to distinguish thorough self-scrutiny from selective display of the failures that were safe to admit.
Response. This is the objection I have no good answer to, and I want to state that rather than manage it.
The partial answer is that §6.2’s failure is not a safe one to admit. It invalidates an amendment on which Paper 4’s remaining structure depends, and it leaves a named defect unrepaired under a self-imposed freeze rather than resolved. Displaying that costs more than it buys. But an adversary can reply that a costly-looking admission is exactly what a credibility strategy would select, and I cannot refute this from inside.
The only real remedy is external replication, and the specifications in this paper exist so that someone who does not accept the conclusions can run them. Whether anyone will is not in my control.
7.4 The thresholds are set to be clearable
The objection. An alpha floor of 0.67 sits at the boundary conventionally used for tentative conclusions, well below the 0.80 usually required for firm ones. A 35% indeterminate ceiling permits a third of the data to be uncodable. A 5% population threshold is low. Each is defensible individually; together they describe a bar chosen so the series can clear it.
Response. Partly right and the right part should change the reporting rather than the thresholds.
The 0.67 floor is adopted from the broader methodology work for consistency and deliberately not tuned to this series, which is the strongest defense available for any threshold. But the objection identifies a real asymmetry: a result clearing 0.67 and not 0.80 is tentative, and the series must say so in the text of the finding rather than in a note. That is a commitment and it is recorded at note 4.
The 5% population threshold I cannot defend as principled. It was set before the funeral-director run but after the research-integrity case, which is weaker preregistration than the series requires elsewhere, and it has never been evaluated against data because no counting has been done.
7.5 The paper protects the series it belongs to
The objection. A guardrail written by the same author, in the same program, published in the same series, is not an independent check. Its function is to make the other four papers look disciplined. A genuine adversary would not have written §3 or §4; he would have written the corpus that refutes them.
Response. Correct as a description of the structural position and the reason the paper is published alongside rather than after. Publishing it simultaneously means an adversary has the tools at the moment of first presentation rather than after the conclusions have circulated. That is the most a same-author guardrail can do.
It does not make the check independent. Nothing written by me can.
8. Falsification Constraint
If the control corpus, built to the specification at §3, exhibits the same asymmetries as the demonstration corpus after matching, then the asymmetries are a property of the genre rather than of the question, and Papers 1 through 3 have no referent and should be withdrawn.
Subsidiary constraints, each independently sufficient:
F1 — Control parity. Adverse/favorable citation asymmetry in the control matches the demonstration → the effect is generic.
F2 — Position dominance. Within-passage asymmetry tracks verse position rather than direction of obligation on position-reversed material → the mechanism is bibliometric.
F3 — Register representation. Self-binding halves well represented in non-polemical registers (H4a fails) → G7 succeeds and the finding migrates to Paper 4.
F4 — Coding failure. Alpha below 0.67 or indeterminate above 35% → published as failed, not recoded.
F5 — Frame instability. The sampling frame at §3.2, enumerated twice at separated dates, does not reproduce → the control is not replicable and the specification fails.
All reported whichever way they fall. F2 and F3 run before the primary analyses, per Paper 3’s fixed order.
9. What Is Not Claimed
This paper does not claim the control corpus exists. It does not. §7.1 concedes the cost and the possibility that it will not be built.
It does not claim the generator list is complete. §7.2 concedes it cannot be.
It does not claim that the self-reported catches at §6 constitute independent verification. §7.3 concedes they do not and that no answer is available from inside.
It does not claim that any result in Papers 1 through 4 has been established. None has. No counting has been performed anywhere in the series, and every number that has appeared in any of the five papers was labeled as invented for the purpose of demonstrating arithmetic.
What it claims is narrower: that the conditions under which the other four papers would be right are specifiable in advance, that the conditions under which they would be wrong are specifiable in advance, and that both have been specified before any evidence was collected.
Notes
- Order of publication. This paper is published with Papers 1 through 4, not after them. An adversary receives the attack tools at the moment of first presentation. This is a design decision and the series’ central procedural commitment.
- The floor principle. Applied throughout: a cumulative case is as strong as its weakest satisfied condition, not as strong as the sum. This governs generator clearance at §2, signature counting in Papers 2 and 4, and the joint case register below.
- Joint case register. Papers 1 through 4 overlap in the demonstration corpus. A case satisfying two categories is entered once, with two descriptions, and the coincidence is never treated as independent confirmation. The register is maintained across the series and its maintenance is checkable.
- Tentative-result reporting. Per §7.4: any result clearing alpha 0.67 but not 0.80 is described as tentative in the text of the finding, not in a footnote.
- The 5% threshold. Conceded at §7.4 as insufficiently preregistered. It should be re-derived from a principle — a candidate is the share below which the region would not change any conclusion the domain’s own rulemaking has thought worth reaching — before Paper 4’s region work is presented as established.
- The uptake class. §6.3 identifies a class of non-closure produced by discretionary declination within a register that formally holds jurisdiction. It is logged and unrepaired under the amendment freeze. Whether it is the same phenomenon as the conditional-cession form in Paper 4’s boundary section — both being formal closure without functional closure — must not be settled while the freeze is in force.
- Amendment discipline. The program has twice adopted an amendment produced by the case it survived. The standing rule: no amendment is adopted until it has been run on a domain selected after the amendment was fixed and not resembling the one that produced it. One such test is currently pending.
- Digitization bias. G8 is the generator least amenable to control and most likely to correlate with the variables of interest, since scanning priorities track institutional prominence. Any result should report the proportion of the frame that was accessible and the direction in which inaccessibility would bias it.
- Scripture. Quotations follow the Authorized Version throughout the series, both because it is the text the corpus used and because the argument turns on what a nineteenth-century reader had before him.
- On what remains undone. The external control corpus (§3.1–3.2), the register corpus (§3.4), the position-reversed passage set (F2), the sampling frame enumeration (F5), and the pending amendment test (note 7). Until these exist, the series is a specification and not a set of findings, and should be cited as one.
References
Barnes, A. (1846). An inquiry into the Scriptural views of slavery. Parry & McMillan.
Bourne, G. (1845). A condensed anti-slavery Bible argument. S. W. Benedict.
Brown, C. G. (2004). The word in the world: Evangelical writing, publishing, and reading in America, 1789–1880. University of North Carolina Press.
Cochran, W. G. (1977). Sampling techniques (3rd ed.). Wiley.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum.
Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465.
Gutjahr, P. C. (1999). An American Bible: A history of the Good Book in the United States, 1777–1880. Stanford University Press.
Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89.
Holifield, E. B. (2003). Theology in America: Christian thought from the age of the Puritans to the Civil War. Yale University Press.
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124.
Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196–217.
Krippendorff, K. (2004). Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research, 30(3), 411–433.
Krippendorff, K. (2018). Content analysis: An introduction to its methodology (4th ed.). SAGE.
Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the growth of knowledge (pp. 91–196). Cambridge University Press.
Lange, J. (1966). The argument from silence. History and Theory, 5(3), 288–301.
Lombard, M., Snyder-Duch, J., & Bracken, C. C. (2002). Content analysis in mass communication: Assessment and reporting of intercoder reliability. Human Communication Research, 28(4), 587–604.
MacRoberts, M. H., & MacRoberts, B. R. (1989). Problems of citation analysis: A critical review. Journal of the American Society for Information Science, 40(5), 342–349.
Mathews, D. G. (1977). Religion in the Old South. University of Chicago Press.
McGrew, T. (2014). The argument from silence. Acta Analytica, 29(2), 215–228.
Meehl, P. E. (1990). Appraising and amending theories: The strategy of Lakatosian defense and two principles that warrant it. Psychological Inquiry, 1(2), 108–141.
Merton, R. K. (1968). The Matthew effect in science. Science, 159(3810), 56–63.
Noll, M. A. (2006). The Civil War as a theological crisis. University of North Carolina Press.
Nord, D. P. (2004). Faith in reading: Religious publishing and the birth of mass media in America. Oxford University Press.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Popper, K. R. (1959). The logic of scientific discovery. Hutchinson.
Proctor, R. N., & Schiebinger, L. (Eds.). (2008). Agnotology: The making and unmaking of ignorance. Stanford University Press.
Rosenthal, R. (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3), 638–641.
Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Stout, H. S. (1986). The New England soul: Preaching and religious culture in colonial New England. Oxford University Press.
Swartley, W. M. (1983). Slavery, Sabbath, war, and women: Case issues in biblical interpretation. Herald Press.
Tise, L. E. (1987). Proslavery: A history of the defense of slavery in America, 1701–1840. University of Georgia Press.
Trouillot, M.-R. (1995). Silencing the past: Power and the production of history. Beacon Press.
Weld, T. D. (1838). The Bible against slavery. American Anti-Slavery Society.
