Abstract
Under current Confrontation Clause doctrine, whether a forensic analyst must testify depends on whether the analyst’s report is testimonial, and whether a report is testimonial depends chiefly on the purpose for which it was made. A laboratory report prepared for a prosecution requires its author; a business record kept for commercial purposes does not. The Supreme Court adopted this line after expressly rejecting reliability as the test for confrontation. This paper asks whether the purpose line nonetheless tracks real differences in reliability between the two classes of evidence. It identifies seven kinds of risk that can corrupt a forensic or documentary result and assesses, for each, whether the purpose line directs confrontation toward it. It concludes that the line tracks one family of risks well: those arising from the adversarial orientation of evidence prepared for litigation, including fabrication and contextual bias. It tracks others poorly or not at all, including instrument error, clerical error, systemic institutional failure, and interpretive overreach in records-based expertise. The purpose line is therefore best understood not as a proxy for reliability in general but as a targeted safeguard against a particular danger, and its gaps must be addressed by other means.
1. Introduction
White Paper 1 identified testimonial status as the axis that does the most work in determining foundation load. Physical evidence accompanied by no testimonial report may enter through a single witness who recognizes it; a record prepared for prosecution requires its author. White Paper 2 showed that the witnesses produced by this requirement frequently form a sequence rather than a corroborating group.
This paper examines the testimonial line itself. The question it poses is narrow but consequential. The Supreme Court has said that the Confrontation Clause is a procedural guarantee, not a reliability rule, and that the constitutional question does not turn on whether a statement is trustworthy.[1] Yet the practical effect of the line is to require live testimony for one broad class of evidence and to dispense with it for another. If the two classes differ in reliability in ways the line captures, the doctrine achieves in practice what it disclaims in principle. If they do not, the line imposes costs on one class of evidence while leaving comparable risks in the other unexamined.
The paper does not argue that the Court’s reading of the Sixth Amendment is historically correct or incorrect. That question belongs to constitutional scholarship and is treated in the monograph. The inquiry here is functional: given the line as drawn, what does it protect against, and what does it miss?
2. From Reliability to Purpose
2.1 The Roberts Framework
For roughly a quarter century before 2004, confrontation doctrine was organized around reliability. Under Ohio v. Roberts, hearsay from an unavailable declarant could be admitted against a criminal defendant if it bore adequate indicia of reliability, which could be inferred where the statement fell within a firmly rooted hearsay exception or shown by particularized guarantees of trustworthiness.[2] Business records and official records, as long-established exceptions, generally satisfied the test. So did many forensic reports, which courts admitted as business or public records.
2.2 The Crawford Turn
Crawford v. Washington abandoned this framework. The Court held that the Clause guarantees a particular method of testing evidence, cross-examination, rather than a particular outcome, reliability, and that allowing judges to dispense with confrontation upon finding a statement reliable substitutes judicial assessment for the procedure the Constitution requires.[3] The Court identified testimonial statements as the core concern of the Clause and left the precise boundaries of that category to later cases.
2.3 The Primary Purpose Test
Subsequent decisions made purpose the organizing principle. Davis v. Washington held that statements to police are testimonial when the circumstances objectively indicate that their primary purpose is to establish or prove past events potentially relevant to later criminal prosecution, rather than to meet an ongoing emergency.[4] Michigan v. Bryant elaborated the inquiry, and Ohio v. Clark applied it to statements made to persons other than law enforcement.[5]
In the forensic context, Melendez-Diaz v. Massachusetts held that certificates of analysis prepared for use at trial were testimonial.[6] The majority distinguished business and public records on the ground that they are created for the administration of an entity’s affairs rather than for the purpose of establishing or proving some fact at trial.[7] Bullcoming v. New Mexico extended the holding to an unsworn but formal blood alcohol report and rejected surrogate testimony.[8] Williams v. Illinois fractured over a DNA profile produced by an outside laboratory, the plurality asking whether the report targeted an accused individual and Justice Thomas asking whether it bore sufficient formality.[9] Smith v. Arizona held that an absent analyst’s statements conveyed by a substitute expert as the basis for the expert’s opinion are offered for their truth, leaving their testimonial status to be determined under the purpose inquiry.[10]
The result is a line that turns on why a statement was made. The line does not ask whether the statement is accurate.
3. The Line as It Falls
Applied to the evidence types classified in White Paper 1, the purpose line produces a clear, if imperfect, division.
On the testimonial side fall laboratory reports prepared at the request of police or prosecutors: drug chemistry certificates, toxicology reports, DNA reports, and pattern evidence reports. Their authors must ordinarily testify, subject to waiver, stipulation, and notice-and-demand procedures.
On the nontestimonial side fall business records such as telephone carrier call detail records, bank statements, and commercial transaction logs, along with many official records kept in the ordinary course. These may be admitted on certification without a live custodian.[11]
A third category has been treated by several lower courts as falling outside the Clause altogether: raw data generated by machines. Courts have held that the printed output of an instrument, such as a gas chromatograph, or automatically generated telephone records are not statements of a person and therefore cannot be testimonial.[12] Under this reasoning, the interpretive statement of the analyst who reads the output is testimonial, but the output itself is not.
In the middle lie autopsy reports and sexual assault examination records, whose purpose is mixed and whose status remains unsettled.
4. A Precedent Within Evidence Law
Before asking whether purpose tracks reliability, it is worth observing that the law of hearsay had already drawn a similar line, and had drawn it on reliability grounds.
In Palmer v. Hoffman, decided in 1943, the Supreme Court held that an accident report prepared by a railroad engineer was not admissible as a business record, because it was prepared not for the systematic conduct of the business but with an eye to litigation.[13] The reasoning was explicitly about trustworthiness. The premise of the business records exception is that an enterprise relies on its routine records and therefore has an incentive to keep them accurate. A document prepared in anticipation of litigation lacks that incentive and may reflect the preparer’s interest in the outcome.
The Federal Rules of Evidence carry forward the same concern. Rule 803(6) permits exclusion of a business record where the opponent shows that its source or the circumstances of its preparation indicate a lack of trustworthiness.[14] Rule 803(8) excludes from the public records exception, in criminal cases, matters observed by law enforcement personnel.[15] In United States v. Oates, the Second Circuit read that limitation to bar the use of a government chemist’s report against a criminal defendant, reasoning that the drafters did not intend adversarial law enforcement reports to substitute for live testimony.[16]
Hearsay law, then, had long treated preparation for litigation as a reliability concern. The purpose line in confrontation doctrine, whatever its constitutional grounding, draws on the same intuition: that a statement made to be used against someone is more likely to be shaped by that use than a statement made for an unrelated purpose.
5. Seven Risks
To assess whether the purpose line tracks reliability, it is necessary to specify what reliability failures look like. The following seven risks account for most of the ways a forensic or documentary result can be wrong.
- Fabrication. The result is invented, as when an analyst reports tests never performed.
- Contextual bias. The result is influenced by knowledge of the investigative hypothesis, such as information that police suspect a particular person.
- Incompetence. The result is wrong because the person producing it lacks the skill or training to produce it correctly.
- Instrument error. The result is wrong because the equipment or software malfunctioned or was improperly calibrated.
- Clerical error. The result is wrong because of mislabeling, transcription error, or data entry mistakes.
- Systemic institutional error. The result is wrong because the institution’s methods, software, or procedures are flawed in ways that affect many results.
- Interpretive overreach. The underlying data is accurate, but the conclusion drawn from it claims more than the data support.
6. Where the Line Tracks Reliability
6.1 Fabrication
The purpose line is well aimed at fabrication. An analyst who knows that a report will be used in prosecution, and who works in close relation to the agency that requested it, may face pressures toward particular results that a clerk recording routine commercial transactions does not. The history of forensic laboratories includes cases of fabrication on a large scale. The West Virginia Supreme Court of Appeals, after an investigation of the state police serology division, concluded that the testimony of a former serologist should be regarded as invalid in the cases he had handled, owing to a pattern of misconduct including false reporting.[17] In Massachusetts, a state chemist’s admitted falsification of drug analyses led the Supreme Judicial Court to order a process that resulted in the dismissal of more than twenty thousand drug convictions.[18]
Cross-examination is not a reliable method of detecting a skilled fabricator in a single case. But the requirement that an analyst appear, swear, and answer personally for a result raises the cost of fabrication and creates a record against which later discovery of misconduct can be measured. The Melendez-Diaz majority relied on precisely this reasoning, observing that confrontation is designed to weed out the fraudulent analyst as well as the incompetent one.[19]
Business records are not immune to fabrication, but routine records ordinarily lack a motive tied to a particular prosecution. Where they are fabricated, the motive is usually commercial or personal and unrelated to the defendant.
6.2 Contextual Bias
The purpose line also tracks contextual bias. Research on forensic decision-making has shown that examiners exposed to case information, such as a suspect’s confession or the investigator’s theory, can reach different conclusions from the same evidence than examiners who lack that information.[20] The risk arises from the adversarial orientation of the work. A report prepared for a prosecution is, by definition, prepared in a context in which such information may be present. A bank’s record of a deposit is not.
The National Research Council’s 2009 report addressed this concern structurally, recommending that forensic laboratories be removed from the administrative control of law enforcement agencies.[21] The purpose line addresses it procedurally, by allowing the defense to question the analyst about what the analyst knew when forming the conclusion.
Taken together, fabrication and contextual bias constitute what this paper calls the adversarial-orientation risk: the risk that a statement made for use against someone will be shaped by that use. The purpose line is a close fit for this risk.
7. Where the Line Does Not Track Reliability
7.1 Incompetence
Incompetence can occur on either side of the line. A laboratory analyst may be poorly trained; so may the employee who configured a carrier’s database or entered transactions into a bank’s ledger. The purpose line requires confrontation of the analyst but not of the database administrator. Its fit with this risk is partial: it reaches incompetence in testimonial reports but not in records.
Studies of wrongful convictions have documented the scale of the problem on the testimonial side. A review of trial transcripts in cases of persons later exonerated by DNA evidence found invalid forensic testimony in a majority of the cases examined, much of it involving overstatement of the significance of results rather than outright fabrication.[22] The finding shows both that confrontation has not been sufficient to prevent incompetent or overstated testimony and that the risk is real where confrontation applies.
7.2 Instrument Error
Instrument error is poorly tracked. Under the machine-output cases, the raw data an instrument produces is not a statement and falls outside the Clause.[23] The analyst who interprets it must testify, but the analyst may be no better placed than anyone else to know whether the instrument was functioning correctly, and cross-examination of the analyst may reveal little. Instrument error is addressed, where it is addressed, by calibration records, maintenance logs, and accreditation requirements, none of which the purpose line reaches directly.
Records-based evidence depends on instruments as well. Carrier data is generated by network equipment and software. Financial records are generated by transaction systems. Errors in these systems are not confronted at all, and the records are admitted on certification.
7.3 Clerical Error
Clerical error, such as a mislabeled sample or a transposed digit, is among the most common failures in both classes of evidence. The purpose line reaches it only where the clerical act is part of a testimonial report. A labeling error at a laboratory intake desk may be explored through the chain of custody witnesses discussed in White Paper 2. A labeling error in a carrier’s cell site list, identifying a tower by the wrong location, is not confronted, although it may be decisive in a location case.
7.4 Systemic Institutional Error
Systemic error is the least well tracked. A flaw in a laboratory’s validation of a method, or in the software used to interpret mixtures, affects every result the laboratory produces with it. Confrontation of the analyst in one case may not reveal it, because the analyst may be unaware of it. Systemic errors in institutional records, such as a carrier’s misconfiguration of tower data, affect every case in which the records are used and are not confronted at all.
Because such errors are institutional rather than individual, the individual accountability that confrontation provides is poorly suited to detecting them. They are detected, when they are detected, through audits, proficiency testing, disclosure obligations, and investigations prompted by anomalies across many cases.
7.5 Interpretive Overreach
Interpretive overreach presents a special case. In Class V evidence, as White Paper 1 described, the foundation is light but the analysis is heavy. The records enter on certification, but the expert who interprets them must testify as a live witness, subject to cross-examination and to the reliability requirements of Rule 702.[24] Confrontation therefore reaches the interpretive stage of records-based evidence even though it does not reach the records.
The same holds for forensic evidence: the analyst who draws a conclusion testifies. The risk of overreach is tracked on both sides, but not by the purpose line. It is tracked because interpretive testimony is live expert testimony regardless of the status of the underlying data.
8. Summary of the Fit
The following table summarizes whether the purpose line directs confrontation toward each risk.
| Risk | Testimonial reports | Nontestimonial records | Fit of purpose line |
|---|---|---|---|
| Fabrication | Confronted | Not confronted; motive usually absent | Good |
| Contextual bias | Confronted | Not confronted; context usually absent | Good |
| Incompetence | Confronted | Not confronted | Partial |
| Instrument error | Analyst confronted; instrument not | Not confronted | Poor |
| Clerical error | Confronted where in report | Not confronted | Poor |
| Systemic error | Individual confronted; system not | Not confronted | Poor |
| Interpretive overreach | Confronted through analyst | Confronted through expert | Tracked, but not by the purpose line |
The pattern is consistent. The purpose line fits the risks that arise from preparing a statement for use against an accused. It does not fit the risks that arise from machines, clerical routine, or institutional design, which are distributed across both sides of the line without regard to purpose.
9. Complications
9.1 Records Produced for Prosecution From Data Not Produced for It
A carrier’s records custodian may compile, at the request of law enforcement, a report of a particular subscriber’s calls and the towers they used. The underlying data was generated for business purposes. The compilation was generated for the prosecution. Courts have generally treated such compilations as business records, on the ground that the data is what matters and the compilation merely retrieves it.[25] But the act of retrieval, including the choice of query, the date range, and the matching of tower identifiers to locations, is performed for litigation and is susceptible to the same clerical and interpretive errors as any other act. The purpose line, applied to the underlying data, can conceal a purpose-driven act of selection.
9.2 The Pipeline Problem
Where a testimonial result is produced by a pipeline, confrontation of one analyst leaves the others unconfronted. Under Melendez-Diaz, not every participant need appear, and under Bullcoming, a surrogate who did not perform or observe the test is insufficient.[26] The analyst who does appear may be able to speak only to his own stage. As White Paper 2 showed, a pipeline is a sequence, and confrontation of one link does not confirm the others. The purpose line identifies the report as testimonial but does not determine which of its many contributors the defendant has a right to confront.
9.3 The Mixed-Purpose Document
Autopsy reports and sexual assault examination records serve both medical and forensic purposes. A purpose test must decide which purpose is primary, and courts applying the same test to similar documents have reached different results. The indeterminacy is not a defect peculiar to these documents. It reflects the fact that purpose, unlike physical nature, is a matter of degree and of perspective.
9.4 Reliability’s Return
Michigan v. Bryant acknowledged that the standard rules of hearsay, designed to identify reliable statements, may be relevant to the primary purpose inquiry.[27] The observation suggests that reliability, expelled from confrontation doctrine as a test, has returned as a consideration. To the extent that purpose is assessed partly by asking whether a statement bears the marks of reliability that the hearsay exceptions recognize, the two inquiries are not fully separable.
10. What Confrontation Accomplishes for Forensic Evidence
The value of the purpose line depends in part on what cross-examination of an analyst actually accomplishes. The Melendez-Diaz dissent argued that analysts are not conventional witnesses, that they typically have no memory of a particular test among thousands, and that their appearance adds cost without adding much scrutiny.[28] The majority responded that confrontation deters fraud, exposes incompetence, and permits inquiry into methodology.[29]
Both positions contain truth. An analyst who performs hundreds of routine tests may indeed recall nothing about a given one, and cross-examination may be limited to confirming the laboratory’s general procedures. At the same time, the requirement of appearance makes the analyst personally accountable, permits questioning about the information the analyst received, and exposes the analyst’s qualifications and error history. Scholarship has also observed that defendants frequently decline to demand the analyst’s appearance, whether from strategy, cost, or inattention, so that the right operates more as a background constraint than as a routine practice.[30]
The accountability confrontation provides is personal. It is well suited to risks that arise from persons, such as fabrication, bias, and individual incompetence, and poorly suited to risks that arise from systems. This is consistent with the pattern in Section 8.
11. Implications
11.1 The Purpose Line as a Targeted Safeguard
The analysis suggests that the purpose line is best understood not as a crude proxy for reliability but as a targeted safeguard against adversarial-orientation risk. On that understanding, the line is neither arbitrary nor comprehensive. It addresses one family of dangers well and leaves others to be addressed elsewhere.
11.2 The Remaining Risks Require Other Instruments
Instrument error, clerical error, and systemic institutional error require instruments other than confrontation: laboratory accreditation, proficiency testing, audit trails, validation studies, disclosure of error rates and corrective action reports, and, for records-based evidence, the opportunity to test the underlying data. The trustworthiness clause of Rule 803(6) provides a limited vehicle for challenging business records, but it places the burden on the opponent, who may lack access to the information needed to meet it.[31] White Paper 6 proposes measures directed at these gaps.
11.3 A Return to Roberts Is Not the Answer
It does not follow from the poor fit of the purpose line with some risks that confrontation should return to a reliability test. The Roberts framework was abandoned in part because judicial assessments of reliability proved inconsistent and because they allowed courts to admit statements on the strength of the very features, such as apparent official regularity, that fabrication and bias exploit.[32] A reliability test would likely have admitted the reports of the analysts described in Section 6.1 without scrutiny. The purpose line, whatever its limits, would have required those analysts to answer personally.
11.4 The Asymmetry Should Be Visible
The practical consequence of the line is an asymmetry in scrutiny. Evidence prepared for prosecution is examined through a procession of witnesses. Evidence drawn from institutional records may enter on paper. Jurors, trial observers, and policymakers are likely to infer that the first class of evidence is more thoroughly tested and the second less trustworthy, or else that the first is more reliable because it is more thoroughly tested. Neither inference follows. The procession reflects the purpose for which the evidence was made, not a measure of how likely it is to be wrong.
12. Limitations
This paper assesses the fit of the purpose line with reliability risks qualitatively. Comparative data on error rates in forensic reports and institutional records are scarce, and the paper does not attempt to quantify how often each risk materializes. It relies on federal doctrine and on the Supreme Court’s decisions; state constitutions and evidence codes may draw the line differently. It does not address the historical question of whether the Framers understood the Clause to reach forensic reports. And its seven-part taxonomy of risk, while intended to be comprehensive for present purposes, is not the only possible division.
13. Conclusion
The Supreme Court adopted the purpose line after rejecting reliability as the measure of confrontation. The line nevertheless has a relation to reliability: it tracks closely the risks that arise when a statement is made for use against an accused, principally fabrication and contextual bias, and it draws on a concern that evidence law had already expressed in excluding litigation-oriented documents from the business records exception. It does not track the risks that arise from instruments, clerical routine, or institutional systems, which afflict testimonial reports and nontestimonial records alike. Nor does it determine which contributors to a pipeline result must appear, a gap that leaves the sequence problem described in White Paper 2 largely untouched.
The line is therefore a partial answer to the question of whom the law should trust and on what terms. It answers that persons working for the prosecution must answer personally for what they say. It does not answer how the law should examine the machines and systems on which both prosecution and commerce rely. The following paper considers what the answer the line does give costs the institutions that must live with it.
Notes
[1] Crawford v. Washington (2004), pp. 61–62.
[2] Ohio v. Roberts (1980), p. 66.
[3] Crawford v. Washington (2004), pp. 61–62, 68–69. The Court described the Clause as commanding that reliability be assessed in a particular manner, by testing in the crucible of cross-examination.
[4] Davis v. Washington (2006), p. 822.
[5] Michigan v. Bryant (2011); Ohio v. Clark (2015).
[6] Melendez-Diaz v. Massachusetts (2009), pp. 310–311.
[7] Melendez-Diaz v. Massachusetts (2009), p. 324.
[8] Bullcoming v. New Mexico (2011).
[9] Williams v. Illinois (2012) (plurality opinion); id. (Thomas, J., concurring in the judgment).
[10] Smith v. Arizona (2024).
[11] Fed. R. Evid. 803(6), 902(11).
[12] United States v. Washington (2007) (raw data generated by laboratory instruments not statements of the technicians); United States v. Lamons (2008) (automatically generated telephone billing data not statements of a person).
[13] Palmer v. Hoffman (1943).
[14] Fed. R. Evid. 803(6)(E). The 2014 amendment clarified that the opponent bears the burden of showing untrustworthiness.
[15] Fed. R. Evid. 803(8)(A)(ii).
[16] United States v. Oates (1977).
[17] In re Investigation of the West Virginia State Police Crime Laboratory, Serology Division (1993).
[18] Bridgeman v. District Attorney for the Suffolk District (2017) established the protocol under which the affected cases were identified and dismissed.
[19] Melendez-Diaz v. Massachusetts (2009), pp. 318–319. The majority cited the National Research Council’s report, released that year, on the state of forensic science.
[20] Kassin, Dror, and Kukucka (2013); Dror and Hampikian (2011).
[21] National Research Council (2009), Recommendation 4.
[22] Garrett and Neufeld (2009).
[23] See note 12. Mnookin (2007) discusses the difficulty of applying testimonial doctrine to expert and instrument-based evidence.
[24] Fed. R. Evid. 702.
[25] The general treatment is discussed in Mosteller (2020). Courts have varied in how they analyze compilations prepared in response to subpoenas, and the monograph’s Chapter 8 surveys the division.
[26] Melendez-Diaz v. Massachusetts (2009), p. 311 n.1; Bullcoming v. New Mexico (2011), p. 652.
[27] Michigan v. Bryant (2011), pp. 358–359.
[28] Melendez-Diaz v. Massachusetts (2009) (Kennedy, J., dissenting).
[29] Melendez-Diaz v. Massachusetts (2009), pp. 318–321.
[30] Metzger (2006) examines notice-and-demand statutes and the conditions under which defendants waive confrontation of forensic analysts.
[31] Fed. R. Evid. 803(6)(E).
[32] Crawford v. Washington (2004), pp. 62–65, collected inconsistent lower-court applications of the Roberts test. Friedman (1998) argued, before Crawford, for a testimonial approach grounded in the procedural character of the right rather than in reliability.
References
Bridgeman v. District Attorney for the Suffolk District, 476 Mass. 298 (2017).
Bullcoming v. New Mexico, 564 U.S. 647 (2011).
Crawford v. Washington, 541 U.S. 36 (2004).
Davis v. Washington, 547 U.S. 813 (2006).
Dror, I. E., & Hampikian, G. (2011). Subjectivity and bias in forensic DNA mixture interpretation. Science & Justice, 51(4), 204–208.
Fed. R. Evid. 702.
Fed. R. Evid. 803.
Fed. R. Evid. 902.
Friedman, R. D. (1998). Confrontation: The search for basic principles. Georgetown Law Journal, 86(4), 1011–1043.
Garrett, B. L., & Neufeld, P. J. (2009). Invalid forensic science testimony and wrongful convictions. Virginia Law Review, 95(1), 1–97.
In re Investigation of the West Virginia State Police Crime Laboratory, Serology Division, 190 W. Va. 321, 438 S.E.2d 501 (1993).
Kassin, S. M., Dror, I. E., & Kukucka, J. (2013). The forensic confirmation bias: Problems, perspectives, and proposed solutions. Journal of Applied Research in Memory and Cognition, 2(1), 42–52.
Melendez-Diaz v. Massachusetts, 557 U.S. 305 (2009).
Metzger, P. R. (2006). Cheating the Constitution. Vanderbilt Law Review, 59(2), 475–538.
Michigan v. Bryant, 562 U.S. 344 (2011).
Mnookin, J. L. (2007). Expert evidence and the Confrontation Clause after Crawford v. Washington. Journal of Law and Policy, 15(2), 791–?.
Mosteller, R. P. (Ed.). (2020). McCormick on evidence (8th ed.). West Academic.
National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press.
Ohio v. Clark, 576 U.S. 237 (2015).
Ohio v. Roberts, 448 U.S. 56 (1980).
Palmer v. Hoffman, 318 U.S. 109 (1943).
Smith v. Arizona, 602 U.S. 779 (2024).
United States v. Lamons, 532 F.3d 1251 (11th Cir. 2008).
United States v. Oates, 560 F.2d 45 (2d Cir. 1977).
United States v. Washington, 498 F.3d 225 (4th Cir. 2007).
Williams v. Illinois, 567 U.S. 50 (2012).
