Understand Transparency in Polling

A Commitment to Transparency

Since the birth of polling, the field has embraced transparency as a core value. Early commercial pollsters openly published details about their methods.1,2,3 The Roper Center was established in 1947 as the first social science data archive for data-sharing purposes. The field adopted professional disclosure standards back in the late 1960s, paving the way for the AAPOR Transparency Initiative in 2014. 5,6 Transparency, then, is not a fad: pollsters have long recognized that it provides the information that is essential for judging a poll’s quality and fit for purpose. 

Disclosure standards are not methodological standards; they do not ensure quality, but rather enable quality to be independently judged and verified. Institutional commitment to transparency – like AAPOR’s Transparency Initiative, Roper Center’s Transparency Project, CNN’s transparency requirements for reporting on surveys, and scientific journals’ data-sharing requirements – create a more trustworthy and robust research environment.

Transparency Is Not Enough

Transparency is necessary to allow evaluation of polls, but it is not sufficient: disclosure means nothing if the reader does not understand what the disclosed information means. Journalists, policy makers, students, and the public all read polling results and should have the tools to understand the potential strengths and weaknesses of any survey. Unfortunately, most people lack any training to understand how to interpret information mandated by disclosure standards. Even those with a solid understanding of basic polling principles may be unaware of, or confused by, certain details in a comprehensive survey report. This is entirely understandable: survey methods are changing rapidly, and it is difficult to keep up with the changing methods and their possible effects. 

To address this issue, in the summer of 2026, the Roper Center and AAPOR supported a W.E.B. DuBois Scholar to develop a guide to understanding the elements of disclosure. Using the AAPOR Transparency Initiative and Roper Center Transparency Project Scoring as a framework, Du Bois Scholar Nancy Toure  (University of Illinois, Chicago) has created this introduction and guide to understanding polling disclosure.

  1. Gallup, G. (1944). A guide to public opinion polls.
  2. Belden, J. (1939). Measuring college thought. The Public Opinion Quarterly, 3(3), 458-462.
  3. Roper, E. (1941). Checks to increase polling accuracy. Public Opinion Quarterly, 87-90.
  4. Geda, C. (1979). Social science data archives. The American Archivist, 42(2), 158-166.
  5. Gollin, AE 1992 ‘AAPOR and the media’, in PB Sheatsley & WJ Mitofsky (eds.) 1992. A meeting place: The history of the American Association for Public Opinion Research, American Association for Public Opinion Research.
  6. Link, M. W. (2015). Presidential address: AAPOR2025 and the opportunities in the decade before US.

By Nancy Toure Ph.D.

Additional contributors to this project include Jonathon Schuldt, Roper Center Executive Director; Kathleen Weldon, Roper Center Director of Data Operations and Communications

Guide to Transparency in Polling

Transparency is necessary to allow evaluation of polls, but it is not sufficient: disclosure means nothing if the reader does not understand what the disclosed information means. The guide is intended to provide a basic overview, in straightforward and comprehensible language, of ways to think about each element of survey disclosure.

The guide is intended to provide a basic overview, in straightforward and comprehensible language, of ways to think about each element of survey disclosure. The guide is not intended to provide value judgments on methodology nor a checklist of settled best practices; in a complex and ever-evolving field with many unresolved debates, this is not possible. A methodology suited to a particular time, place, and purpose might be entirely inappropriate in another context. Instead, for each piece of information that AAPOR and Roper have identified as important for transparency and disclosure, the guide offers:

  • A plain language definition 
  • Why it matters for interpreting results
  • What effective disclosure looks like
  • What incomplete disclosure looks like
  • Strengths when well-disclosed
  • Limitations to keep in mind
  • Variations that exist
  • Related elements
  • Resources to learn more

This guide should be considered a living document, updated over time to reflect changes in the field. 

One limitation is worth noting. In standard methodological statements, lack of disclosure may be hard to differentiate from a lack of items to disclose. For example, an organization may have used a subcontractor, an online survey router, or an incentive for participation and not disclosed that information – or, it may simply not have used any of these and therefore did not disclose them. The AAPOR Transparency Initiative and Roper Center Transparency Project provide a structure for reporting that ensures comprehensiveness of information and demonstrates a commitment to full disclosure. In some cases, organizations have adopted the elements of these disclosure standards as an outline for their survey information in releases, providing an excellent model for clarity in reporting.

1. Plain-Language Definition

The survey organization is the entity that conducted the fieldwork: the company, university, government agency, or other body that designed the data collection procedures, recruited respondents, administered the questions, and produced the raw dataset. This is often distinct from the entity that paid for the survey (the sponsor) or released the findings publicly. For example, a news organization might publish poll results from survey data that were collected by an outside research firm that was hired by the news organization. In this case, the news outlet is the sponsor, but the research firm is the survey organization.1

Knowing the survey organization matters because different types of organizations operate under different professional norms, incentive structures, and levels of public accountability. Academic survey labs, commercial polling firms, government statistical agencies, news organization polling units, and in-house advocacy research teams each bring distinct priorities and constraints to the research process. AAPOR's Transparency Initiative, launched in 2014 and now the leading industry standard, requires organizations releasing survey findings to publicly disclose who conducted the research alongside who sponsored it.2,1

2. Why It Matters for Interpreting Results

Knowing the survey organization is useful for evaluating data quality because it conveys information about the professional standards guiding the research. Organizations that are members of AAPOR's Transparency Initiative have pledged to disclose certain methodological details whenever they release findings, including sample design, data collection modes, dates, and response rates. Organizations that are not members of the initiative may still adhere to rigorous standards, but users cannot verify this without additional investigation.2

Research on polling performance shows that individual organizations can develop characteristic 'house effects,' which are systematic tendencies in their results attributable to their specific procedures and practices. Knowing which firm collected the data provides useful information about what patterns to expect and how to contextualize findings against surveys by other organizations.3

Users who ignore this element risk treating different polls as equivalent, when they are unlikely to be. The organization conducting fieldwork controls crucial design decisions (how the sample is drawn, how interviewers are trained, how quality checks are performed) that shape data quality independently of a sponsor's research interest. For example, willingness to accept a 'don't know' response can result in different results across survey organizations depending on how interviewers are trained.4

3. What Effective Disclosure Looks Like

Effective disclosure names the organization specifically and in a form that allows independent verification: a full legal or registered name rather than an abbreviation, a trade name, or a vague description. According to AAPOR's Transparency Initiative disclosure checklist, publicly released survey results should identify both who sponsored the research and who conducted it as two distinct pieces of information. Effective disclosure includes the conducting organization's full name, its organizational type (academic, commercial, government, etc.), and whether a subcontractor performed any of the data collection.1

Well-documented records typically state the conducting organization by name in the methodology notes or topline document, distinguish it from the client or sponsor, and provide enough detail for a researcher to independently verify the organization's standing. Records from AAPOR Transparency Initiative members are especially useful because membership requires ongoing commitment to the full disclosure checklist.1

4. What Incomplete Disclosure Looks Like

Generic descriptions such as 'a national research firm' or 'our research team' hinder independent verification. Listing only the sponsoring organization when a separate firm actually conducted the fieldwork is another common gap. Readers usually cannot tell this just by reading the report; both lack of disclosure and lack of anything to disclose can look the same. Users should consider whether an organization has promised to share this kind of detail, for example by joining a group like AAPOR's Transparency Initiative. Roper Center's Transparency Score will also indicate whether the organization has confirmed the lack of further information to disclose with the archive.1,5

An organization name with no website, affiliation, or indication of AAPOR Transparency Initiative membership leaves users unable to verify the organization's track record. A well-known brand name used without clarifying whether a subsidiary or subcontractor did the actual fieldwork is also a concern, and can be difficult to detect1

AAPOR's Code specifies that good professional practice requires disclosure sufficient to allow independent review and verification of research claims, which requires identification of the field organization conducting the survey.1 

5. Strengths When Well-Disclosed

Polling scholars, methodologists, and data journalists evaluate poll quality substantially based on the reputation and track record of the survey organization. Thorough disclosure enables the kind of informed interpretation that experts themselves apply.4

Clear identification supports accountability. If an organization's methodology turns out to be flawed or its results systematically biased, the named entity can be identified and followed up.4

6. Limitations to Keep in Mind

An organization's name does not reveal whether they followed AAPOR standards on a particular study, whether their methodology has changed over time, or whether quality controls on a specific project matched their usual practices.1

Some organizations conduct both advocacy-oriented work and independent research under the same identity. Users benefit from also examining who sponsored the specific study.1

Well-known organization names are sometimes used to lend credibility to studies where a lesser-known subcontractor actually performed the fieldwork.1 This can be very difficult to detect if not explicitly disclosed.

Organizational reputation is a lagging indicator. An organization may have an excellent historical track record while current practices have changed due to budget pressures, staff turnover, or shifts in data collection technology.4

The survey organization field is most useful when considered alongside other methodological disclosure elements rather than as a standalone quality signal.1

7. Types of Survey Organizations You May Encounter

Academic and university-based organizations

Academic survey centers typically operate under Institutional Review Board oversight and adhere to academic publishing norms, including peer review expectations and open data commitments.6

Commercial polling firms

Commercial firms operate under contractual relationships with clients, meaning their disclosure practices can vary considerably depending on whether the client authorizes full transparency.1,7

Government statistical agencies

Government statistical agencies operate under legal mandates that require rigorous methodological documentation and public access. Their surveys are generally considered thoroughly documented, with authority typically derived from federal statute and oversight bodies such as the Office of Management and Budget.8,9

News organization polling units

These units combine journalistic editorial practices with research operations. This can provide strong public accountability but may also involve editorial decisions about which findings to report.4,10

In-house polling by advocacy organizations or think tanks

Because the same organization is both advocating for a cause and producing the research on it, studies touching on that cause warrant closer scrutiny of the specific study's design.1

Co-sponsored fieldwork

Sometimes two or more organizations collaborate on sponsoring and conducting research. In such cases, it is useful to check which organization was operationally responsible for the fieldwork itself.2

8. Related Elements

Survey Organization is most closely paired with External Survey Sponsor and Grant Funding Source. Together these three fields answer who did the work, who paid for it, and who commissioned it. Discrepancies or overlaps among them are often the most diagnostic signal for assessing potential conflicts of interest.1

Survey Organization also connects directly to Data Collection Mode and Sample Design, since the conducting organization's specific capabilities (whether it maintains a probability-based panel, employs trained interviewers, or relies on opt-in recruitment) determine what modes and sampling strategies are available for a given study.1

9. Further Reading

AAPOR Code of Professional Ethics and Practices, April 2021, Section III. aapor.org/standards-and-ethics/disclosure-standards/

Jackman, S. (2005). Australian Journal of Political Science 40(4):499-517. doi:10.1080/10361140500302472

Hillygus, D.S. (2011). Public Opinion Quarterly 75(5):962-981. doi:10.1093/poq/nfr054

Edwards, Dillman & Smyth (2014). Public Opinion Quarterly 78(3):734-750. doi:10.1093/poq/nfu027

AAPOR Transparency Initiative member list. aapor.org/standards-and-ethics/transparency-initiative/

Roper Center Data Provider list: https://ropercenter.cornell.edu/membership/our-data-providers

 

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 2. aapor.org/standards-and-ethics/disclosure-standards/

2. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

3. Jackman, S. (2005). Pooling the Polls Over an Election Campaign. Australian Journal of Political Science, 40(4), 499-517.

4. Hillygus, D.S. (2011). The Evolution of Election Polling in the United States. Public Opinion Quarterly, 75(5), 962-981.

5. Roper Center Transparency Score. ropercenter.cornell.edu

6. AAPOR Institutional Review Boards guidance. aapor.org/standards-and-ethics/institutional-review-boards/

7. British Polling Council Rules of Disclosure. britishpollingcouncil.org

8. U.S. Census Bureau Statistical Quality Standards. census.gov/about/policies/quality/standards.html

9. Cecco, K. & Cohen, S. (2004). Standards for Statistical Surveys in the U.S. Statistical System. BLS working paper.

10. Kaiser Family Foundation. The Role of Public Opinion Polls in Health Policy (Health Policy 101). kff.org

1. Plain-Language Definition

Funding disclosure tells the reader who paid for a survey. Most media polls, for example, have a sponsoring media organization and a field organization paid by the media organization. However,  a single academic or non-profit survey can involve as many as three distinct parties: the grant funder that supplied the original money, the external sponsor that commissioned the specific survey, and the survey organization that designed and fielded it. Because the grant funder and the external sponsor often overlap, and because AAPOR discloses them together under a single standard, this guide treats them as one combined element while noting where the two genuinely differ.1

A grant funding source is the organization that provided money, through a formal grant award, to support the research. Grants are typically awarded through a competitive application process, and the funder may have had no direct hand in designing the specific survey questions or interpreting the results. If the money arrived as a grant and the funder did not commission or shape the instrument, it belongs under grant funding source.1

An external survey sponsor, by contrast, is the organization that commissioned a particular survey, especially if it shaped what got asked. If a single organization funded, designed, and fielded a poll, no separate sponsor is listed. The practical difference is one of involvement: sponsors usually have more direct influence over what questions are asked, how they are framed, and which results are publicized, while grant funders provide financial support with less direct control over design. Overlap is common. For example, a grant can be awarded to a sponsoring organization that then hires a survey firm.1

The type of sponsor matters because it signals the kind of stake the sponsor may have in the outcome. Corporate or industry sponsors commission research tied to commercial interests. Political organizations and campaigns sponsor polling tied to electoral or messaging goals. Advocacy and nonprofit sponsors fund work aligned with a cause. Government agencies sponsor surveys to inform policy. Media organizations sponsor polls for news coverage. Academic institutions sponsor research as part of their scholarly mission. Identifying which type is involved is the first step in judging whether the sponsor had a reason to prefer a particular result.1,2

2. Why It Matters for Interpreting Results

Funders shape research agendas by deciding which questions are worth studying, whether the money arrives as a grant or a direct commission. A funder with a stake in a particular outcome may support surveys on some topics but not others.3

The federal government recognized this risk early. The first rules on scientific conflicts of interest, issued together by the Department of Health and Human Services and the National Science Foundation in 1995 (now codified at 42 CFR Part 50 Subpart F), were written to ensure that the design, conduct, and reporting of grant-funded research would not be biased by investigators' conflicting financial interests.4 Grant funders also shape agendas more quietly through priority-setting: a study of five major public-health funders found that they move potential topics through four phases (idea generation, analysis, socialization, and selection), often involving researchers but less often the public or advocacy groups. In other words, funders influence what gets studied even when they do not touch methodology.3

On the sponsor side, there is strong evidence across many research fields that who pays for a study can affect the conclusions it reaches. A 2003 systematic review in JAMA found that industry-sponsored research was significantly more likely to reach conclusions favorable to the sponsor, and was also associated with restrictions on publication and data sharing.5 A 2017 Cochrane systematic review confirmed that manufacturer-sponsored drug and device studies produced more favorable conclusions than studies funded by other sources, even after controlling for methodology.6 One researcher describes this broad pattern as evidence that bias may exist rather than proof that it does in any particular case; resolving any individual case requires further analysis of the methods and interpretation.7

Sponsor identity can even affect who answers in the first place. In an experiment, surveys carrying an in-state sponsor drew higher response rates than those carrying an out-of-state sponsor -  in one case, 12 to 15 percentage points higher. Different sponsors produced somewhat different respondent profiles by age and political affiliation.8 A reader who ignores the sponsor can therefore miss a source of differences in who chose to participate, unrelated to the survey organization's methods.8

3. What Effective Disclosure Looks Like

Thorough disclosure goes beyond naming a funder. It gives the full name of the funding organization (not just 'a federal agency'), states the funder's role (whether it provided financial support only or also had input on design), and lists every funder when the work was co-funded. 1,9

Federal grants set the highest bar. Since January 2023, a major federal data-sharing policy has required all grant applications that generate scientific data to include a plan for sharing it, a major expansion from the old policy that applied only to grants of $500,000 a year or more.9 The practical effect is that the underlying data from grant-funded surveys should eventually become public. Federal award records show what complete documentation looks like: a single public record lists the award number, recipient, program manager, dates, dollar amount, principal investigator, and a full research abstract, which is everything needed for independent verification.10 The federal conflict-of-interest regulation goes further still, requiring institutions to maintain a public policy on financial conflicts of interest and to report the grant number, investigator, entity, nature and value of any conflict, and the management plan.4

Good sponsor disclosure begins with naming the sponsor clearly on any public release. AAPOR folds this into a single standard that names the sponsor, the organization that conducted the survey, and, where it differs, the original funding source.1 One professional association of pollsters makes sponsorship a Level-1 disclosure requirement that members should carry in every public report,11 and the code of a major international polling association adds that respondents should be told who the sponsor is on request, and that each sponsor must be informed when one survey serves several clients.12 A model sponsor disclosure therefore names the sponsor, states whether it had a role in question design or approval, identifies the conducting organization, and notes any co-sponsors.13

4. What Incomplete Disclosure Looks Like

Several issues to be aware of apply to both grant and sponsor funding: no funder or sponsor named at all, even for academic research; vague language such as 'funded by a private foundation' with no name; and no statement of whether the funder or sponsor had any role in question design. A grant-specific concern is naming a funder but omitting the grant number, which makes verification difficult. Any missing piece of important information should make you cautious about other pieces of information in the disclosure as well.1,14

AAPOR treats omitting funder information on public results as a disclosure violation, since its standards require enough information for independent review.1 The scale of the problem is documented: one study found that about two-thirds of peer-reviewed journals published no positive financial-interest disclosures at all, and only 0.5 percent of articles carried personal ones.14 The Roper Center's Transparency Score reflects this directly: studies that omit funding receive lower scores.15 For grant-funded work, the consequences can be regulatory: if an investigator fails to disclose a conflict, the institution must conduct a retrospective review and, if bias is found, notify the funding agency and submit a mitigation plan, with funding suspension possible for noncompliance.4

A concern distinctive to sponsors is identity concealment through an intermediary. Some corporations fund charities or industry groups that in turn support research and promote industry-favorable positions. One well-documented case involved an industry-funded group working to influence research, conferences, messaging, and policy on behalf of its corporate members, to the point that the World Health Organization withdrew from official relations with it.16 The lesson is to be wary of a seemingly independent group whose funding traces back to interested parties. 

A further caution relates to the publication of survey results in academic journals. Although the widespread requirements for conflict-of-interest disclosure by authors would be expected to ensure the highest level of transparency, having a disclosure policy does not necessarily result in meaningful disclosure. A 2025 scoping review found that conflict-of-interest publication policies, though widespread, are inconsistently enforced and not especially effective on their own.17

5. Strengths When Well-Disclosed

A disclosed federal grant signals likely data availability, because the 2023 federal data-sharing mandate requires the underlying data to become publicly accessible, and a disclosed grant number lets anyone look up the original award and read the funder's own description of the research goals through publicly searchable federal award databases.9,10 No comparable public database exists for privately sponsored research. Federally funded work is also subject to legally mandated bias review through the conflict-of-interest regulation, and it tends to carry fewer restrictions on publication and data sharing than industry-sponsored research.4,6

On the sponsor side, the mirror image of the funding-effect evidence is itself informative: non-industry funding is associated with less outcome distortion than industry sponsorship, so knowing a survey was independently sponsored carries weight in its own right.6 Membership in AAPOR's Transparency Initiative functions as a trust marker. You can check whether an organization is a member on AAPOR's website to verify their commitment to disclosing the full funding chain.1 Whenever funding is clearly disclosed, it helps a reader judge whether the funder or sponsor had a stake in the topic.1

6. Limitations to Keep in Mind

A funder is not the same as the organization that designed the survey and may have had no role in question wording. Even so, funding is not bias-free: funders shape agendas through priority-setting, with strategic programs proposing specific topics and even investigator-driven programs filtered through funder criteria.3 Grant funding tends to cover a defined period, whereas other kinds of sponsor relationships may be ongoing or a one-off.3

Most importantly, funding source is probabilistic, not deterministic. It flags a risk, not a guaranteed problem, and the absence of a listed grant funder does not prove independence, since the work may be commercially funded without disclosure. Researchers caution that strong overall evidence for a funding effect does not let you conclude any single study is biased.5,7 On the sponsor side, two further limits are worth keeping in mind: a sponsor's real identity can be hidden behind an intermediary,16 and sponsor effects operate partly through how respondents perceive who is asking, not only through wording. Experiments have shown the named sponsor changed who chose to respond.8 A faithfully disclosed sponsor can therefore still influence results through participation alone.8

7. Types of Funding You May Encounter

Federal government grants are the most consistently transparent category: award details are publicly searchable and recipients are bound by both data-sharing and conflict-of-interest rules.9,4 Private foundation grants reflect a foundation's programmatic priorities, which shape what gets studied without necessarily directing methodology; they are less publicly verifiable, though some foundations publish grant lists and basic financials can often be found through publicly available foundation databases.3 University or institutional funding is generally independent but less thoroughly documented and can operate through a number of different mechanisms.1

Among sponsors, the type carries different implications. Corporate or industry sponsors have the strongest funding-effect evidence behind them. Systematic reviews have found their studies markedly more likely to reach favorable conclusions, partly because industry tends to support study designs that favor positive results.5,6 Government-agency sponsors tend to obtain higher response rates.8 Advocacy and nonprofit sponsors can be straightforward, but as documented cases show, a seemingly independent sponsor may itself be funded by interested industry parties, so the label alone does not settle the question of independence.16 Political-campaign and media sponsors are also common, and each shapes a poll's purpose and presentation in characteristic ways.1

Some surveys simply do not have an external funder or sponsor. Advocacy groups, trade associations, and media companies often commission surveys directly using their own operating budgets. This is not necessarily a problem, but it shifts attention to the survey organization itself and whether it has incentives tied to the outcome. Co-funded studies and multi-organization collaborations, particularly between media and academic partners, are common. Because each funder may have different interests, good disclosure spells out each one's role.1,8

8. Related Elements

AAPOR groups grant funding source with two related elements in a single disclosure item, and all three must be read together: the survey organization that conducted the fieldwork and the external sponsor that commissioned the survey. The most consequential link is to full question wording, because funder and sponsor priorities can influence how questions are framed.1

Sponsor identity also shapes who participates, not only what is asked. Experiments found that in-state and government-affiliated sponsors drew higher response rates than out-of-state or unknown ones.8 It is worth distinguishing knowing the funder or sponsor from knowing the field organization: the field organization tells you about methodology, while the funder or sponsor tells you about potential bias and research agenda. They are different kinds of information.1,8

9. Further Reading

AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A. aapor.org/standards-and-ethics/disclosure-standards/

Bekelman, J. E., Li, Y., & Gross, C. P. (2003). Scope and impact of financial conflicts of interest in biomedical research. JAMA, 289(4), 454-465.

Lundh, A., et al. (2017). Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews.

Edwards, M. L., Dillman, D. A., & Smyth, J. D. (2014). Public Opinion Quarterly, 78(3), 734-750.

NIH Data Management and Sharing Policy (2023). sharing.nih.gov/data-management-and-sharing-policy

Jamieson, K. H., et al. (2023). Protecting the integrity of survey research. PNAS Nexus, 2(3).

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 2. aapor.org/standards-and-ethics/disclosure-standards/

2. Roper Center Transparency and Acquisitions Policy. ropercenter.cornell.edu/roper-center-transparency-and-acquisitions-policy

3. Mador, R., et al. (2018). How research funders set priorities. Health Research Policy and Systems, 16:49.

4. 42 CFR Part 50 Subpart F (Financial Conflicts of Interest), §§50.601-50.607. ecfr.gov

5. Bekelman, J. E., Li, Y., & Gross, C. P. (2003). Scope and impact of financial conflicts of interest in biomedical research. JAMA, 289(4), 454-465.

6. Lundh, A., et al. (2017). Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews.

7. Krimsky, S. (2013). Do financial conflicts of interest bias research? PLoS Medicine, 10(2).

8. Edwards, M. L., Dillman, D. A., & Smyth, J. D. (2014). Public Opinion Quarterly, 78(3), 734-750.

9. NIH Data Management and Sharing Policy (2023). sharing.nih.gov/data-management-and-sharing-policy

10. NSF Award Search. nsf.gov/awardsearch

11. Miringoff, L. M. & Carvalho, B. L. National Council on Public Polls. In Encyclopedia of Survey Research Methods.

12. WAPOR Code of Professional Ethics and Practices (2021), Section IV.

13. Pew Research Center (Kennedy, C., Popky, D., & Keeter, S., 2023). How Public Polling Has Changed in the 21st Century.

14. Krimsky, S. & Rothenberg, L. S. (2001). Conflict of interest policies in science and medical journals. Science and Engineering Ethics, 7(2), 205-218.

15. Roper Center Transparency Score. ropercenter.cornell.edu

16. U.S. Right to Know. Case study of the International Life Sciences Institute (ILSI). usrtk.org

17. Graham, S. S., et al. (2025). Research on policy mechanisms to address funding bias and conflicts of interest in biomedical research: a scoping review. Research Integrity and Peer Review, 10(1):6.

 

1. Plain-Language Definition

The universe, also called the target population, is the entire group of people whose opinions or behaviors a survey is trying to measure. It might be all adults living in the United States, all registered voters in Ohio, all parents of school-age children in Italy, or any other clearly defined group. When a survey reports that '62 percent of Americans think,' the universe is U.S. adults, and every number in the report is meant to represent that group, unless the report specifically indicates that a result is based on a subpopulation.1,2

Geographic coverage is the physical area from which the survey collected its data. A survey of 'U.S. adults' should draw respondents from all 50 states and the District of Columbia. A survey of 'adults in the Midwest' covers a smaller area. Geographic coverage is closely tied to the universe because the boundaries of the area help define who is and is not part of the target population.3,4

AAPOR combines these two concepts into a single disclosure item called 'Population Under Study,' which asks researchers to describe both the target population and the geographic coverage of the study. That is why this guide treats them together, because they are two sides of the same question: who exactly does this survey represent?3

Defining the universe precisely matters because survey results can vary depending on which group is being described. For example, polls about political topics may report results for all U.S. adults, registered voters, or likely voters, and these groups can have noticeably different opinions. Only about two-thirds of adults are registered to vote, and registered voters tend to be more politically engaged than those who are not. Only about a third of adults voted in the last midterm election, so a report based on likely midterm voters describes a much smaller slice of the population.1

2. Why It Matters for Interpreting Results

If you do not know the universe, you cannot judge what the survey results mean. A finding that '58 percent support the policy' could describe all adults, only people who voted in the last election, or only people with internet access. These are very different groups, and the same percentage can mean different things depending on which one was surveyed.1,2

Geographic coverage matters for similar reasons. A survey described as 'national' that excludes Alaska and Hawaii, or that draws most of its respondents from urban areas, does not truly represent the entire country. For many years after Alaska and Hawaii joined the union in 1959, some national surveys continued to exclude these two states from their samples for logistical reasons; this practice has been documented in academic work that reconstructs major national polling datasets from that era. Readers can sometimes check for this directly: where the underlying dataset is available, it generally includes a state variable, which can show whether respondents from these states were included. When a survey's sampling frame (the list or method used to contact potential respondents) does not fully cover the geographic area it claims to represent, the result is called coverage error. Coverage error means some people in the target population had no chance of being included in the survey.5,6

This is not just a theoretical concern. Telephone surveys that rely on area codes to target a specific city or county can include people who live outside the intended area, because cell phone numbers are not tied to where a person lives. Mail surveys that use address lists can miss people in very rural areas or in new housing developments not yet on the postal delivery file. 7,8A well-designed mail survey can partly guard against this by asking respondents to confirm their address or residence as part of the questionnaire itself, which lets researchers verify that responses actually came from within the intended geographic area.

Researchers who study survey quality emphasize that clearly specifying the target population and ensuring the sample adequately covers it are among the most fundamental requirements for a trustworthy survey. Without this, even perfect question wording and high response rates cannot make the results meaningful.4,9

3. What Effective Disclosure Looks Like

Effective disclosure names the target population precisely and in plain language. Instead of 'a representative sample,' a well-disclosed survey says something like 'a nationally representative sample of 5,733 U.S. adults ages 18 and older.' It specifies any eligibility criteria: age, citizenship, language, internet access, or other characteristics that determined who could participate.3,10

For geographic coverage, effective disclosure states the exact boundaries of the area covered. If a survey covers the entire United States, it says so. If it covers only the 48 contiguous states, it notes that Alaska and Hawaii are excluded. If it is a state or local survey, it names the state, county, or metropolitan area.3,4

Effective disclosure also describes any groups within the target population that the sampling method could not reach. For example, most U.S. household surveys exclude people living in institutions (prisons, nursing homes, military barracks), and online surveys cannot reach adults without internet access unless the survey organization provides that access. Noting these exclusions is important because excluded groups may hold different views than those who were included.2,5

Reporting guidelines for survey research in academic journals recommend that the target population, the sampling frame, and any exclusions be described clearly enough for another researcher to evaluate how well the sample represents the population it claims to describe.9,11

4. What Incomplete Disclosure Looks Like

The most common problem is vague language. Descriptions like 'a survey of Americans' or 'a nationwide poll' do not tell you whether the target population is all adults, registered voters, likely voters, or some other group. Without this detail, readers may assume the results apply more broadly than they actually do.1,3

Another issue is failing to describe exclusions. If a survey is described as representing 'U.S. adults' but was conducted entirely online, it effectively excludes the roughly 7 percent of American adults who do not use the internet. If this exclusion is not mentioned, readers may not realize the results leave out a portion of the population.5,10

On the geographic side, vague descriptions like 'a regional survey' or 'conducted in the southern United States' provide too little information. Which states were included? Were rural areas covered, or only metropolitan areas? Without specifics, there is no way to judge whether the sample truly represents the region it claims to cover.4,8

A more subtle problem occurs when the geographic label does not match the actual sample. A survey labeled 'statewide' that drew most of its respondents from one metropolitan area would overrepresent urban views and underrepresent rural ones, even if the overall sample size looks adequate.6,7

5. Strengths When Well-Disclosed

When the universe and geographic coverage are clearly described, readers can immediately judge whether the survey results apply to the group they care about. A journalist writing about national opinion can confirm the poll actually covers the whole country. A policy researcher studying rural health can check whether rural areas were included in the sample.1,3

Clear population definitions also make it possible to compare surveys meaningfully. If one poll reports results for 'all adults' and another for 'registered voters,' a reader who knows this can avoid the mistake of treating the two as directly comparable. This kind of careful comparison is exactly what experts do when evaluating conflicting poll results.1,4

Transparent disclosure of exclusions lets readers factor in potential blind spots. If a survey notes that it could not reach adults without internet access, an informed reader knows the results may underrepresent older, lower-income, or rural populations who are less likely to be online.5,10

6. Limitations to Keep in Mind

Even well-disclosed universe definitions have limits. A survey that targets 'all U.S. adults' still cannot reach every single person. Some people are homeless, some live in institutions, some do not speak the languages the survey is offered in, and some simply cannot be contacted by any available method. The target population is an ideal; the actual sample is always an approximation of it.2,5

Geographic coverage can be misleading if the sampling method does not match the stated area. A survey that claims national coverage but uses a sampling frame that underrepresents certain regions may not truly reflect the whole country, even if some respondents from every state happen to be included.6,8

The universe definition itself can shape results. Defining the target population as 'likely voters' rather than 'all adults' can dramatically change the findings on political questions, because likely voters tend to be older, wealthier, and more politically engaged than the general adult population. The choice of universe is not neutral. It reflects decisions about whose opinions count for the purpose of the survey.1 Likely voter screens vary across pollsters and can meaningfully shift results. The New York Times Upshot polling project demonstrated this directly by giving identical raw polling data to four outside pollsters, who applied their own likely-voter and weighting choices and arrived at results ranging from Clinton +4 to Trump +1 on the same underlying interviews.13

Some target populations are particularly difficult to reach for a number of reasons. In some cases the population is small and widely dispersed across the country, such as people with a rare disease. In other cases the population faces specific phone or internet access challenges, such as people experiencing unstable housing or extreme poverty. For populations that are highly concentrated in certain areas, such as ethnic groups with a large presence in a small number of cities or states, researchers can poll those geographies in proportion to their share of the target population, a technique sometimes called geographic oversampling.14

Historical limitations

When Hawaii and Alaska gained statehood in 1959, most surveys were still being conducted with area probability methods, and the new states were rarely included. This exclusion continued into the early telephone sample period for some organizations. If a dataset is available, researchers can usually use a state variable to verify which states were included.

Gallup’s polling before the adoption of probability sample in 1952 relied heavily on previous election results in setting quotas, particularly for surveys focused on politics. Although these surveys were intended to be national adult surveys and described as such, this sampling approach created systematic bias that was meaningful for some topics. More detailed information on early U.S. Gallup sampling is available on every pre-1952 survey at Roper Center. Adam Berinsky and Eric Schickler developed new weights to make these surveys more representative of the full national adult population; these datasets are also available at Roper Center.15 

7. Types of Universes and Geographic Scopes You May Encounter

Common universe definitions

All adults (ages 18 and older) is the broadest commonly used universe for public opinion surveys. Registered voters narrows the group to those on voter rolls. Likely voters narrows it further to people the survey organization predicts will actually vote, based on past behavior, stated intentions, or both. Each of these produces a different respondent pool and potentially different results, even on the same questions.1

Specialized universes are common in health, education, and policy research. A survey might target parents of children under 18, adults with a specific medical condition, employees of a particular industry, or residents of public housing. The more specific the universe, the more important it is that the disclosure says exactly who qualified.2,3

Common geographic scopes

National surveys aim to cover the entire country. International surveys cover multiple countries, and each country may use different modes and sampling methods, which means geographic coverage can vary from country to country within the same study.10,12

State and local surveys cover a single state, county, city, or metropolitan area. These require especially clear disclosure because the boundaries may not be obvious. A 'New York City survey' might cover all five boroughs, only Manhattan, or everyone living in the NY-Newark-Jersey City Metropolitan Statistical Area (MSA), and the distinction matters.7,8

8. Related Elements

Universe and geographic coverage connect directly to sample design, because the sampling method determines whether the target population is actually covered. A telephone survey using random-digit dialing covers a different slice of the population than an address-based mail survey, even if both target the same universe.4,8

These elements also connect to survey mode. Online-only surveys cannot reach people without internet access, which narrows the effective universe. Phone surveys face challenges reaching people who do not answer calls from unknown numbers. Understanding the mode helps you judge whether the stated universe matches the population that had a chance of being included.5,6

Finally, universe and geographic coverage are closely linked to the survey organization and sponsor, because the organization's capabilities (whether it maintains a nationwide panel, employs interviewers in multiple regions, or relies on online recruitment) determine what universes and geographic areas it can realistically cover.3

References

1. Pew Research Center (2018). Defining the Universe Is Essential When Writing About Survey Data. pewresearch.org

2. Lavrakas, P. J. (Ed.) (2008). Encyclopedia of Survey Research Methods: Target Population entry. Sage Publications.

3. AAPOR Transparency Initiative Disclosure Elements (revised April 2021), Item 4: Population Under Study. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

4. AAPOR Best Practices for Survey Research, Section 2: Designing Your Sample. aapor.org/standards-and-ethics/best-practices/

5. Lavrakas, P. J. (Ed.) (2008). Encyclopedia of Survey Research Methods: Coverage Error entry. Sage Publications.

6. Groves, R. M. & Lyberg, L. (2010). Total Survey Error. Public Opinion Quarterly, 74(5), 849-879.

7. Lavrakas, P. J. (Ed.) (2008). Encyclopedia of Survey Research Methods: Geographic Screening entry. Sage Publications.

8. Harter, R., et al. (2016). Address-Based Sampling. AAPOR Task Force Report.

9. Bennett, C., et al. (2011). Reporting guidelines for survey research. PLoS Medicine, 8(8).

10. Pew Research Center. U.S. Survey Methodology. pewresearch.org/u-s-survey-methodology/

11. Sharma, A., et al. (2021). CROSS checklist for reporting of survey studies. Journal of General Internal Medicine, 36, 3179-3187.

12. Pew Research Center. International Survey Mode and Sample Design. pewresearch.org/methods/international-survey-research/survey-mode-and-sample-design/

13. Cohn, N. (2016, September 20). We gave four good pollsters the same raw data. They had four different results. The New York Times. https://www.nytimes.com/interactive/2016/09/20/upshot/the-error-the-polling-world-rarely-talks-about.html

14. Chen, S. & Kalton, G. (2015). Geographic Oversampling for Race/Ethnicity Using Data from the 2010 U.S. Population Census. Journal of Survey Statistics and Methodology, 3(4), 543–565.

15. Berinsky, A. J., Powell, E. N., Schickler, E., & Yohai, I. B. (2011). Revisiting Public Opinion in the 1930s and 1940s. PS: Political Science and Politics, 44(3), 515–520. 

1. Plain-Language Definition

Data collection dates are the start and end dates indicating when a survey was actually collecting answers from respondents. In research terminology this is often called the "field period," meaning the window of time when interviewers were making calls, online questionnaires were open, or mail surveys were being returned. This is different from when the survey was designed or when the results were published.1,2

The field period can be as short as a few hours for an online opinion poll or as long as several months for a large government study. The length of the field period depends largely on the purpose of the survey. A poll designed to capture the public's reaction to a breaking news event needs a short window so the responses reflect opinions in the immediate aftermath. General health or social surveys often take longer because the topics are less time-sensitive.1

AAPOR's Transparency Initiative requires that publicly released survey results include the dates of data collection (for example, 'data collection from January 15 through March 10 of 2019').2

2. Why It Matters for Interpreting Results

Knowing when a survey was fielded lets you place the results in context. Public opinion can shift quickly in response to major events like a presidential debate, a natural disaster, an economic crisis, or a court ruling. If you do not know that a poll was conducted in the days after a major news event, you might treat the results as reflecting a normal period when they actually captured a moment of heightened emotion or attention.1,3

Field period length also affects data quality. In general, the longer a survey is in the field, the greater the opportunity to achieve a higher response rate and a more representative sample, in part because early responders can sometimes be different from later responders. Longer field periods also allow for follow-up attempts and targeting of groups that are harder to reach.3 However, there are exceptions: many public opinion surveys deal with topics where attitudes can change in a matter of days, with political polls being a good example. In those cases, a long field period can mean that opinions shifted during collection, blurring the picture.3

Shorter field periods can also create challenges. When there is less time to make multiple contact attempts, there is greater potential for a sample that overrepresents people who are easy to reach, which can introduce bias. Statistical weighting can help correct for this, but it is not a perfect fix.1

For anyone comparing results across multiple surveys (a common practice in tracking public opinion over time), knowing the data collection dates for each survey is essential. You can only meaningfully compare different surveys if you know when each was fielded and whether anything significant happened between them.3,4

3. What Effective Disclosure Looks Like

Effective disclosure provides both the start date and the end date of data collection. A well-documented methodology statement might read: “The data collection field period for this survey was March 23 to March 29, 2026” or “ballots were mailed from September 1-September 5, 1990, and returned ballots were accepted until November 15” or “interviews were conducted January 2-25, 2003 in the UK and January 2-23 in Japan." "This level of detail allows a reader to check whether any notable events occurred during that window.5

Reporting guidelines for survey research in academic journals call for describing all relevant dates.6 For example, a checklist developed for standardizing survey reporting recommends providing information about the survey's time frame, including periods of recruitment, exposure, and follow-up days.7

For surveys that use multiple modes of data collection (such as online and phone together), effective disclosure also notes when each mode was active, since different modes may have been in the field at different times. When a survey involves multiple waves (repeated rounds of data collection from the same or similar samples), each wave's dates should be reported separately.2,5

4. What Incomplete Disclosure Looks Like

The most common problem is providing a start date without an end date, which makes it impossible to know how long the survey was in the field or whether any significant events occurred before collection ended. Another challenge is a rolling poll, that is, ongoing polls conducted daily over a longer time period in which results are compiled for some subset of dates, which do not clearly indicate which dates correspond to which reported results, making it hard to connect specific findings to a specific moment in time.3

Vague descriptions like 'conducted in early 2025' or 'fielded last spring' are not sufficient for meaningful interpretation. Readers need exact dates to check what was happening in the world while data were being collected. Another warning sign is when a methodology statement discusses the survey design in detail but omits the field dates entirely. Transparency standards emphasize that disclosure must be detailed enough for independent evaluation of the research.2,3

Unusually long field periods without explanation can also be a concern. If a political opinion survey was in the field for three months, that is atypical and worth questioning, since opinions on political topics can shift dramatically over that time frame. Similarly, an unusually short field period, such as a single day for a complex topic, may suggest the survey did not allow enough time for callbacks to hard-to-reach respondents.1,3

Historical disclosure

It should be noted that current transparency standards requiring exact field dates did not always exist. Surveys conducted from the 1930s through the late 1960s often reported only the month of fieldwork, or gave vague descriptors such as “early this summer” or “in recent months.” While contextual information in the original reporting may provide additional clues about field periods, this information may in some cases be lost. These older surveys may still be of some value, particularly if the attitudes measured were likely to be fairly stable and not responsive to specific events. However, without exact field dates, caution should be used in analyzing such data.

5. Strengths of Effective Disclosure

When data collection dates are clearly reported, readers can place the survey results alongside a timeline of real-world events and judge for themselves whether those events might have influenced responses. This is one of the simplest but most powerful tools for evaluating any poll.3,4

Well-disclosed dates also allow researchers to compare surveys conducted at different times on the same topic, which is essential for tracking trends in public opinion. If two surveys ask the same question but were fielded six months apart, the difference in results might reflect genuine changes in opinion rather than methodological differences.2,3

Clear and accurate disclosure of data collection dates supports the broader goal of transparency in survey research. AAPOR's best-practices guidance emphasizes that because there are many different ways to run surveys, it is important to be transparent about how a survey was run and analyzed so that people know how to interpret and draw conclusions from it.8

6. Limitations to Keep in Mind

As noted above, long field periods introduce more variability because opinions may shift during data collection, especially for surveys that span major events. Readers should consider the field period length in light of how time-sensitive the survey topic is.1,3

Surveys fielded immediately after major events capture a volatile moment that may not represent steady-state opinion. A poll taken in the first 48 hours after a crisis may look very different from one taken a week later, even if both are technically well-conducted.3

Even with perfect date disclosure, the dates alone do not tell you everything. You also need to consider how long the field period was relative to the survey mode. A mail survey that took six weeks is normal, while an online survey that took six weeks might be unusual. Dates are most useful when considered alongside other disclosure elements like mode and response rates.1,2

7. Common Types of Data Collection Schedules 

Single-day collection is common for overnight opinion polls, which aim to capture a quick snapshot of public reaction. These are most often conducted in relation to a planned event, such as a political debate or speech, since survey development and field operations can be arranged to follow the event immediately. Quick-turnaround surveys can also be conducted in response to unexpected major events, but these are more operationally complex and therefore less common. These are often conducted by phone or online and involve a large number of initial contact attempts s to compensate for the limited time available for callbacks.1

Multi-day windows of three to ten days are the most common schedule for standard opinion polls. This allows enough time for follow-up attempts while keeping the data collection period short enough that opinions are unlikely to shift dramatically.3

Extended fieldwork periods lasting weeks or months are typical of large government surveys and academic studies. These longer windows allow for more thorough follow-up and usually produce higher response rates, but they also mean that results may span different conditions or events. Longer field periods are also often needed to target harder-to-reach or low-incidence populations, such as caregivers, people with certain chronic illnesses, or those currently looking for work.3

Rolling or continuous collection is used in tracking polls, where interviews are conducted continuously and results are reported as moving averages. These require especially clear disclosure about which dates correspond to which reported figures.3

Multiple-wave collection involves the same or similar samples being surveyed at different points in time. Each wave has its own field dates, and good disclosure reports all of them separately.2

Baseline or benchmark collection is another important type. In election polling, a benchmark survey is typically commissioned when a candidate decides to seek office, collecting standard information about the candidate’s public image, positions on issues, and the demographics of the electorate in order to provide a baseline for evaluating the progress of a campaign. One challenge with benchmark surveys is their timing: the earlier the survey is conducted, the less likely respondents are to know much about the candidates, and the more likely it is that political and economic conditions will change substantially before Election Day.10 Beyond elections, the same logic applies to public education campaigns, planned policy announcements, or other events that could influence public opinion. A poll conducted before the campaign or event establishes a starting point, allowing researchers to measure its effect by comparing results before and after. Researchers analyzing historical media polls should be aware that certain topics may lack baselines because they are most often polled only when they have become newsworthy, which may not reflect attitudes during less salient times. For example, polling questions about airline safety are frequently asked after major crashes, which may temporarily make the public less trusting than during other periods.

8. Related Elements

Data collection dates are tightly linked to mode of data collection, because some modes inherently require longer field periods than others. Mail surveys take longer than online surveys simply because of the time needed for delivery and return. Phone surveys fall somewhere in between. Knowing the mode helps you judge whether a given field period is typical or unusual.2,9

Dates also relate to sample size and response rates: a longer field period generally allows more contact attempts and higher response rates, which in turn affects how representative the sample is. If you see a very short field period combined with a high sample size, the organization likely used a mode like online panels that can collect data very quickly.3,9

9. Further Reading

Beullens, K., et al. (2022). Journal of Survey Statistics and Methodology, 10(1), 161-182. academic.oup.com/jssam/article/10/1/161/6237202

Baker, R., et al. (2016). Evaluating Survey Quality in Today's Complex Environment. AAPOR report. aapor.org/wp-content/uploads/2022/11/AAPOR_Reassessing_Survey_Methods_Report_Final.pdf

Groves, R. M. & Lyberg, L. (2010). Total Survey Error: Past, Present, and Future. Public Opinion Quarterly, 74(5), 849-879.

References

1. Lavrakas, P. J. (Ed.) (2008). Encyclopedia of Survey Research Methods: Field Period entry. Sage Publications.

2. AAPOR Transparency Initiative Disclosure Elements (revised April 2021), Item 7: Dates of Data Collection. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

3. Baker, R., et al. (2016). Evaluating Survey Quality in Today's Complex Environment. AAPOR report.

4. Jamieson, K. H., et al. (2023). Protecting the integrity of survey research. PNAS Nexus, 2(3).

5. Pew Research Center (2026). Methodology: U.S. Role in the World survey. pewresearch.org/global/2026/04/28/methodology-us-role-world/

6. Bennett, C., et al. (2011). Reporting guidelines for survey research. PLoS Medicine, 8(8).

7. Sharma, A., et al. (2021). CROSS: Consensus-based checklist for reporting of survey studies. Journal of General Internal Medicine, 36, 3179-3187.

8. AAPOR Best Practices for Survey Research, Section 5: Analyzing and Reporting Results. aapor.org/standards-and-ethics/best-practices/

9. Groves, R. M., et al. (2009). Survey Methodology, 2nd ed. Wiley.

10. Asher, H. (2017). Polling and the Public: What Every Citizen Should Know (9th ed.), Chapter 7: Polls and Elections. CQ Press.

1. Plain-Language Definition

The justification for claims of representativeness is the description of the design choices a survey uses to make its sample resemble the larger population it is meant to describe. It answers a simple question: on what basis does the survey claim that the people who answered can stand in for the target universe? These choices can include how the sample was drawn (probability or non-probability), what frame, list, or panel it came from, and how the results were adjusted or weighted afterward to match known population figures.1

In AAPOR's Transparency Initiative disclosure scheme, this information lives mainly in the items covering how the sample was generated and recruited and how the data were weighted.1 Some archives, such as the Roper Center's iPoll, store representativeness as a distinct, per-study field alongside sample and weighting information.9 In plain terms, pollsters try to make a sample look like the population through some combination of sampling, quotas, and weighting, and this element records details about which of those tools were used and how.8

2. Why It Matters for Interpreting Results

The justification shapes what you can reasonably infer from a survey's results, because a claim of representativeness is only as informative as the design behind it, and the type of justification changes what the results can support.2

Probability-based samples can allow researchers to make statistically meaningful claims about a larger population within a calculable margin of error. Probability samples have, therefore, been relied upon for decades for the purpose of generating population-level inferences.  

At the same time, recent analyses show that some large, non-probability samples that are carefully designed, quota-balanced, and quality-filtered can align closely with estimates from probability samples and government-derived benchmarks, including longitudinal, state-level estimates of COVID-19 infection and vaccination rates tracked over more than two years. However, experts caution that the accuracy of non-probability surveys depends on several design choices, such as quotas on variables linked to the outcomes and steps to reduce nonresponse among low-trust respondents, rather than weighting alone, which in one analysis made little difference on top of the quota design.10 

Combining probability and non-probability samples within a single design is itself an active area of methodological research, rather than a strict either-or choice.11 While the performance of different sampling approaches is an ongoing area of scholarly inquiry and extensive debate within the field of survey research, knowing how a given survey justifies its representativeness helps you calibrate how much confidence to place on its estimates.

This disclosure requirement also reframes representativeness as something to examine rather than assume. AAPOR's survey-quality framework treats coverage, sampling, and nonresponse as questions a reader is meant to ask actively, because each can affect how closely a sample matches its universe.6 A stated justification gives you the material to ask those questions; its absence leaves you without a basis for judging the fit between sample and population.

3. What Effective Disclosure Looks Like

Effective disclosure states the basis for the representativeness claim rather than simply asserting it. According to AAPOR's disclosure elements, thorough documentation names the sample source and its population coverage and, where weighting is used, lists the variables and the sources of the benchmark figures.1 When those benchmarks are named (for example, the Census Bureau's American Community Survey or Current Population Survey), a reader can look them up independently.1

Well-documented records specify the sampling approach, the frame or panel, the weighting targets and their sources, and any acknowledged limitations. A Roper Center Transparency Score report illustrates this, placing fields for the justification for representativeness, the weighting benchmark source, and the variables used for weighting together as Core Elements of disclosure.9 For online and panel-based studies, AAPOR's guidance describes concrete metrics that disclosure can report to characterize how well a sample represents its target.7

4. What Incomplete Disclosure Looks Like

Incomplete disclosure asserts that a sample is representative without stating the model, frame, or benchmarks behind the claim. AAPOR's task force on non-probability sampling suggested that the label "representative" is difficult to support for opt-in sources when no adjustment model is stated.2 More recent AAPOR guidance reiterates and extends this point for online samples specifically.7 Reporting only a response rate, or offering no information at all, is a limited signal of representativeness, since a single figure does not describe how closely the sample matches the population.4

A related gap is vague or absent information about nonresponse and weighting. AAPOR's quality framework treats missing nonresponse and adjustment details as a point to watch, because without them a reader cannot tell how the achieved sample was shaped.6 In each case, the issue is not the use of a particular method but the absence of enough detail to evaluate the claim.

5. Strengths When Well-Disclosed

A well-specified justification can be checked rather than taken on trust. When a survey states its weighting variables and benchmark sources, a reader can verify those benchmarks against public data.1 Some approaches to representativeness can even be quantified: measures such as representativeness indicators (R-indicators) allow a well-specified claim to be expressed as a number and examined independently.4

Requiring an explicit justification also prompts providers to articulate and document their strategy rather than leave it implicit. Documented recruitment and weighting - for example address-based recruitment into a probability panel followed by random selection within it - make the chain of reasoning from sample to population easier to follow.8 A clear statement that a sample is not representative is itself useful information, because it tells a reader how to use the results.

A non-representative sample is not inherently a flawed or illegitimate one; its value depends on the purpose of the study. Survey experiments that compare a randomly assigned control group to a treatment group, for example, do not require the overall sample to be representative of the general population to support valid conclusions about the effect being tested, since the comparison of interest is between the two groups rather than to the population.13 Similarly, an initial survey of a nonrepresentative but hard-to-reach population, such as people with a rare condition or a stigmatized behavior, can supply information that helps researchers design a larger, more expensive representative study later, much as a focus group can. In some cases, non-probability approaches are the only practical or affordable way to gather any information on such populations at all.13

6. Limitations to Keep in Mind

Representativeness is multidimensional. A sample can match the population on demographics while still differing on attitudes or behaviors, so demographic balance does not by itself guarantee representativeness on the measures a survey is about.3 Post-stratification weighting improves demographic balance but cannot fully correct self-selection on characteristics related to the survey topic; in benchmark studies, weighting improved some non-probability surveys and worsened others.3

This limitation is not unique to non-probability samples. Probability-based surveys also rely on weighting to correct for coverage gaps and nonresponse, and that adjustment carries a similar risk: a weighting variable only reduces bias to the extent it is correlated with both the likelihood of responding and the outcome being measured, so weighting a probability sample can leave some estimates improved and others largely unchanged.14 In certain cases, weighting a probability sample can decrease the accuracy of estimates. Response rates are also a limited guide: across dozens of studies, the response rate was a weak predictor of nonresponse bias, meaning a high rate does not by itself confirm representativeness and a lower one does not rule it out.5 Finally, the status of representativeness claims for non-probability samples remains an active area of discussion in the field.2

7. Approaches to Justifying Representativeness You May Encounter

Justifications vary in how they connect the sample to the population, and each choice carries different implications for interpretation.7

Probability sampling with weighting. A probability-based sample adjusted to known population figures offers the most fully specified justification and the clearest chain of inference.2

Weighting to documented benchmarks. Results are adjusted to named external figures such as the ACS or CPS, with the variables and sources documented so they can be checked.1

Matching or propensity-score adjustment. For non-probability samples, statistical modeling is used to make the sample resemble a reference population.7

Benchmarking against external sources. Estimates are compared against known outside figures to show how well they align.4

Quota sampling with demographic targets. The sample is filled to preset demographic quotas intended to mirror the population.8

No explicit justification, or a stated disclaimer. Some records provide no justification, and others explicitly note that the results are not representative, an absence that is itself useful information for the reader.2

Historical context: Representativeness in polling from the beginning of the field in the 1930s to 1950s in the U.S. depended heavily upon quota-based samples to match population demographics, with some minor adjustments occasionally made by weighting the final sample. The field turned to probability sampling from the 1950s to the 1970s, when the advent of random-digit-dial telephone polling made probability sampling more affordable and feasible for even smaller polling projects. Probability became the primary justification for representativeness, with some weighting applied to improve accuracy. As response rates fell and online and non-probability panels became common, justifications shifted to more commonly include model-based adjustment, matching, and benchmarking rather than the sampling frame alone. This is worth keeping in mind when comparing older and newer studies.2

Note that the progression of polling methods outside the United States did not follow the same path as in the United States. While some countries like Australia adopted probability even earlier, in Great Britain and some other areas in-person quota polling remained dominant until replaced with online nonprobability polling.12

8. Related Elements

This element sits alongside the weighting elements, the sampling procedure, the sampling frame, and the estimated noncovered population, since representativeness is built from all of these together.1 AAPOR's disclosure elements spread the components across neighboring items, sampling, sample size and precision, weighting, and the limitations statement, so a full picture usually requires reading several fields together.1 AAPOR's quality framework groups coverage, sampling, and nonresponse as the three error sources that jointly determine whether a sample represents its universe.6

References

1. AAPOR Transparency Initiative — Disclosure Elements (revised April 2021), Items 5 and 9. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. Baker, R., et al. (2013). Report of the AAPOR Task Force on Non-Probability Sampling. aapor.org

3. MacInnis, B., Krosnick, J.A., Ho, A.S., & Cho, M.J. (2018). The Accuracy of Measurements with Probability and Nonprobability Survey Samples: Replication and Extension. Public Opinion Quarterly, 82(4), 707–744.

4. Schouten, B., Cobben, F. & Bethlehem, J. (2009). Survey Methodology, 35(1), 101–113.

5. Groves, R.M. & Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias: A Meta-Analysis. Public Opinion Quarterly, 72(2), 167–189.

6. AAPOR (2016). Evaluating Survey Quality in Today's Complex Environment. aapor.org

7. AAPOR (2023). Data Quality Metrics for Online Samples: Task Force Report. aapor.org

8. Roper Center for Public Opinion Research (Cornell). Polling Fundamentals. ropercenter.cornell.edu

9. Roper Center. What is Roper iPoll?, and iPoll study record (Axios/Ipsos American Health Index, Wave 1, #31120137). ropercenter.cornell.edu

10. Quintana-Mathé, A., Uslu, A.A., Radford, J., Druckman, J.N., Lunz Trujillo, K., Safarpour, A., et al. (2026). Using Opt-In Non-Probability Surveys to Estimate COVID-19 Infection and Vaccination Rates. Journal of Survey Statistics and Methodology, smaf041.

11. Wiśniowski, A., Sakshaug, J.W., Perez Ruiz, D.A., & Blom, A.G. (2020). Integrating Probability and Nonprobability Samples for Survey Inference. Journal of Survey Statistics and Methodology, 8(1), 120–147.

12. Taylor, H. (1995). Horses for Courses: How Survey Firms in Different Countries Measure Public Opinion with Very Different Methods. Journal of the Market Research Society, 37(3), 1-9.

13. Rohr, Björn, Barbara Felderer, Henning Silber, Jessica Daikeler, Joss Roßmann, and Jette Schröder. 2024. When Are Non-Probability Surveys Fit for My Purpose? Mannheim: GESIS – Leibniz Institute for the Social Sciences (GESIS Survey Guidelines), version 1.0. https://doi.org/10.15465/gesis-sg_en_050.

14. Little, R.J. & Vartivarian, S. (2005). Does Weighting for Nonresponse Increase the Variance of Survey Means? Survey Methodology, 31(2), 161–168.

1. Plain-Language Definition

The mode of data collection is simply the method used to gather answers from respondents. The most common modes are telephone (landline or cell phone); in-person (face-to-face interviews), mail (paper questionnaires sent and returned by post), and online (web-based surveys completed on a computer, tablet, or phone). If a survey used a method that does not fit these standard categories (for example, text-message surveys or interactive voice response systems), that method should be described separately.1,2

Knowing the mode matters because each method has its own strengths and weaknesses that can affect who responds, how they respond, and the quality of the data. For much of the twentieth century, face-to-face interviewing was the dominant mode in public opinion polling specifically; mail played a larger role in survey research more broadly, such as academic and government surveys, than it did in political and media polling. Telephone surveys rose to prominence in the 1970s and 1980s because they were faster and cheaper than sending interviewers to people's homes. Since the 1990s, internet surveys have increasingly challenged telephone as the dominant mode, again largely because of speed and cost advantages.2

Today, many surveys use more than one mode. For example, sending an initial invitation by mail with a link to an online questionnaire, then following up by phone with people who did not respond. These are called mixed-mode or multi-mode surveys, and they are designed to combine the strengths of different methods while offsetting each one's weaknesses.1,2,6

2. Why It Matters for Interpreting Results

The mode of data collection can influence survey results in ways that go beyond logistics. Different modes produce different kinds of errors, and understanding the mode helps readers judge which errors are most likely to be present.2,3

First, mode affects who can be reached. An online-only survey cannot include people without internet access. A telephone survey using only landline numbers will miss the growing share of adults who rely exclusively on cell phones. A mail survey will miss people at addresses not on the postal delivery file. These gaps are called coverage errors. They mean some members of the target population had no chance of being included.2,3

Second, mode affects who actually participates. Response rates vary significantly by mode. In-person surveys have historically achieved the highest response rates, followed by telephone, then mail, with online surveys often having the lowest rates among general-population surveys. Lower response rates increase the risk that the people who chose to participate are different in meaningful ways from those who did not.2,3

Third, mode affects how people answer questions. Research shows that people are more likely to give socially desirable answers (saying what they think they should versus what they really think) when speaking to a live interviewer compared to filling out a questionnaire privately. Sensitive topics like drug use, racial attitudes, or sexual behavior tend to produce different results depending on whether an interviewer is present.2

For mixed-mode surveys, these effects become more complex. If some respondents answered by phone and others answered online, the analyst needs to consider whether any differences between those groups are due to the mode itself or to genuine differences in opinion. This is an active area of research, and there is no single solution that works for all situations.2,6

3. What Effective Disclosure Looks Like

Effective disclosure names every mode used and, for mixed-mode surveys, explains how respondents were assigned to each mode. A well-documented methodology section might say: 'Surveys were conducted via self-administered web survey or by live telephone interviewing. Online panelists received an email invitation and up to two reminders. Phone panelists received a prenotification postcard and could receive up to six calls from trained interviewers.' This level of detail lets readers judge how thorough the data collection process was.4,5

AAPOR's Transparency Initiative requires disclosure of the mode of data collection as a core element. For mixed-mode designs, good practice means reporting the sample size and response rate for each mode separately, not just an overall figure. This matters because collapsing different modes into a single response rate can hide important differences in data quality.1,2

For online surveys specifically, simply saying 'conducted online' is not enough. There are several ways to recruit people for online surveys (through probability-based panels, opt-in volunteer panels, social media ads, website pop-ups, and more), and each has very different implications for data quality. Effective disclosure describes the panel or recruitment method, not just the mode.2,3

4. What Incomplete Disclosure Looks Like

The most basic problem is not naming the mode at all. Without this information, readers cannot assess any of the potential biases discussed above. A description like 'a national survey of 1,000 adults' that omits how those adults were contacted is missing a critical piece of the picture.1

Another common issue is using vague labels. Saying a survey was 'conducted by telephone' without specifying whether landline, cell phone, or both were used leaves important questions unanswered, since landline-only surveys systematically underrepresent younger adults and certain minority groups. Similarly, 'conducted online' without details about the panel or recruitment method gives no basis for evaluating the sample's representativeness.2,3

For mixed-mode surveys, failing to report which respondents used which mode is a significant omission. If 80 percent of respondents answered online and 20 percent by phone, the survey's error profile is dominated by the properties of the online mode, not the phone mode. Without this breakdown, readers cannot evaluate the data properly.2

Omitting information about interviewer-assisted versus self-administered components within a single survey is also a concern. Some face-to-face surveys include sections where respondents enter answers privately on a laptop (for sensitive questions), and those self-administered sections may produce different response patterns than the interviewer-administered portions.2

5. Strengths When Well-Disclosed

When the mode is clearly described, readers can immediately identify the most likely sources of error and factor them into their interpretation. An experienced poll reader who sees ‘self-administered online survey’ knows to ask different questions about coverage and response style than one who sees ‘live telephone interviewing.’ 

Detailed mode disclosure also enables meaningful comparison between surveys. If two polls on the same topic produce different results, knowing that one was conducted by phone and the other online helps explain the discrepancy. The difference might be a mode effect rather than a genuine shift in opinion.2

For mixed-mode surveys, disclosing the breakdown by mode lets analysts test whether the mode itself is driving differences in responses, which supports more careful and honest reporting of results.2

6. Limitations to Keep in Mind

Mode labels can be deceptively simple. Two surveys both described as 'online' can be very different. One might use a probability-recruited panel while the other uses a convenience sample of volunteers. The mode label alone does not capture these crucial differences.2,3

Mode effects are not constant. As people become more comfortable with newer technologies and as survey designers learn to reduce mode differences through careful questionnaire design, the gap between modes may narrow for some types of questions. But for sensitive topics where interviewer presence matters, mode differences are likely to persist.2

Several long-running trend series have changed mode partway through their history. The General Social Survey has tracked American attitudes since 1972, mostly through in-person interviews. In 2022, NORC began collecting it using in-person, phone, and web modes together for the first time. NORC's director of the GSS said the organization is still researching how mode affects responses and trends in GSS estimates over time.9

Gallup faced the same question for its Gallup Poll Social Series, which includes questions dating to the 1930s. Starting in 2021, Gallup tested phone versus web administration side by side. The test showed that switching to web could change results on many of Gallup's long-tracked questions. Gallup chose to keep the series on the phone to avoid breaking those trends.7

Pew Research Center saw a concrete example of this problem. From 2009 to 2014, Pew's phone polls found that about two-thirds of Americans said there was a lot of discrimination against gay and lesbian people. In 2014, Pew asked the same question online and got 48 percent. A side-by-side experiment showed this drop was caused by the mode change, not a real shift in opinion. The phone results from that same experiment matched the earlier phone trend closely.8

Mixed-mode designs are intended to reduce overall error, but they can also introduce new complications. If different modes produce systematically different answers to the same question, combining them requires statistical adjustments that may not fully resolve the problem. 2

One source of these differences is visual display. Mail surveys can share the questionnaire itself as a file in their reporting, letting readers see exactly what respondents saw. Online surveys rarely do the same. AAPOR's disclosure standards list providing screenshots for self-administered computer-assisted surveys as "strongly encouraged, though not required," and even where encouraged, this level of detail must only be made available on request rather than published upfront. In practice, screenshots remain uncommon due to the complexities of capturing every variation in display for a complex survey with branching questions accessed via desktop, tablet, or smartphone. Without that information, researchers combining data across modes have no way to check whether spacing or layout contributed to the differences they are trying to statistically adjust for. The field is still developing best practices for handling these challenges.7

7. Types of Modes You May Encounter

Interviewer-administered modes

Telephone surveys use trained interviewers who call respondents on landline phones, cell phones, or both. They can include computer-assisted telephone interviewing (CATI), where the interviewer reads questions from a screen and enters responses directly. In-person or face-to-face surveys involve interviewers visiting respondents in their homes or at designated locations. These often include computer-assisted personal interviewing (CAPI) and may incorporate self-administered sections for sensitive questions.1,2

Self-administered modes

Mail surveys send paper questionnaires that respondents fill out and return. Online or web surveys are completed on a computer, tablet, or smartphone. Interactive voice response (IVR) systems, sometimes called 'robo-polls,' use automated telephone systems that play recorded questions and collect responses through keypad entries, with no live interviewer involved.1,2 A live interviewer also usually lets respondents volunteer "don't know" even when it isn't offered as an option. Self-administered surveys instead either make "don't know" an explicit choice, which can increase its use, or require an answer to continue, making a genuine "don't know" impossible to record.

Mixed-mode designs

Many modern surveys combine two or more modes. Common combinations include mail invitation with a web response option, or web as the primary mode with phone follow-up for nonrespondents. Multi-mode designs can be sequential (offering modes one after another to encourage response) or concurrent (giving respondents a choice of modes from the start).1,2

8. Related Elements

Mode connects directly to data collection dates, because different modes require different amounts of time. Mail surveys inherently take longer than online surveys because of postal delivery times. Phone surveys can be faster than mail but slower than online, depending on how many callbacks are needed. Knowing the mode helps you judge whether the field period was typical.1,2

Mode also relates to sample design and response rates. The choice of mode determines which sampling frames are available (for example, only address-based surveys can use the postal delivery file) and strongly influences how many people respond. These connections mean that mode disclosure is most useful when read alongside the other methodological details.2,3

References

1. AAPOR Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th ed. aapor.org

2. Couper, M. P. (2011). The Future of Modes of Data Collection. Public Opinion Quarterly, 75(5), 889-908.

3. Groves, R. M. & Lyberg, L. (2010). Total Survey Error. Public Opinion Quarterly, 74(5), 849-879.

4. Pew Research Center (2026). Methodology: U.S. Role in the World survey.

5. AAPOR Transparency Initiative Disclosure Elements (revised April 2021). aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

6. Enns, P.K., Barry, C.L., Druckman, J.N., Garcia-Rios, S., Wilson, D.C., & Schuldt, J.P. (2024). The Need for a Recurring Large-Scale Benchmarking Survey to Continually Evaluate Sampling Methods and Administration Modes: Lessons from the 2022 Collaborative Midterm Survey. arXiv preprint arXiv:2407.06090.

7. Gallup. (2024, January 25). Gallup poll methodology. Gallup News. https://news.gallup.com/opinion/methodology/608690/gallup-poll-methodology.aspx

8. Kennedy, C., & Deane, C. (2019, February 27). What our transition to online polling means for decades of phone survey trends. Pew Research Center. https://www.pewresearch.org/short-reads/2019/02/27/what-our-transition-to-online-polling-means-for-decades-of-phone-survey-trends/

9. NORC at the University of Chicago. (2023, May 17). 2022 General Social Survey data released to public [Press release]. https://www.norc.org/research/library/2022-general-social-survey-data-released-to-public.html

10. American Association for Public Opinion Research. (2022, December 1). Disclosure standards. AAPOR. https://aapor.org/standards-and-ethics/disclosure-standards/

1. Plain-Language Definition*

The sample size is the total number of people who completed the survey, counted before any statistical adjustments (weighting) are applied. This unweighted count is the standard way sample size is reported, and it is the starting point for judging how accurate the survey's results are likely to be.1

A note on terminology: while the vast majority of public polling organizations use “sample size” to refer to the number of people who actually completed the survey, some academic and government surveys use “sample size” to mean the number of people originally targeted for contact, and use “completes” to indicate the number who actually took the survey. Throughout this piece, “sample size” refers to the total number of completed interviews.

Think of it this way: because a survey talks to a sample of people rather than every single person in the population, survey results are unlikely to exactly match what you would get if you could talk to everyone. In other words, surveys always come with a margin of error. A larger sample generally means smaller margins of error and more precise estimates, while a smaller sample means wider margins of error and less certainty.1,2

While bigger is generally better, there are diminishing returns. Increasing the sample from 100 to 1,000 dramatically reduces the margin of error, but going from 1,000 to 2,000 only shrinks it by about one percentage point. This is why many surveys have sample sizes of around 1,000 to 1,200 respondents. Beyond a certain size, each additional respondent buys you very little added precision.1

2. Why It Matters for Interpreting Results

Sample size matters because it directly relates to the margin of sampling error, which is the plus-or-minus figure reported alongside most poll results. A standard national survey of about 1,000 adults typically has a margin of error of around +/– 3 percentage points. That means if the survey finds 58 percent approve of some policy proposal, the true figure is likely to fall somewhere between 55 and 61 percent.1,2

However, it is important to remember that the overall sampling margin of error applies only to the full sample. When results are broken down by subgroups (such as by race, age, or political party), each subgroup has a smaller number of respondents and therefore a larger margin of error. For example, in a survey of 1,000 adults, the Hispanic subsample might include only about 160 people. Because of this, results for Hispanics come with a larger margin of error of around +/– 7 percentage points. Simply put, estimates for small subgroups can be quite unreliable even within a large overall survey.2

Raw sample size can also be misleading. A survey with 10,000 respondents sounds impressive, but if those respondents were all volunteers who opted into the survey by choosing to click on an online link, the sample may be less representative of the general population than a carefully designed probability sample of 1,000 people in which everyone had an equal chance of being invited to participate.3

Statistical weighting, which adjusts the data to match known population characteristics, can further complicate the picture. Weighting is essential for avoiding biased results, but it has the side effect of increasing the margin of error. The effective sample size (the sample size after accounting for the extra variability introduced by weighting) is always smaller than the raw count. One study found that applying statistical weights to a 2,000-person sample increased the average margin of error from +/– 1.3 points (unweighted) to +/– 1.9 points.4

3. What Effective Disclosure Looks Like

AAPOR's Transparency Initiative requires that publicly released survey results report sample sizes by sampling frame, and for probability surveys, include estimates of sampling error along with a note about whether the reported margin of error has been adjusted for design effects from weighting or clustering.5

A model disclosure reports the total number of completed interviews, distinguishes between different modes if more than one was used (for example, 2,217 online plus 3,516 paper), provides the margin of error (typically at the 95 percent confidence level, by convention), notes the design effect, and makes subgroup sample sizes available. It also describes the weighting procedure and whether weights were trimmed to prevent any single respondent from having an outsized influence on the results.6

For surveys that oversample certain groups (intentionally interviewing more people from a particular subgroup to improve precision for that group), effective disclosure notes the oversampled group, its sample size, and any limitations. For example, a survey might note that it includes 325 Asian adults in the U.S., and also that these respondents are primarily English-speaking and may not fully represent the broader Asian adult population.6

Because of the heightened error that accompanies estimates for smaller subgroups, some organizations have adopted practices to help readers avoid unjustified conclusions from the data. For example, Pew Research Center displays margins of error directly in charts to make sure readers understand the limits of precision for smaller groups.7  The Roper Center’s Roper iPoll demographic crosstab feature does not even display results for subgroups of fewer than 100 respondents, given the very large error intervals that accompany any such results.

4. What Incomplete Disclosure Looks Like

The most basic problem is not reporting the sample size at all or reporting it without clarifying whether it is weighted or unweighted. Without the unweighted count, readers cannot calculate their own margins of error or judge the precision of the estimates.5

Another common issue is reporting the overall sample size but not providing subgroup counts. If a report breaks results down by race, age, gender, or other categories, readers need to know how many respondents are in each group to judge whether those breakdowns are reliable.5,7

For non-probability surveys (such as opt-in online polls), a particular concern is presenting a large sample size as evidence of reliability without acknowledging that traditional margins of error do not apply in the same way to non-probability surveys. AAPOR's guidance states that non-probability surveys should only provide precision measures if they are clearly defined and accompanied by a description of how the underlying model was specified and validated.3,5

Omitting information about design effects is also a gap. If a survey reports a margin of error that has not been adjusted for weighting, it is overstating its precision, claiming more accuracy than can be justified by the data.4,5

5. Strengths When Well-Disclosed

A clearly reported sample size lets any reader gauge the statistical precision of the survey's findings. It is one of the most straightforward pieces of information to evaluate. You do not need advanced statistical training to understand that a survey of 2,000 people is likely to offer more precision than one of 200.1,2

When oversampling is disclosed, readers can understand why certain subgroups have more respondents than their share of the population would suggest, and they can see that the survey organization took extra steps to ensure reliable estimates for those groups.6

Additional reporting of the effective sample size (after accounting for design effects) gives readers the most honest picture of precision, rather than the potentially misleading raw count. This is the kind of responsible reporting that AAPOR's Transparency Initiative members are required to practice.4,5

6. Limitations to Keep in Mind

Sample size is only one factor in survey quality. A well-designed survey of 1,000 people using probability sampling can produce more accurate results than a poorly designed survey of 50,000 volunteers. Focusing too much on sample size alone can lead readers to overvalue large but unrepresentative samples.3

The raw sample size does not account for the effects of weighting. After weighting, the effective sample size may be substantially smaller, meaning the survey results are less precise than the raw count implies. Readers should look for the design-effect-adjusted margin of error rather than calculating precision from the raw count alone.4

Sample size also does not address problems of nonresponse or coverage. Even a very large sample has blind spots if certain types of people systematically did not respond or were not reachable by the survey method used. These errors cannot be measured by looking at the sample size.2,3

Survey researchers have introduced the concept of total survey error to account for these kinds of additional error sources that go beyond error associated with sample size.

7. Variations You May Encounter

Total N is the overall number of respondents who completed the survey. The number might be given in a general methodological statement, such as “interviews were conducted with 1,200 adults”, or sometimes with the notation “n=1,200”, Total N is the most commonly reported figure and is the basis for the sampling margin of error.5

Unweighted N versus weighted N: the unweighted count is the actual number of people interviewed, while the weighted count adjusts those numbers to represent the population. These can differ substantially, especially if certain groups were oversampled or if the weighting procedure was aggressive.4,5

Effective N is the sample size adjusted for the extra variability introduced by weighting. It is always smaller than the unweighted count, and it gives the most honest picture of a survey's precision.4

Subgroup N refers to the number of respondents within a particular demographic or other breakdown. Subgroup Ns are critical for evaluating whether reported differences between groups are statistically meaningful or might simply reflect small sample sizes.2,7

Oversample N is the extra number of respondents added from a particular subgroup to improve precision for that group. When oversampling is used, the extra respondents are weighted back to their correct proportion of the population for overall estimates.6

8. Related Elements

Sample size is closely linked to weighting. The weights applied to each respondent determine the effective sample size, and the variables used for weighting (demographics, political attitudes, and so on) influence how much the raw count differs from the effective count.4,5

Sampling procedure is also closely connected. The design of the sample (whether probability-based or non-probability, stratified or clustered) determines whether traditional margin-of-error calculations are valid in the first place. A large non-probability sample does not have the same statistical properties as a smaller probability sample.3,5

Finally, sample size relates to the sampling margin of error that is reported alongside the results. The sampling margin of error should incorporate the design effect from weighting; if it does not, the reported precision is overstated.4,5

9. Further Reading

AAPOR Education Resources: Margin of Sampling Error / Credibility Interval. aapor.org/Education-Resources/Election-Polling-Resources/Margin-of-Sampling-Error-Credibility-Interval.aspx

Pew Research Center (2016). Understanding the Margin of Error in Election Polls. pewresearch.org/short-reads/2016/09/08/understanding-the-margin-of-error-in-election-polls/

Baker, R., et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143.

Pew Research Center (2018). Variability of Survey Estimates. pewresearch.org/methods/2018/01/26/variability-of-survey-estimates/

References

1. AAPOR Education Resources: Margin of Sampling Error / Credibility Interval. aapor.org

2. Pew Research Center (2016). Understanding the Margin of Error in Election Polls. pewresearch.org

3. Baker, R., et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143.

4. Pew Research Center (2018). Variability of Survey Estimates. pewresearch.org/methods/2018/01/26/variability-of-survey-estimates/

5. AAPOR Transparency Initiative Disclosure Elements (revised April 2021), Item 8: Sample Sizes. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

6. Pew Research Center (2024). 2023 NPORS Methodology. pewresearch.org/internet/2024/01/31/social-media-2023-national-public-opinion-reference-survey-npors-methodology/

7. Pew Research Center (2021). Why Pew Research Center Will Display Margins of Error in Some Graphics. pewresearch.org/decoded/2021/10/25/

1. Plain-Language Definition

Sampling procedure has two connected parts. The first covers the overall method used to select participants: how the researcher decided who would be contacted or given a chance to respond. The second specifically covers the last step in that process, when there are multiple stages, in which an individual respondent is chosen from within a household, panel, or other group. These are treated together because they answer closely related questions and often appear together in methodology reports.1

The most important distinction in Element 16 is between probability sampling, where every eligible person has a known and non-zero chance of being selected, and non-probability sampling, where selection depends on availability, self-selection, or the researcher's judgment. Address-based sampling from the USPS Delivery Sequence File, random-digit dialing of telephone numbers, and online probability panels initially recruited using one of these methods are common probability approaches. Opt-in online panels, quota samples, river samples (recruiting from website visitors), and convenience samples are common non-probability approaches.2

Respondent selection covers who exactly gets interviewed within the sampled unit. Methods range from formal probability techniques (the Kish grid) to quasi-probability shortcuts (asking for the adult with the next or last birthday), to non-random choices (whoever answers the phone or clicks the link).3

2. Why It Matters for Interpreting Results

Probability sampling provides an established statistical basis for generalizing from the sample to a population and for calculating a margin of error, Non-probability sampling does not permit the same margin-of-error calculation, since there is no defined probability of selection to draw on. Reports of non-probability samples that quote a classic margin of error are technically overstating what the method delivers.2 

However, sampling is complex, and no approach can guarantee perfectly accurate results. Proponents of both probability and non-probability methods have argued that their designs can support representativeness. It is also worth remembering that margin of error, when it can be calculated, addresses only sampling error and is one component of Total Survey Error, a framework that also accounts for coverage, nonresponse, measurement, and processing error. 11

The AAPOR Task Force on Non-Probability Sampling made the necessity of disclosure explicit: when non-probability sampling methods are used, there is a higher burden on the researcher to describe how the sample was drawn, how the data were collected, and how inferences were made. Non-probability results depend on modeling assumptions, and users need enough procedural detail to evaluate those assumptions.2 This is a matter of degree rather than kind: probability sampling methods vary too, but not nearly as much as non-probability methods do, and their standard approaches are widely known and well documented. Because non-probability approaches vary far more from one implementation to the next, they place a comparatively higher burden on researchers to spell out exactly what was done.

The respondent selection stage matters for a different reason: it determines whose voice gets heard within each sampled unit. In household telephone surveys, if a survey simply interviews whoever answers the phone, women and older adults tend to be overrepresented because they are more likely to be at home and to pick up. Systematic within-household selection reduces this bias.3,4 In online panel surveys the 'respondent selection stage' is often just 'whoever clicks the link first,' which is worth noting explicitly rather than treating as if the survey had no such stage.

3. What Effective Disclosure Looks Like

Effective disclosure states plainly whether the sample is probability, non-probability, or a combination. It names the frame or list (e.g., 'USPS Delivery Sequence File,' 'the vendor's opt-in panel,' 'a purchased list of registered voters in Ohio'), identifies any panel provider or sample vendor involved, and describes how participants were contacted, recruited, or intercepted. For multi-stage designs, each stage is documented, including the final within-unit selection method. Any use of quotas, incentives, or refusal-conversion procedures is stated.1

The Pew Research Center's American Trends Panel methodology is a widely cited model. Pew documents the sampling frame (USPS Delivery Service File), the recruitment mechanism (address-based sampling by mail), the within-household selection rule (the adult with the next birthday), the panel management procedures, and the weighting approach. A reader can walk through the full sampling process from target population to individual respondent.5

4. What Incomplete Disclosure Looks Like

Common gaps include describing a sample as 'an online panel' without saying whether the panel is probability-based or opt-in; reporting a margin of error for a non-probability sample without noting the model-based caveat; naming a vendor without describing how that vendor recruited participants; and omitting the within-household or within-panel selection method altogether.1

Vague phrasings such as 'a national research firm surveyed 1,000 adults' leave users unable to determine whether the design supports the reported claims. Reports that omit the sampling frame, the mode of recruitment, or the within-unit selection stage fall well below the AAPOR Transparency Initiative standard for immediate disclosure.1,6

For panel-based studies, incomplete disclosure often does not specify how panel members were originally recruited, how attrition is managed, or whether the reported respondents were a random subset of the full panel. 'Pre-recruited panelist' is not sufficient by itself.7

5. Strengths When Well-Disclosed

Clear disclosure of the sampling procedure lets users match the design to the claim. A probability sample with documented within-household selection can support inferences about the target population, complete with sampling error estimates.5,8

Full procedural disclosure also allows comparison across surveys on the same topic. Two polls showing different numbers on the same question can often be explained by differences in sampling procedure (e.g., probability panel versus river sample) or in respondent selection (e.g., next-birthday versus no selection).3

Detailed disclosure also supports methodological improvement across the profession. When organizations describe what they did in enough detail for others to replicate or critique it, weaknesses can be identified and corrected in future studies.1

6. Limitations to Keep in Mind

Disclosure alone does not fix problems in the underlying design. All samples now grapple with declining response rates, and even the strongest sampling frame cannot include people who refuse to participate. Weighting can adjust for some of these gaps but does not eliminate them.5,7

The AAPOR Task Force reached its assessment of non-probability accuracy in 2013.2 More recent research complicates a blanket preference for probability sampling: Pew Research Center's 2023 benchmarking study found that while opt-in samples had roughly twice the average error of probability-based panels overall, all three probability-based panels consistently overestimated 2020 voter turnout by 8 to 9 percentage points, whereas the opt-in samples came within 1 to 3 points of the actual benchmark.¹² This suggests neither method is uniformly reliable across all measures, and a single accurate estimate on one variable should not be taken as evidence that either approach is broadly superior.

Within-household selection methods carry their own trade-offs. The Kish grid is close to true probability selection but is intrusive and can increase refusal rates. Birthday methods sacrifice some randomness for shorter conversations, and they depend on the informant knowing everyone's birthday, which can be unreliable in larger households.3,9

Historical disclosure

It is also worth noting that the most commonly used sampling methods have varied substantially across countries and across the history of polling, as demonstrated in a 1995 survey of polling firms that described the most commonly used methodologies in a number of countries.13 Therefore, the assumptions that might be made about U.S. polling methodologies for a particular decade cannot be extrapolated to polling from the same time frame in the UK or Brazil.

Within the United States, non-probability, quota-based sampling was almost exclusively used before the early 1950s, at which point Gallup moved to in-person probability sampling. Over the next twenty years, other organizations followed, with the arrival of Random Digit Dialing, an affordable probability-based telephone polling method, becoming the dominant sampling method by the late 1970s. Inexpensive non-probability online panels appeared in the late 1990s and have become more common over the decades since. Online panels utilizing probability methods for recruitment were first attempted in the early 2000s and have been increasingly adopted as telephone response rates have dropped.

Historical disclosure on an individual survey basis has substantial gaps, as professional transparency standards were not in place for the earliest years of polling. While quota sampling can be assumed for virtually all U.S. polls in the thirties and forties, the slow adoption of probability methods and the common use of mixed probability and non-probability methods for the 1950s through the 1970s makes research on the typical techniques of particular survey firms imperative for historical researchers.

7. Types of Sampling Procedures You May Encounter

Probability-based sampling

Selects participants with known, non-zero probabilities from a defined frame. Common variants include random-digit dialing (RDD), address-based sampling (ABS), and probability-recruited online panels.2,5

Non-probability sampling

Includes opt-in online panels (large pools of self-selected volunteers), river sampling (recruiting directly from website visitors), quota sampling (filling category targets without random selection within category), snowball or respondent-driven sampling (participants recruit others), and convenience samples. These methods are increasingly common because of cost and speed, but their statistical properties depend heavily on modeling assumptions.2

Mixed or hybrid designs

Some studies combine a probability sample with a non-probability boost sample (often to reach small subgroups), or combine multiple frames such as landline and cell phone. When mixed, each component should be described separately.7,10

Within-household selection methods

The Kish grid (formal probability selection using age and gender), full enumeration (list all adults, pick randomly), last-birthday and next-birthday methods (quasi-probability), the Troldahl-Carter family and its variants (asking for the oldest or youngest man or woman), Youngest Male / Oldest Female (YMOF), and no selection (whoever answers). Each has documented trade-offs among statistical rigor, respondent burden, and cost.3,9

Online panel selection

For online panel surveys, the 'final stage' is often self-selection into an active survey wave from within the panel. While important to consider, whether panel members were originally recruited through probability methods matters more than the mechanics of that final stage.2,7

8. Related Elements

Sampling Procedure is most closely paired with Sampling Frame, which describes the universe from which the sample was drawn. It also connects to Data Collection Mode, since some sampling procedures are workable only in certain modes, and to Sample Size and Margin of Error, which depend on whether the method is probability-based. Weighting exists specifically to compensate for known limitations in the sampling procedure.1

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 5. aapor.org/standards-and-ethics/disclosure-standards/

2. Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology 1(2), 90-143.

3. Gaziano, C. (2005). Comparative Analysis of Within-Household Respondent Selection Techniques. Public Opinion Quarterly 69(1), 124-157.

4. Denk, C.E., Guterbock, T.M., & Gold, D. (1996). Within-household respondent selection: effects on handoffs and cooperation. Paper presented at AAPOR Annual Conference.

5. Pew Research Center. About the American Trends Panel. pewresearch.org/the-american-trends-panel/

6. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

7. AAPOR Task Force on Transitions from Telephone Surveys to Self-Administered and Mixed-Mode Surveys. aapor.org/standards-and-ethics/reports/

8. Pew Research Center (2019). Growing and Improving Pew Research Center's American Trends Panel. pewresearch.org/methods/

9. Kish, L. (1949). A Procedure for Objective Respondent Selection within the Household. Journal of the American Statistical Association 44(247), 380-387.

10. AAPOR Task Force on Address-Based Sampling (2016). aapor.org/standards-and-ethics/reports/

11. Groves, R.M. & Lyberg, L. (2010). Total Survey Error: Past, Present, and Future. Public Opinion Quarterly, 74(5), 849–879.

12. Mercer, A. & Lau, A. (2023). Comparing Two Types of Online Survey Samples: Opt-In Samples Are About Half as Accurate as Probability-Based Panels. Pew Research Center.

13. Taylor, H. (1995). Horses for courses: How different countries measure public opinion in very different ways. The Public Perspective, 6(2), 3–7. https://ropercenter.cornell.edu/sites/default/files/2018-07/62003.pdf

1. Plain-Language Definition

The sampling frame is the list, database, or method used to identify and reach potential respondents. It is the practical operationalization of the target population. If the target population is 'adults living in the United States,' the sampling frame is the specific tool the researcher used to try to reach them: for example, a national list of residential addresses, a list of cellphone and landline numbers, or a pre-recruited online panel.1

The frame is not the same as the population. The population is who the researcher wants to represent. The frame is what the researcher actually had access to. Every frame excludes someone. Address-based frames miss people without stable addresses. Landline frames miss cell-only households. Online panel frames miss people who are not on that panel. The size of the gap between the frame and the population, and whether the excluded people differ systematically from those included, is the central question this element helps a reader assess.1,2

2. Why It Matters for Interpreting Results

The sampling frame determines who has any chance at all of being in the sample. Anyone missing from the frame has zero chance of selection, regardless of how good the rest of the design is. If people missing from the frame differ from people on it in ways related to the topic being measured, the results will be biased, and no amount of weighting or oversampling can fully repair the damage. This is called coverage error.2

Coverage error can change over time for a particular sampling frame as the result of technological or social shifts. The classic example is the erosion of the landline telephone frame. Landline coverage of U.S. households was high through the 1990s but has fallen dramatically as households have moved to cell-only service. As of the late 2010s, more than half of U.S. adults lived in cell-only households.3 A landline-only frame today would systematically miss younger, lower-income, and more mobile households, biasing any survey that used it. Address-based sampling (ABS) from the USPS Delivery Sequence File emerged in part to address this coverage gap.4

For non-probability designs, the same logic applies to the effective frame. An opt-in online panel is, in effect, a list of people who found their way onto the panel. Those people are not a random slice of the population; they tend to be more digitally engaged, and often differ from non-panelists on attitudes and behaviors that are relevant to survey topics. Users need to know the frame in order to reason about who was reachable and who was not.5

3. What Effective Disclosure Looks Like

Effective disclosure names the specific frame used, describes how it was constructed or licensed, and notes any known coverage gaps relative to the target population. For an address-based sample, it would identify the vendor (or that the USPS Delivery Sequence File is the base), and disclose whether unusual address types (P.O. boxes, drop points, seasonal addresses) were included or excluded. For a telephone sample, it would specify landline, cellphone, or dual-frame, and identify the sample vendor.1,4

For an online panel, effective disclosure identifies the panel by name, states whether the panel is probability-based or opt-in, and explains how members were originally recruited to the panel. If the researcher used a sample router or multiple panels blended together, that is disclosed too.5

The AAPOR Transparency Initiative requires that a probability-based sample specification include the name of the supplier of the sample or list and the nature of the list (e.g., 'ABS,' 'registered voters in Texas in 2018,' 'pre-recruited panel'), plus the coverage of the target population, including any segments not covered by the design.1,6

4. What Incomplete Disclosure Looks Like

Common gaps include using generic terms like 'online panel' or 'nationally representative sample' without naming the frame or the vendor; describing a telephone sample without specifying landline, cellphone, or both; and reporting coverage in vague terms ('covers most U.S. adults') without stating what is excluded.1

Reports that omit the sampling frame entirely, or that leave a reader unable to distinguish between a probability-based address frame and an opt-in panel, fall well below the AAPOR Transparency Initiative standard for immediate disclosure. When two different frames were combined (such as address-based mail plus web recruitment) and only one is disclosed, that is also incomplete.6

For non-probability designs, 'the sample was drawn from an online panel' without more information is not enough. The reader needs to know how that panel was assembled to have any basis for evaluating who is on it.5

5. Strengths When Well-Disclosed

A clearly named frame lets users evaluate coverage against the stated target population. For probability-based frames such as ABS with the USPS Delivery Sequence File, coverage of the U.S. residential population is very high, and users can be reasonably confident that most households had a chance of selection.4,7

Full frame disclosure enables comparison across surveys. A study using ABS and a study using an opt-in panel are not directly comparable even if they report the same sample size on the same question, and knowing the frame tells the reader that.2,5

Frame disclosure also enables cumulative learning. When researchers document which frames they used and what coverage they achieved, the profession as a whole can track how frames are performing over time and when new approaches are needed.1,3

6. Limitations to Keep in Mind

No frame is perfect. Every frame has some undercoverage and often some overcoverage as well (e.g., duplicate addresses, defunct phone numbers, or panel members who no longer participate). The best-documented frame still misses somebody.2,4

Frame coverage also changes over time. Landline frames were high-coverage in the 1990s and are low-coverage today. ABS frames continue to improve as USPS records improve. Online penetration grows, but not everyone is online, and the composition of the online population continues to shift. A frame that was appropriate for a study five years ago may not be appropriate now.3

Coverage bias is hardest to detect when frame gaps correlate with the outcome being measured. If the excluded people are similar to the included people on the topic of interest, there may be little bias even with substantial undercoverage. If they differ, the bias can be large. Users cannot usually determine which situation they are in without additional data.2

Frame disclosure is most useful when considered alongside other methodological elements. Frame quality interacts with mode of contact, recruitment procedures, and weighting, and cannot be fully assessed in isolation.1

7. Types of Sampling Frames You May Encounter

Address-based sampling (ABS)

Uses lists of residential mailing addresses derived from the USPS Computerized Delivery Sequence file. Coverage of U.S. households is very high and continues to improve. Now widely considered the strongest general-purpose frame for U.S. household surveys.4,7

Random-digit dialing (RDD)

Generates telephone numbers within known residential exchanges. Historically the dominant frame for U.S. surveys. Landline-only RDD now covers a shrinking share of households, so most current RDD designs use dual-frame (landline plus cellphone) sampling. Response rates have declined sharply over the past two decades.3

Probability-based online panels

Panels whose members were recruited through probability sampling (typically ABS or RDD), invited to join, and provided with internet access when needed. Examples include Pew's American Trends Panel and NORC's AmeriSpeak. Support the same kinds of statistical inferences as fresh probability samples if managed rigorously.8

Opt-in (non-probability) online panels

Pools of volunteers who signed up to take surveys, often recruited through online advertising or referrals. Fast and inexpensive, but coverage is inherently self-selected. Representativeness depends on modeling assumptions rather than random selection.5

Registered voter files

Official lists of registered voters, often used for election polling. Coverage of the general adult population is partial (unregistered adults are excluded, while new voters and the recently moved may not be captured depending on the frequency with which the updated list is made available), but coverage of the target population of interest (voters or likely voters) can be strong.1

Customer or membership lists

Frames provided by the study sponsor (e.g., employees of a company, subscribers to a service, members of an association). Coverage of the specific target population can be complete if the list is well-maintained, but generalization beyond the list is limited.

Social media user bases

Frames constructed from users of a particular platform, such as Facebook or LinkedIn. Coverage is limited to platform users, who differ demographically from the general population, and researchers usually cannot verify the frame independently.5

No formal frame (river sampling and routers)

Some designs recruit respondents directly from website visitors or through sample routers that direct visitors into whichever survey needs them. There is no pre-defined list; the effective frame is 'people who happened to be online and encountered the invitation.' Coverage properties are difficult to characterize.5

8. Related Elements

Sampling Frame is most closely paired with Population Under Study (Element 4), which defines who the researcher intends to represent. It also connects directly to Sampling Procedure , since the frame determines what sampling procedures are possible; to Data Collection Mode, since some frames work only with certain modes (email addresses for web, phone numbers for CATI); and to Weighting, which often adjusts specifically for known frame undercoverage.1

9. Further Reading

AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 5. aapor.org/standards-and-ethics/disclosure-standards/

AAPOR Task Force on Address-Based Sampling (2016). Address-Based Sampling. aapor.org/standards-and-ethics/reports/

Groves, R.M. & Lyberg, L. (2010). Total Survey Error: Past, Present, and Future. Public Opinion Quarterly 74(5), 849-879.

AAPOR Task Force on Transitions from Telephone Surveys. aapor.org/standards-and-ethics/reports/

Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology 1(2), 90-143.

AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 5. aapor.org/standards-and-ethics/disclosure-standards/

2. Groves, R.M. & Lyberg, L. (2010). Total Survey Error: Past, Present, and Future. Public Opinion Quarterly 74(5), 849-879.

3. AAPOR Task Force on Transitions from Telephone Surveys to Self-Administered and Mixed-Mode Surveys. aapor.org/standards-and-ethics/reports/

4. AAPOR Task Force on Address-Based Sampling (2016). aapor.org/standards-and-ethics/reports/

5. Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology 1(2), 90-143.

6. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

7. Shook-Sa, B.E., Currivan, D.B., McMichael, J.P., & Iannacchione, V.G. (2013). Extending the Coverage of Address-Based Sampling Frames. Public Opinion Quarterly 77(4), 994-1005.

8. Pew Research Center. About the American Trends Panel. pewresearch.org/the-american-trends-panel/

9. Iannacchione, V.G. (2011). The Changing Role of Address-Based Sampling in Survey Research. Public Opinion Quarterly 75(3), 556-575.

10. Blumberg, S.J. & Luke, J.V. (2018). Wireless Substitution: Early Release of Estimates from the National Health Interview Survey. National Center for Health Statistics.

1. Plain-Language Definition

Survey weighting is a set of statistical adjustments applied to raw survey responses so that the final results better represent the population the survey is trying to describe. Three closely related disclosure elements capture different parts of this process.1

The Weight Variable is the specific column in the dataset that analysts must apply when running any calculation. Without it, every respondent counts equally even if the sample over- or under-represents certain groups. The Weighting Benchmark Source is the authoritative external dataset that supplies the population proportions against which the sample is calibrated. The most common sources for U.S. surveys are the U.S. Census Bureau's American Community Survey and Current Population Survey. The Variables Used for Weighting are the demographic or other dimensions (such as age, gender, race, education, past vote, or region) on which the sample is adjusted to match those benchmarks.1,2

Knowing all three parts matters because they form an inseparable cluster. A weight variable has no meaning without knowing what benchmark it targets and which dimensions it adjusts on. AAPOR's Transparency Initiative treats weighting as a single disclosure requirement, asking researchers to describe how the weights were calculated, including the variables used and the sources of the weighting parameters.1,3

2. Why It Matters for Interpreting Results

Without weighting, a survey sample almost never mirrors the population exactly. If, for example, college graduates are over-represented in the raw data, unweighted estimates of any attitude correlated with education will be skewed. Proper weighting corrects these imbalances, but only along the dimensions included.4

Knowing which variables were weighted, what benchmark supplied the targets, and which dataset column to apply lets a secondary analyst verify the correction, reproduce the published results, and judge whether important dimensions were left unweighted.1

Users who do not apply weights to weighted data will get incorrect estimates. Disclosed weighting variables allow users to check whether key dimensions for their analysis were balanced. A named benchmark source is verifiable, meaning users can check the American Community Survey or Current Population Survey values used against publicly available population tables.1,2

Research on non-probability sampling has shown that calibration methods such as post-stratification and raking can reduce bias but rarely eliminate it entirely. Transparency about the adjustment process is essential for users to assess how much confidence the weighting warrants.5

3. What Effective Disclosure Looks Like

Effective disclosure names each weight variable, states the benchmark source with its vintage year, and lists every demographic or attitudinal variable used in the weighting procedure. It also specifies the weighting method (for example, raking or post-stratification) and, ideally, whether weight trimming was applied and at what threshold.1

Well-documented records provide enough detail for a secondary analyst to understand and reproduce the weighting procedure. For example, a methodology statement might read: "Poststratification variables included age, gender, census division, race/ethnicity, education, and 2024 presidential vote. Weighting variables were obtained from the 2025 Current Population Survey and the final results for 2024 presidential vote turnout and vote choice."6

For panel surveys with multi-stage weight construction, best practice distinguishes between design weights (which correct for sampling probabilities), nonresponse adjustments, and calibration weights (which rake to population benchmarks). Each stage should be documented so users understand how the final weight was built.7,8

If multiple weight variables exist in the dataset (for example, one for the total population and another for registered voters), effective disclosure explains when to use each one.1

4. What Incomplete Disclosure Looks Like

Incomplete disclosure typically says something like "data were weighted to reflect the general population" without naming which variables, which benchmark, or which weight column to use.1

Common gaps include omitting the benchmark vintage year, bundling everything into a single vague sentence, failing to distinguish between design weights and post-stratification weights when multiple weight variables exist, and not disclosing whether weights were trimmed.1,4

Without the weight variable name, users may apply the wrong weight or no weight at all. Multiple weight variables without documentation about which to use for which analysis cause confusion. Some providers use proprietary benchmarks with no public documentation, making independent verification impossible.1

Lack of weighting information is itself informative, because it means either the data are unweighted (which should be stated explicitly) or the organization chose not to disclose its procedures.3

5. Strengths When Well-Disclosed

When these three elements are well-disclosed, users can independently verify that the weighting is sound, reproduce the published findings, and assess whether the variables that matter most for their own analysis were included in the calibration.1

A named weight variable allows users to locate and apply the correct weight in the dataset. Multiple weight variables for different analytical purposes increase flexibility. Distinguishing between design weights and post-stratification weights is methodologically important because each corrects for a different source of error.1,4

Well-established benchmark sources such as the American Community Survey and Current Population Survey are updated regularly and have known margins of error. Naming the source allows users to assess whether the benchmark is appropriate for the target population. Listing all weighting variables allows users to evaluate whether the calibration covered the dimensions most relevant to their analysis.2,9

Transparent weighting documentation also signals organizational commitment to methodological rigor and supports cumulative trust across the research community.3

6. Limitations to Keep in Mind

Weighting can correct for measured imbalances but cannot fix problems it does not know about. If people who choose to take a survey differ from non-respondents in ways not captured by the weighting variables, the bias remains.4,5

Weighting on many variables with a small sample can produce extreme individual weights that inflate variance. Trimming those weights to reduce variance reintroduces a different kind of bias. These are not reasons to distrust weighted surveys, but they are reasons to examine the weighting documentation carefully.10

Political variables such as vote choice and party identification are increasingly used as weighting variables, but no consensus benchmark exists for them. Different organizations make different choices about whether and how to weight on partisanship, which can produce different results from the same underlying data.4 Election surveys of likely voters sometimes use these types of variables to weight, because the target population is not the general public, but the people who show up to vote. The demographics of voters can be very different from the general public.  For this reason, some election polls do not weight at all. 

Benchmarks cover demographics well but cannot account for attitudinal or behavioral self-selection. Weighting cannot correct for questionnaire effects or sponsorship bias. The weighting cluster is most useful when considered alongside other methodological disclosure elements rather than as a standalone quality signal.1,5

Historical disclosure

Non-probability polls from the early years of polling generally did very little weighting after fieldwork was completed. Instead, they used benchmarks based on Census or electoral information to set quotas by demographics. For the largest U.S. survey organizations, the quotas used are well-documented, but for smaller organizations this information may have been lost. After probability polling became dominant in the United States, but before telephone polling emerged, “times at home” weighting was heavily used to weight respondents based on their self-reported frequency of being home at the time of the interview, with those rarely home weighted more.

Before the 1970s, those organizations that did weight survey data sometimes used a method known as “card-weighting.” The punchcards for those respondents who were to be weighted more heavily were run through the card-counting machine multiple times, making the count of records in the final dataset higher than the final. Roper Center identifies such surveys whenever possible, but such weighting was not always documented. Completely identical respondent records should be a sign that there may be unreported card-weighting.

7. Variations That Exist

Practitioners make a range of legitimate choices across all three elements, and those choices affect how the resulting data should be interpreted.1

Weight variable variations. Some datasets include a single weight variable for all analyses. Others include multiple weight variables (for example, one for the total population and one for registered voters). Some datasets are released unweighted. Design weights correct for sampling design but not nonresponse. Calibration or raking weights iteratively adjust to match multiple population margins simultaneously. Trimmed weights cap extreme values at a threshold to reduce variance.1,4

Benchmark source variations. The most common benchmark is the U.S. Census Bureau's American Community Survey. The Current Population Survey is also widely used, especially for political surveys. Other sources include the Decennial Census, registered voter files (for voter surveys), the CDC National Health Interview Survey (for health surveys), exit polls, and proprietary benchmarks that are internally derived and not publicly documented.2,6

Weighting variable variations. Common weighting dimensions include age (often in bands), gender, race and ethnicity, education (typically four categories), geographic region or state, and urbanicity. Political surveys sometimes weight on party registration or past vote. Health surveys may weight on insurance status or chronic conditions. Raking (iterative proportional fitting) allows simultaneous adjustment across multiple variables. Education-by-race interactions are increasingly used but not widely understood by general audiences.4,8

8. Related Elements

These three weighting elements form a tightly coupled cluster, but they also connect outward to several other disclosure items. The weighting cluster interacts with Sample Design, since design weights depend on the selection probability. It connects to Sample Size, since small samples constrain how many weighting variables can be used without producing extreme weights. It relates to Margin of Error, since weighting inflates the effective design effect. It also intersects with Response Rate and Disposition Codes, since nonresponse adjustment is often a component of the final weight.1,3

If an analyst is assessing a survey's overall quality, the weighting cluster should be read alongside these elements.3

References

1. AAPOR Transparency Initiative Disclosure Elements (April 2021), Item 9. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. Roper Center Glossary of Terms: Weighting. ropercenter.cornell.edu

3. AAPOR Code of Professional Ethics and Practices, April 2021, Section III. aapor.org/standards-and-ethics/

4. Kalton, G. & Flores-Cervantes, I. (2003). Weighting Methods. Journal of Official Statistics, 19(2), 81-97.

5. Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143.

6. NORC AP-NORC Center for Public Affairs Research, December 2025 Methodology Statement. apnorc.org

7. Ipsos KnowledgePanel Methodological Overview. ipsos.com

8. NORC AmeriSpeak Technical Overview. norc.org/amerispeak/

9. Gallup Panel Methodology. gallup.com/174158/gallup-panel-methodology.aspx

10. Battaglia, M., Hoaglin, D., & Frankel, M. (2009). Practical Considerations in Raking Survey Data. Survey Practice, 2(5).

1. Plain-Language Definition

The response rate tells you what fraction of the people who were supposed to be in the survey actually ended up participating. AAPOR defines six slightly different ways to calculate it (RR1 through RR6), depending on how strictly you count cases whose eligibility is unknown. A higher response rate is generally preferred, but a low rate does not necessarily imply that the results are biased.1

Knowing the response rate matters because it is the most widely cited single indicator of data collection effort and potential nonresponse error. AAPOR's Transparency Initiative requires that organizations releasing survey findings report response rates calculated according to AAPOR Standard Definitions, or explain why a standard rate cannot be computed.2 A standard response rate cannot be computed for a non-probability-based survey: see Completion/Participation rate.

Disposition codes are the labels a survey team assigns to every sampled case at the end of data collection. They record what ultimately happened with each case: was the person interviewed, did they refuse, were they never reached, or were they found to be ineligible? AAPOR's Standard Definitions document provides a shared vocabulary so that different organizations classify these outcomes in the same way. This makes it possible to calculate and compare response rates across surveys.1

Knowing the disposition codes matters because they are the raw material behind every response rate calculation and provide more detailed information about the basis for the rate. AAPOR's Transparency Initiative requires that organizations releasing survey findings disclose enough information about case dispositions for users to independently assess data collection quality.2

Roper Center’s Transparency Project scoring requires either a response rate and indication of what formula was used to calculate it or disposition codes sufficient to calculate a response rate for surveys where a response rate is appropriate. Both are encouraged. For non-probability-based surveys where a response rate cannot be calculated, the Transparency Project requires a completion or participation rate and an explanation of how it was calculated. 

2. Why It Matters for Interpreting Results

When a large share of sampled people do not respond, and when those non-respondents differ systematically from respondents on the topic being measured, the survey estimates may be biased. However, research has shown that the relationship between response rates and actual bias is weak and inconsistent. A meta-analysis of 59 studies found that nonresponse rate explained only about 11% of the variation in nonresponse bias across different survey estimates.3

Users who ignore the response rate risk treating all surveys as equivalent regardless of how thoroughly data collection was conducted. The response rate, while imperfect, provides at least a baseline indicator of effort.5

Disposition codes go further, giving users the final status of every sampled case. Disposition codes reveal whether a low completion count reflects high refusal, low contact, or widespread ineligibility. Each of these has different implications for potential bias.1

Full disposition codes allow calculation of all six AAPOR response rate variants. Transparency about non-contacts and refusals allows nonresponse bias assessment. Any rigorous claim about response rate methodology requires this level of detail.1,6

Disposition codes can reveal information not shown by the final response rate. For example, a survey with a 30% response rate due mostly to non-contact implies different risks than one with a 30% rate due mostly to refusal. Disposition codes, therefore, are preferred by advanced researchers. For many users, however, a calculated response rate provides a simpler comparison across surveys. 

3. What Effective Disclosure Looks Like

Effective disclosure names the specific AAPOR response rate variant used (for example, RR3), provides the numerical rate, and shows full set of final disposition codes organized by the categories in AAPOR Standard Definitions (complete, partial, refusal, non-contact, ineligible, and unknown eligibility), along with the count or proportion of cases in each category.. For multi-stage designs such as online probability panel surveys, cumulative rates that account for each stage of recruitment and participation represent best practice.1,5

Well-documented records also discuss limitations of the reported rate. For example, they note whether the rate was weighted, whether partial interviews were included, and whether any nonresponse bias analysis was conducted.2

For online opt-in panels, effective disclosure acknowledges that a traditional response rate cannot be calculated and instead reports a participation or completion rate with a clear explanation of its denominator.7

4. What Incomplete Disclosure Looks Like

Common problems include reporting a response rate without naming the AAPOR variant, reporting a proprietary or non-standard rate that inflates the number, or claiming a response rate for an opt-in online panel where the concept does not meaningfully apply. Another red flag is omitting any mention of nonresponse without explanation.1,7 Another common gap is reporting a rate without clarifying whether partial interviews were counted as completions.1

Reporting only a cooperation rate (which excludes non-contacts from the denominator) while calling it a response rate overstates the survey's reach. Similarly, excluding cases of unknown eligibility from the denominator without justification produces a rate that appears higher than warranted.1

Lack of response rate or disposition codes is itself informative. It typically means either the organization did not calculate one or chose not to disclose it.2

5. Strengths When Well-Disclosed

A clearly reported AAPOR response rate enables standardized comparison across surveys. Users can evaluate whether one study's data collection was more thorough than another's and adjust their confidence accordingly.1

For probability-based online panels, cumulative response rate reporting across recruitment stages allows users to see where nonresponse accumulated and whether follow-up efforts successfully reached underrepresented groups. For example, NORC's AmeriSpeak panel reports that nonresponse follow-up recruitment reduced total absolute bias by 5 to 21 percentage points compared to initial-stage recruits alone.7

Thorough disclosure supports the kind of informed interpretation that methodologists themselves apply when evaluating poll quality.5When disposition codes are fully disclosed, they give users an unusually detailed view of data collection quality. Users can calculate any of the six AAPOR response rate variants themselves, compare rates across studies on a standardized basis, and identify specific stages of data collection (contact, cooperation, eligibility screening) where problems arose.1

Full disposition reporting supports accountability. If a survey organization reports an inflated response rate, other researchers can check the underlying codes and recalculate.2,7 

For probability-based panels, disposition code transparency across recruitment stages allows users to evaluate where nonresponse accumulated and whether follow-up efforts successfully reached underrepresented groups.8

6. Limitations to Keep in Mind

Response rate has well-documented limitations as a sole indicator of data quality. The most important is that a low rate does not necessarily produce bias, and a high rate does not guarantee representativeness. A survey with a 60% response rate and one with a 10% rate can have similar or reversed levels of bias depending on the correlation between nonresponse and the survey variables. Research has shown that nonresponse rate by itself explains only about 11% of the variation in nonresponse bias across studies.4Over-emphasis on the rate can lead organizations to pursue costly contact attempts with diminishing returns, or to pressure interviewers in ways that compromise data quality.3,4

Online opt-in panels have no meaningful response rate in the traditional sense because there is no defined sampling frame from which a denominator can be constructed. Using the term "response rate" in that context is misleading.9 Applying disposition codes designed for probability samples to non-probability contexts can create a misleading impression of rigor.9

Response rates for most surveys have declined significantly since the 1990s. Telephone surveys in the 1980s and 1990s routinely achieved response rates of 50 to 70 percent. By the 2010s, rates for major RDD surveys had fallen to single digits. This decline prompted methodological debate about whether response rate remained a useful quality indicator, leading AAPOR to commission multiple task force reports exploring alternatives.3,10

The response rate is most useful when considered alongside other methodological disclosure elements rather than as a standalone quality signal.2

7. Variations That Exist

Response Rate Variations

AAPOR defines six response rate formulas (RR1 through RR6), four cooperation rate formulas (COOP1 through COOP4), three contact rate formulas (CON1 through CON3), and three refusal rate formulas (REF1 through REF3). The key difference among the RR variants is how they handle cases of unknown eligibility. RR1 is the most conservative, counting all unknowns as eligible non-respondents. RR6 uses an estimated eligibility proportion.1 Among the surveys in the Roper Center archive, the most commonly used formula is RR3.

For online panels where a traditional response rate cannot be computed, AAPOR recommends reporting a participation rate, defined as the number of respondents who provided a usable response divided by the total number of initial personal invitations requesting participation. The AAPOR Response Rate Calculator (Version 5.1) is a publicly available spreadsheet that computes all rate variants from disposition code inputs. There is also an online interactive version.  1,11

Disposition Code Variations

The AAPOR Standard Definitions recognize different disposition code tables for different survey frames (list samples, address-based samples, random-digit-dial telephone, and online panels). Within each table, practitioners choose among layered subcodes that provide operational detail beyond the minimum required for rate calculation.1

Subtypes of disposition include completed interview, partial complete, refusal, break-off, non-contact, ineligible, unknown eligibility, and language barrier. The 10th edition of Standard Definitions reorganized these codes by sampling frame rather than by data collection mode, reflecting how surveys are designed today.1

Some organizations also use establishment survey disposition codes, which account for multi-respondent units and partial data submissions. AAPOR released a separate first-edition guide for establishment surveys in 2025.12

8. Related Elements

Response Rate and Disposition Codes are \ connect to Completion or Participation Rate, since those alternative metrics are used when traditional disposition codes cannot be meaningfully assigned. The Sampling Frame determines which cases appear in the denominator, and the Survey Mode dictates which disposition code table applies.1,2

9. Further Reading


AAPOR Standard Definitions, 10th edition (2023). aapor.org/standards-and-ethics/standard-definitions/

Groves, R.M. & Peytcheva, E. (2008). Public Opinion Quarterly, 72(2), 167-189. doi:10.1093/poq/nfn011

Baker, R. et al. (2013). Journal of Survey Statistics and Methodology, 1(2), 90-143. doi:10.1093/jssam/smt008

National Academies (2013). Nonresponse in Social Science Surveys. nationalacademies.org

AAPOR Response Rate Calculator V5.1. aapor.org/standards-and-ethics/standard-definitions/

 

References

1. AAPOR Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th edition (2023). aapor.org/standards-and-ethics/standard-definitions/

2. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

3. Groves, R.M. & Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias. Public Opinion Quarterly, 72(2), 167-189.

4. National Academies (2013). Nonresponse in Social Science Surveys: A Research Agenda. nationalacademies.org

5. Hillygus, D.S. (2011). The Evolution of Election Polling in the United States. Public Opinion Quarterly, 75(5), 962-981.

6. AAPOR Standard Definitions Webinar Slides (2024). aapor.org

7. Montgomery, R., Dennis, J.M. & Ganesh, N. Response Rate Calculation Methodology for Recruitment of a Two-Phase Probability-Based Panel: The Case of AmeriSpeak. amerispeak.norc.org/research/

8. NORC AmeriSpeak Technical Overview. Norc.org/amerispeak

9. Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143.

10 Groves, R.M. (2006). Nonresponse Rates and Nonresponse Bias in Household Surveys. Public Opinion Quarterly, 70(5), 646-675.

11. AAPOR Response Rate Calculator V5.1. aapor.org/standards-and-ethics/standard-definitions/

12. AAPOR Standard Definitions for Establishment Surveys, 1st edition (2025). aapor.org

13. NORC AmeriSpeak Technical Overview. norc.org/amerispeak/

1. Plain-Language Definition

A completion or participation rate tells you what share of people who were invited to take a survey actually finished it. Unlike a traditional response rate, which is built from disposition codes and calculated using AAPOR formulas (covered in the Disposition Codes and Response Rate elements), this metric is used for surveys where there is no probability-based sampling frame and therefore no standard denominator. It measures cooperation among those who were contacted, not coverage of a defined population.1

The completion or participation rate and its calculation details are two sides of the same coin. The rate alone is just a number. The calculation details alone have nothing to explain. This document treats them as one integrated topic.1 Both are required together in Roper Center’s Transparency Project.

The calculation details matter because two surveys can report identical completion rates yet measure fundamentally different things if their denominators differ. One organization might divide completed interviews by total invitations sent. Another might divide by only those invitations that were opened or confirmed delivered. Without the formula, the rate has no context.1,2

2. Why It Matters for Interpreting Results

For a growing share of survey research, especially studies conducted through online opt-in panels, traditional AAPOR response rates cannot be computed because there is no defined sampling frame. In these cases, the completion or participation rate is the primary available measure of data collection effort. AAPOR and the ISO have both recommended that non-probability surveys report participation rates rather than response rates to avoid confusion with probability-based metrics.1,2

Users who ignore this element risk comparing surveys on different metrics without realizing it. A study reporting a 60% "completion rate" based on delivered invitations is not comparable to one reporting a 60% "response rate" based on a full probability sample. The denominator defines what the rate means.1

Understanding the calculation details also helps users assess whether the reported rate is inflated. Excluding non-delivered invitations, screening failures, or break-offs from the denominator can substantially raise the reported number without reflecting any improvement in actual data collection quality.2,3

3. What Effective Disclosure Looks Like

Effective disclosure reports the completion or participation rate alongside a clear explanation of how the numerator and denominator were defined. It specifies whether the numerator includes only complete interviews or also counts partial interviews. It specifies whether the denominator includes all invitations sent, only confirmed-delivered invitations, or only those who opened the survey link.1

Well-documented records use AAPOR's cooperation rate formulas (COOP1 through COOP4) when applicable, naming the specific formula used. They also note the break-off rate separately so users can see how many people started the survey but did not finish.1,4

For panel-based surveys, effective disclosure reports the rate at each stage of the process: the panel recruitment rate, the survey invitation rate, and the survey completion rate. This allows users to see where attrition occurred and to calculate a cumulative figure if needed.5

4. What Incomplete Disclosure Looks Like

Incomplete disclosure reports a rate without explaining the denominator. A "75% completion rate" is meaningless if the user cannot determine whether it was calculated from total invitations, delivered invitations, or only those who clicked into the survey.1

Another common gap is labeling a cooperation rate as a "response rate," which overstates the survey's reach by excluding non-contacts from the denominator. Using the term "response rate" for a non-probability panel survey is itself a form of incomplete disclosure because it implies a probability-based calculation that was not performed.2

Omitting the break-off rate is also a concern. High break-off rates can signal questionnaire design problems, excessive survey length, or panel fatigue, all of which affect data quality independently of the final completion number.4

5. Strengths When Well-Disclosed

When the completion rate and its calculation details are fully disclosed, users can meaningfully compare participation levels across non-probability surveys. This is particularly valuable in a research landscape where opt-in panel surveys are increasingly common.2

Detailed rate reporting also supports quality assessment within a single study. If the break-off rate is high at a particular question, this may indicate a measurement problem that could bias estimates on that topic.4

Transparent reporting of stage-by-stage rates for panel surveys enables users to identify where the biggest losses occurred and to judge whether the final sample is likely to differ in important ways from the initial invited sample.5

6. Limitations to Keep in Mind

A completion rate, even when well-disclosed, measures only cooperation among those who were reached. It says nothing about whether the people who were invited to participate in the first place are representative of the broader population. Coverage error in the underlying panel is a separate and often larger concern.2

Participation rates from different panel providers are difficult to compare even with full disclosure, because each provider defines its invitation pool differently and may use different pre-screening, routing, or deduplication procedures before an invitation is counted.2,4

As with traditional response rates, completion rates do not directly measure bias. A high completion rate does not guarantee representative results, and a low rate does not necessarily produce biased estimates.3

7. Variations That Exist

Rate variations include AAPOR cooperation rates COOP1 through COOP4, the International Organization for Standardization (ISO)-defined participation rate, view rate (proportion of panel members who saw the invitation), completion rate (proportion who finished among those who started), and screening completion rate (proportion who completed a screener among those invited).1,4

Calculation variations differ primarily in how the denominator is defined. Some organizations count all invitations sent. Others exclude undeliverable invitations, quota-filled cases, or screened-out respondents. Some weight the rate by base weights to account for unequal selection probabilities. The AAPOR Standard Definitions recommend that any rate be accompanied by the formula used and the disposition code counts that went into it.1

8. Related Elements

Completion and Participation Rate serve as substitutes for the Response Rate when a traditional response rate cannot be computed. They draw on the same disposition code framework described in the Disposition Codes element, though they apply it differently. These elements also connect to the Sampling Frame and Sampling Procedure elements, since these determine whether a response rate or participation rate is the appropriate metric.1,2

References

1. AAPOR Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th edition (2023). aapor.org/standards-and-ethics/standard-definitions/

2. Baker, R. et al. (2013). Summary Report of the AAPOR Task Force on Non-Probability Sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143.

3. Groves, R.M. & Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias. Public Opinion Quarterly, 72(2), 167-189.

4. Callegaro, M. & DiSogra, C. (2008). Computing Response Metrics for Online Panels. Public Opinion Quarterly, 72(5), 1008-1032.

5. NORC AmeriSpeak Technical Overview. norc.org/amerispeak/

1. Plain-Language Definition

Survey language refers to which language or languages a survey was made available in, such as English only, English and Spanish, or additional languages, and whether the questionnaire, interviewer scripts, and respondent-facing materials were translated. This element captures not just whether translation occurred, but which languages were offered and how translation quality was managed.

AAPOR's Transparency Initiative disclosure checklist pairs language with data collection mode, requiring disclosure of "the language(s) offered or included" as part of the method and mode description. AAPOR's Best Practices guidance addresses language from the interviewer-training perspective, specifying that when a survey is administered in languages other than English, interviewers should demonstrate language proficiency and cultural awareness, and training should address how to conduct non-English interviews appropriately. Together, these standards establish that language is not a minor logistical detail but a core component of methodological transparency.1,2,3

2. Why It Matters for Interpreting Results

Language coverage determines who could participate in a survey, which in turn shapes what the data can credibly represent. English-only surveys systematically exclude non-English-speaking populations, introducing coverage error that may go unnoticed unless the survey's language design is disclosed. In the United States, approximately one in eight people speaks a language other than English at home, and about half of Spanish-speaking adults are considered limited English proficient. When a survey's target population includes these communities, as most general-population surveys intend to, fielding in English alone means that a meaningful segment of the population had no realistic opportunity to participate.5

Research on surveying Latino populations in the United States illustrates how language design choices shape sample composition in concrete ways. Pew Research Center has documented that using a fully bilingual interviewing staff, where all interviewers can conduct interviews in either English or Spanish, produces more interviews with Spanish-dominant and foreign-born Hispanics compared to a modified bilingual approach, where mostly English-speaking interviewers arrange callbacks from bilingual staff when a Spanish-speaking respondent is encountered. These two interviewing models yield samples with measurably different demographic profiles, particularly on language dominance and nativity. While appropriate weighting can reduce the resulting differences on many survey estimates, the underlying sample composition is shaped by the language infrastructure before any statistical adjustment occurs.4

Experimental research on multilingual survey administration reinforces the point. A study of a local address-based survey in Nebraska found that including Spanish-language materials had mixed effects on response rates and sample composition, sometimes increasing participation among Spanish speakers, sometimes producing no effect, and occasionally triggering a backfire effect where non-Spanish-speaking recipients disengaged. These findings suggest that the decision to offer additional languages interacts with other features of survey design (such as mode and geographic targeting) in ways that are difficult to predict without knowing the specific language protocol used.5

At a broader level, the AAPOR/WAPOR Task Force on Quality in Comparative Surveys has identified language and cultural comparability as a defining challenge for survey quality in multinational, multiregional, and multicultural research contexts. The pursuit of data quality in cross-cultural surveys is simultaneously the pursuit of comparability, and language is one of the most consequential dimensions along which comparability can succeed or fail. When language details are disclosed, users can assess whether the sample meaningfully represents all populations relevant to the research question. When they are not, users are left to assume, often incorrectly, that the survey's findings apply equally to populations who may never have had the opportunity to participate.6

3. What Effective Disclosure Looks Like

Effective disclosure names every language in which the survey was fielded, describes the translation process used, and ideally reports language-specific sample sizes or response distributions. The minimum standard, as set by AAPOR's Transparency Initiative checklist (Item 6), is to identify the languages offered or included alongside the mode of data collection. But best practice goes considerably further.1

Pew Research Center's 2022-2023 Asian American Survey provides a strong model of thorough language disclosure. Its methodology documentation names all six fielding languages (English, Simplified Chinese, Traditional Chinese, Hindi, Korean, Tagalog, and Vietnamese), describes the web interface's language-toggle functionality that allowed respondents to switch languages via a dropdown on each page, and reports that an independent linguist review compared the translated instruments to the English source as an additional quality-control step. Paper surveys were formatted in all six languages, and recipients believed to be more likely to use a specific language, based on supplemental information in the sampling frame or address location, were sent a paper screener in that language alongside an English version. This level of detail allows users to evaluate both the breadth of language coverage and the quality of the translation process.8

Similarly, well-documented surveys report the exact language split in the achieved sample. For example, KFF regularly reports the exact number of respondents who were interviewed in English and Spanish, as in this example from their 2026 survey: “The survey was conducted May 4 – May 26, 2026, online and by telephone among a nationally representative sample of 25,873 U.S. adults in English (n=25,422) and in Spanish (n=451).” This kind of language-specific sample-size reporting enables users to assess whether the language groups are large enough for meaningful subgroup analysis and whether the proportions are plausible given the target population.9

4. What Incomplete Disclosure Looks Like

Several patterns signal inadequate language disclosure. The most common is English-only fielding with no discussion of whether the population's language composition is relevant to interpreting the results. When a survey targets the general U.S. adult population but says nothing about language, users cannot determine whether non-English speakers were excluded by design or simply not mentioned. Another red flag is generic language such as "available in multiple languages" without naming which ones or specifying how many respondents used each language.

Incomplete translation disclosure is equally problematic. Research on survey translation quality has documented that a single-pass vendor translation without adjudication, bilingual-staff review, or cognitive testing is a common but incomplete baseline. A study of practical translation approaches found it necessary to add four additional quality-control steps beyond the initial vendor translation: an internal review by native-speaking survey researchers, client adjudication of recommended revisions, bilingual interviewing staff review for grammatical and spelling errors, and bilingual staff testing of the programmed survey. The fact that these additional steps were not in the original budget or timeline for either project studied suggests that minimal translation processes are widespread. A methodology report that simply states "the survey was translated into Japanese" without describing the translation method leaves users unable to assess whether the translated instrument measures the same constructs as the original.10

A related gap occurs when methodology reports do not specify which interviewing model was used for bilingual data collection. Without knowing whether a fully bilingual or modified bilingual staffing approach was employed, users cannot assess whose responses the data actually represent, since these models produce measurably different sample profiles.4

5. Strengths When Well-Disclosed

When disclosed well, language coverage tells users directly whether all members of the target population  were represented and how. This is one of the most straightforward elements to evaluate: if the survey names its fielding languages and reports language-specific sample sizes, users can immediately assess whether the data can support claims about specific populations.

Reporting language-specific sample sizes allows analysts to check whether findings hold across language groups. Research on Latino public opinion has demonstrated that foreign-born and U.S.-born Hispanics show significant differences of opinion on some key issues, not always consistently in the same ideological direction, making it important to know whether each subgroup is adequately represented in the data. A survey that discloses it achieved, say, 679 Spanish-language and 821 English-language interviews gives users a concrete basis for evaluating whether the Spanish-language subsample is large enough for the analysis being presented.4

Named translation processes allow methodologists to evaluate measurement equivalence. A survey that describes using a team-based translation approach with independent review and cognitive testing provides much stronger grounds for cross-language comparison than one that mentions translation without specifying the method. Multi-language fielding, when paired with detailed documentation, strengthens the credibility of a survey's claim to represent its stated target population.8

6. Limitations to Keep in Mind

Translation quality affects measurement validity in ways that are not visible from language disclosure alone. A poorly translated item may not be equivalent to its original language counterpart, meaning that apparent differences between language groups could reflect translation artifacts rather than genuine population differences. Research on translation equivalence has identified three kinds of equivalence that a translation can fail to achieve: semantic equivalence (same meaning), conceptual equivalence (same construct), and normative equivalence (similar responses from target populations). These forms of equivalence are rarely tested formally outside of dedicated methodological studies, so users generally cannot assume that the translated version measures the same thing as the original.7

Partial translation, where some items are translated but others are not, creates inconsistent measurement that is easy to miss. If a survey translates its main substantive questions into multiple languages but leaves demographic or attitudinal battery items in the original language only, respondents who selected the alternative language version may encounter items they cannot fully understand, producing measurement error that is invisible in the final dataset.

Offering a language is also not the same as respondents understanding or using it as intended. Research on respondent language choice has found that in one study, nearly three out of four respondents (73%) in households where some Spanish was spoken actually requested English-only materials. Even among those who spoke mostly Spanish at home, only 56% requested bilingual materials. This finding is a reminder that language-of-interview does not map neatly onto language ability, and that the relationship between offering a language and actually reaching the intended population is more complex than it may appear.11

In addition, those conducting research across countries where multiple languages are spoken have to balance the advantages of a range of translations with potential incomparability of results. For example, AfroBarometer provides guidance to its local partners: “In principle, every language group that is likely to constitute at least 5% of the sample should have a translated questionnaire. In practice, because of the complications and costs introduced by too many versions of the questionnaire, it is desirable to limit the number of local language translations to no more than six, and preferably fewer.” 12  

The AAPOR/WAPOR Task Force on Quality in Comparative Surveys has cataloged data-quality challenges specific to multilingual and cross-cultural survey research as an ongoing limitation of the field, noting that the pursuit of comparability across languages and cultures remains a formidable and incompletely solved challenge.6

7. Variations You May Encounter

The most common U.S. bilingual survey design is English and Spanish, but surveys targeting specific populations may field in many additional languages. Pew Research Center's Asian American Survey, for example, was fielded in six languages to reflect the linguistic diversity of its target population. Beyond the number of languages, surveys vary in how translation is produced and how language administration works in practice.8

Translation approaches range from professional vendor translation reviewed by a committee to in-house translation without a documented review process. Research on practical translation methods has described an adapted committee approach, a single translation with an internal review team, as a workable middle ground for survey organizations with limited budgets and timelines. The most rigorous approaches involve team-based translation with independent linguistic review, cognitive testing with members of the target population, and quantitative evaluation of translated items through test-retest methods.10

Interviewer staffing models also vary. In the fully bilingual model, all interviewers assigned to areas with high concentrations of the target language group can conduct interviews in either language. In the modified bilingual model, most interviewers speak only English, and when they encounter a respondent who prefers another language, they arrange for a bilingual interviewer to call back. These two approaches produce measurably different sample profiles, with the fully bilingual method reaching more respondents who are dominant in the non-English language.4

In web-based surveys (CAWI), language administration may involve automatic language detection, respondent-selected language switching via a dropdown or toggle, or separate survey links for each language. In computer-assisted personal interviewing (CAPI) or telephone interviewing (CATI), the interviewer typically selects the language based on the respondent's preference.8

8. Related Elements

Survey Language connects to several other disclosure elements. It is closely related to Universe and Geographic Coverage, since the relevance of language coverage depends on the population and geography being studied. A general-population survey of a linguistically diverse metropolitan area raises different language expectations than a survey of a specialized professional population that operates primarily in English.

Language is also connected to Sampling Frame, since frames built from surname lists or geographic strata are often used specifically to reach non-English-speaking segments. Pew Research Center's Asian American Survey, for example, tied its sampling frame (an address-based sample supplemented with surname list frames for Chinese, Filipino, Indian, Korean, and Vietnamese households) directly to its choice of six fielding languages. The connection between who is sampled and what languages are offered is not incidental; the two design choices are interdependent.8

Survey Language is also worth cross-checking against Survey Organization. An organization's documented capacity for multilingual data collection, whether it maintains trained bilingual interviewers, has in-house translation review capability, or routinely fields in multiple languages, is itself a signal about how seriously language coverage was likely handled for a given study.13

9. Further Reading

AAPOR/WAPOR Task Force Report on Quality in Comparative Surveys. wapor.org/resources/aapor-wapor-task-force-report-on-quality-in-comparative-surveys/

AAPOR Best Practices for Survey Research. aapor.org/standards-and-ethics/best-practices/

Pew Research Center. (2015). "The Unique Challenges of Surveying U.S. Latinos." pewresearch.org/social-trends/2015/11/12/the-unique-challenges-of-surveying-u-s-latinos-2/

RTI International. "Quantitative Evaluation of Response Scale Translation," Chapter 4. rti.org/sites/default/files/6-surveylanguage-chapter4.pdf

Agans, R. P. et al. "Test-Retest Approach to Evaluating Survey Translation." PMC 7473424.

References

1. AAPOR Transparency Initiative Disclosure Checklist, Item 6. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. AAPOR Best Practices for Survey Research, Interviewer Training section. aapor.org/standards-and-ethics/best-practices/

3. AAPOR Code of Professional Ethics and Practices, Section III.A. aapor.org/standards-and-ethics/disclosure-standards/

4. Pew Research Center. (2015). The Unique Challenges of Surveying U.S. Latinos. pewresearch.org/social-trends/2015/11/12/the-unique-challenges-of-surveying-u-s-latinos-2/

5. Survey Practice. The Effects of Spanish-Language Materials in a Local Area ABS Mixed-Mode Survey. surveypractice.org/article/125781

6. AAPOR/WAPOR Task Force Report on Quality in Comparative Surveys. wapor.org/resources/aapor-wapor-task-force-report-on-quality-in-comparative-surveys/

7. Agans, R. P. et al. Test-Retest Reliability Approach to Evaluating Survey Translation Quality. PMC 7473424.

8. Pew Research Center. (2023). Asian American Identity Methodology. pewresearch.org/race-and-ethnicity/2023/05/08/asian-american-identity-methodology/

9. KFF. (2026) Examining LGBTQ+ Adults’ Experiences with Health Care Costs and Access. https://www.kff.org/public-opinion/examining-lgbtq-adults-experiences-with-health-care-costs-and-access/#daae4d64-96a4-4146-9c37-660571948e70

10. Survey Practice. Development of a Practical Translation Approach for Survey Research Projects. surveypractice.org/article/127842

11. Survey Practice. Spanish Respondents' Choice of Language: Bilingual or English. surveypractice.org/article/3021

12. AfroBarometer. (2025). Round 10 Survey Manual. https://www.afrobarometer.org/survey-resource/round-10-survey-manual/

13. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 2. aapor.org/standards-and-ethics/disclosure-standards/

1. Plain-Language Definition

Full question wording means the complete text of every question that produced reported results, in the order it was asked, with every element of context that shaped what respondents actually saw or heard. That includes the target question itself, the response options, any introductory or transitional language, any preceding questions that could have primed the respondent, any interviewer instructions (such as 'do not read options'), any definitions or explanations that appeared before the question, and any visual aids (show cards, images, sliders, screenshots) that were part of the presentation.1

This is not the same as reporting a paraphrase or a summary of what the survey asked about. The exact wording matters because small changes in phrasing, response format, or context can produce measurably different results.2

2. Why It Matters for Interpreting Results

Survey answers are shaped by the specific words used to ask the question, the order the questions appeared in, the response options offered, and the way the questionnaire was presented or the interview was conducted. Decades of methodological research have documented that question order effects, response-scale effects, and wording effects are large and predictable enough to matter for how results should be interpreted.2,3

Consider a simple example. A yes/no question and a five-point Likert scale question about the same topic will produce different distributions of answers. A question that offers 'agree/disagree' typically yields different results than a question with item-specific response categories, even when the underlying concept is the same. A question that is asked after a related question can sometimes “prime” respondents and yield different results than the same question asked in isolation.2

Without access to the full wording, users cannot tell whether a reported result reflects the underlying opinion or an artifact of question design. This is especially important when a survey is being cited to support a policy claim or when two surveys on the same topic report different numbers. Very often, differences across surveys turn out to be differences in how the question was asked.2,4

3. What Effective Disclosure Looks Like

Effective disclosure provides the complete questionnaire, in the order administered, with every element the respondent encountered. That includes screening questions used to establish eligibility, transitions and preambles, exact target question text, all response options exactly as offered, and any 'do not read' or 'read aloud' interviewer instructions.1

Any visual aids used (show cards, images, response sliders, illustrations) should be included or clearly described. For self-administered surveys, screenshots or facsimiles of the screens as respondents saw them are useful, especially where layout or visual presentation affects responses.1,2

KFF’s topline documents are a widely cited model. Each report links to a downloadable topline containing exact question text, response options, and skip logic. Users can walk through the full instrument as respondents experienced it and evaluate any specific wording or ordering concern.5

In addition, surveys fielded in multiple languages should make the translations used available. 

4. What Incomplete Disclosure Looks Like

The most common gap is providing only a summary or paraphrase of question topics, rather than the exact wording. Reports that describe results by saying 'respondents were asked about their views on X' without showing the actual question fall well short of the AAPOR standard.1,6

Other frequent gaps include providing question text but omitting response options; showing questions out of order or without preceding context; leaving out screening questions that shaped who was in the sample; and hiding interviewer instructions that changed what respondents actually heard. Visual aids are rarely disclosed but can substantially affect responses when they are used, and their absence from the disclosed materials is a real gap.2

'Available on request' is not equivalent to proactive disclosure. The AAPOR Transparency Initiative expects that measurement tools will be disclosed at the time results are released, not held back until someone specifically asks.6,7

5. Strengths When Well-Disclosed

A full questionnaire allows users to evaluate context effects, question-order effects, and any potentially leading language. When a poll produces a surprising result, the first thing an experienced reader wants to see is the exact question, in context, so they can judge whether the wording is doing the work.2,3

Full disclosure also supports replication and comparison. Two surveys reporting different numbers on the same topic can be evaluated to see whether the difference is in the world or in the wording. Researchers designing new studies can build on established question sets rather than reinventing measures.2

Where sponsorship pressure is a concern, question wording is often where it shows up. A full instrument lets users check whether the phrasing of the questions could plausibly have favored the sponsor's preferred conclusion.1,4

6. Limitations to Keep in Mind

Even a fully disclosed questionnaire cannot answer every methodological question. The written wording does not capture how an interviewer inflected the question aloud, how quickly a respondent moved through a self-administered survey, or how attentive respondents were at the time of the interview. Full wording is necessary for evaluation, but not sufficient.2

Question wording is also only part of the measurement story. Even a well-worded question can produce misleading results if the sample is unrepresentative, the mode is a poor fit for the topic, or the response rate is very low.1,3

As AI tools become more common in questionnaire design, a new consideration is emerging: AI-drafted or AI-edited questions may have subtle systematic patterns that are hard to spot in any individual item. AAPOR now recommends that use of AI in question writing be disclosed alongside the questionnaire itself.8

7. Variations You May Encounter

Full questionnaire with exact wording

The gold standard. Every question, every response option, every instruction, in the order asked, disclosed at the time results are released. This is what the AAPOR Transparency Initiative expects.1,7

Topline only

A report showing question wording and results side by side, in a shortened format. If the topline includes all questions in order with full text and response options, it can be equivalent to full disclosure. When it lists only selected questions or omits instructions, it is incomplete.5

Summary of question topics only

A methodology note describing what the survey was about without providing the actual questions. Not adequate for evaluating results.1

'Available on request'

Wording provided only when someone asks. This falls short of the Transparency Initiative expectation of proactive disclosure at release.6

Partial disclosure

Some questions shown, others withheld (often for commercial confidentiality). This prevents evaluation of context effects and question order.1,2

Visual aids

Show cards used in phone or in-person interviews, images or graphics used in web surveys, response sliders, screenshots of self-administered layouts. These are frequently omitted from disclosure but can meaningfully affect responses when used.2

Interviewer instructions

'Do not read options aloud,' 'probe once if uncertain,' 'accept multiple answers.' These change what respondents heard and how their answers were recorded. Should appear alongside the question text.1

8. Related Elements

Full Question Wording is most closely paired with Data Collection Mode (Element 6), since visual aids and interviewer instructions are mode-dependent. It also connects to External Survey Sponsor (Element 3), because sponsor influence, where present, is often most visible in how questions are worded, and to Population Under Study (Element 4), since screening questions establish who counts as part of the target population.1

For studies using AI in questionnaire development, this element also connects to newer disclosure requirements around AI involvement in the research process.8

9. Further Reading

AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 3. aapor.org/standards-and-ethics/disclosure-standards/

Schaeffer, N.C. & Dykema, J. (2020). Advances in the Science of Asking Questions. Annual Review of Sociology 46, 37-60.

AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

AAPOR (2026). Responsible AI Integration in Survey Research. aapor.org/standards-and-ethics/reports/

Pew Research Center topline questionnaires. pewresearch.org

Tourangeau, R., Rips, L.J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge University Press.

Krosnick, J.A. & Presser, S. (2010). Question and Questionnaire Design. In Handbook of Survey Research (2nd ed.).

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A, Item 3. aapor.org/standards-and-ethics/disclosure-standards/

2. Schaeffer, N.C. & Dykema, J. (2020). Advances in the Science of Asking Questions. Annual Review of Sociology 46, 37-60.

3. Tourangeau, R., Rips, L.J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge University Press.

4. Krosnick, J.A. & Presser, S. (2010). Question and Questionnaire Design. In Handbook of Survey Research (2nd ed.), edited by P.V. Marsden and J.D. Wright.

5. KFF. Public Opinion. https://www.kff.org/topic/public-opinion/

6. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

7. AAPOR Transparency Initiative Disclosure Elements (April 2021). aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

8. AAPOR (2026). Responsible AI Integration in Survey Research. aapor.org/wp-content/uploads/2026/05/Responsible-AI-Integration-In-Survey-Research.pdf

9. Willis, G.B. (2004). Cognitive Interviewing: A Tool for Improving Questionnaire Design. Sage.

10. Dillman, D.A., Smyth, J.D., & Christian, L.M. (2014). Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method. Wiley.

1. Plain-Language Definition

An external sample provider is the organization that supplied respondents or the list of potential respondents to the organization conducting the survey. This may be a panel company, a list vendor, or a sampling-frame supplier, an entity distinct from both the survey sponsor and the organization that administered the questionnaire and collected the data. Common examples include opt-in panel vendors, probability-based panel providers, address-based sampling vendors, and list vendors for phone or mail surveys.

The proportion of sample provided is a companion element that captures how much of the total achieved sample came from each provider. A survey might draw 100% of its sample from a single vendor, split it evenly across two, or blend respondents from multiple sources with supplemental oversamples from a specialist provider. When the sample comes from multiple sources, the proportions determine which provider's characteristics most strongly influence the data.

AAPOR's Transparency Initiative disclosure checklist addresses both elements. Item 5 requires disclosure of how the sample was generated and recruited, including the name of the supplier of the sample or list if a frame, list, or panel is used. Item 8 requires sample sizes by sampling frame if more than one frame was used, which is the closest analog to proportion disclosure. AAPOR's Code of Professional Ethics further specifies that if the results reported are based on multiple samples or multiple modes, the relevant disclosure items will be disclosed for each, effectively mandating per-source reporting.1,2

These two elements are treated together in this guide because they are closely related: knowing which vendors supplied the sample is most useful when paired with information about how much of the sample each vendor contributed.

2. Why It Matters for Interpreting Results

Different sample providers use different recruitment methods, maintain panels of different compositions, and apply different quality-control standards. Knowing which provider supplied the sample, and how much of it, lets users evaluate likely coverage gaps, demographic skews, and the overall credibility of the data, as the choice of sample providers  can have measurable consequences for data quality 

A 2013 study 

A 2013 study found approximately 15-25% cross-panel overlap among six of the seven vendors, meaning that respondents enrolled in multiple panels simultaneously. This overlap means that blended samples drawn from multiple vendors may not be as independent as researchers assume, a finding that users cannot evaluate without knowing which vendors were involved.5

When multiple vendors contribute to one sample, understanding the proportions is essential for assessing whether one vendor's characteristics dominate the results. As AAPOR's Task Force on Data Quality Metrics has noted, in blended sampling, the sample purchaser decides from what panels and in what ratios to source for a specific survey. If 80% of a sample comes from a lower-quality vendor, the overall data quality reflects that vendor's characteristics far more than a 50/50 split would. Without this information, users cannot determine whether observed findings reflect the population or the particular panel composition.6

3. What Effective Disclosure Looks Like

Effective disclosure names each sample provider specifically, not generically, states whether the provider is probability-based or opt-in, reports the number or proportion of completed interviews from each provider, and ideally provides enough information for a user to look up the provider's own methodology documentation.

AAPOR's Transparency Initiative checklist provides the disclosure-standard phrasing. It requires that if a frame, list, or panel is used, the description should include the name of the supplier and the nature of the list. The checklist also requires an explicit statement of whether the sample comes from a probability-based or non-probability methodology. A well-documented example from a university survey center's Transparency Initiative filing names specific vendors, such as "Dynata, LLC (mail survey)" and "Qualtrics, LLC (online survey)," demonstrating the level of specificity the standard expects.1,8

Proportion disclosure is equally important. A Pew Research Center's benchmarking study, though it masked vendor names for methodological reasons, modeled good proportion disclosure by describing one opt-in sample as "a blend, with about three-fifths sourced from a single opt-in panel and the remainder sourced from three sample aggregators." This level of detail allows users to assess whether the blending proportions might influence the results and to evaluate the design's overall quality profile.4

4. What Incomplete Disclosure Looks Like

Several patterns signal inadequate disclosure of sample provider information. The most common is the use of generic labels such as "an online panel provider" or "a national research panel" without naming the company. This prevents users from independently verifying the provider's recruitment methods, panel composition, or track record. Listing only the survey organization when a separate vendor actually supplied the sample is another frequent gap, parallel to the problem identified in survey organization disclosure, where listing only the sponsor when a separate firm conducted the fieldwork obscures the chain of methodological responsibility.9

When multiple vendors are used, disclosing that fact without specifying proportions or identifying which vendor provided which portion falls short of AAPOR's Code, which states that if results are based on multiple samples, the relevant disclosure items will be disclosed for each. A report that says "multiple vendors were used" without per-vendor breakdowns leaves users unable to assess whether vendor-specific quality differences are driving the results.2

The AAPOR Task Force on Data Quality Metrics has noted that transparency around blended sampling practices remains limited, suggesting that incomplete vendor disclosure is a recognized field-wide concern rather than an isolated oversight.4,6

5. Strengths When Well-Disclosed

When well-disclosed, a named vendor can be independently researched. Many sample providers have publicly documented recruitment processes and known coverage properties, allowing users to assess the likely strengths and limitations of the sample without relying solely on the survey report's own description. Named vendors can also be cross-referenced in methodological research and benchmarking studies, much as named survey organizations can be looked up in AAPOR's Transparency Initiative member list.

Multi-vendor disclosure with proportions allows users to assess blending effects on data quality. If a user knows that a survey drew 60% of its sample from a probability-based panel and 40% from an opt-in source, the user can form reasonable expectations about the type of statistical claims that are supported by the data and can evaluate whether the weighting techniques described in the methodology report are appropriate for that design.  Finally, the ability to evaluate specific vendors against empirical benchmarks is possible only when the vendor identity is disclosed.5

6. Limitations to Keep in Mind

A vendor's panel composition may not match the stated survey universe, and the vendor's practices may have changed since the most recent publicly available documentation. Just as organizational reputation is a lagging indicator for survey organizations, where current practices may differ from historical track records due to budget pressures, staff turnover, or technological shifts, the same caution applies to sample providers.9

Blended samples from multiple vendors can introduce heterogeneous quality levels and systematic differences across subgroups. Cross-panel overlap, respondents belonging to multiple panels simultaneously, is common and not always detectable. A comparative study of seven vendors found approximately 15-25% overlap among six of them, likely because panelists enrolled with multiple companies. The AAPOR Task Force on Data Quality Metrics has noted that the amount of overlap across panels in terms of multiple memberships is unclear, and while some research suggests that multi-panel membership does not always affect survey-taking behavior, the uncertainty itself represents a risk that users should be aware of.5,6

Without proportion data, users cannot assess whether vendor-specific effects are driving the overall results. If one vendor's respondents have higher rates of straightlining, faster completion times, or lower screener-pass rates, these quality differences will disproportionately influence the survey's findings in proportion to that vendor's share of the sample. 

7. Types of Sample Providers You May Encounter

Opt-in panel vendors

Opt-in or volunteer panels recruit members through website banner ads, email solicitations, social media, and partner websites. Members self-select into the panel, typically motivated by incentives, the opportunity to express opinions, or entertainment.  AAPOR's 2010 Report on Online Panels noted that such panels represent a substantial departure from traditional probability-based methods. The methods used by such panels continue to evolve, and there is substantial variation in their internal practices for recruitment, quality control, and weighting.

Probability-based panel providers

Probability-based panels recruit members using traditional random-sampling methods, most commonly address-based sampling from the U.S. Postal Service's Delivery Sequence File.. These panels generally have far fewer members than opt-in panels but offer known coverage properties and calculable selection probabilities. Some provide Internet access to recruited members who lack it, reducing the coverage gap associated with online-only data collection.4

Panel aggregators and marketplaces

Panel aggregators draw respondents from many opt-in sample sources that have agreed to make their sample available. Rather than maintaining their own panel, aggregators act as intermediaries, routing respondents from multiple panels to surveys based on quota requirements. This approach can fill sample targets quickly but makes the provenance of individual respondents less transparent.4

List vendors for phone or mail surveys

These vendors supply sampling frames such as random-digit-dial telephone samples, voter registration lists, consumer databases, or address lists. They differ from panel vendors in that they supply a list of potential contacts rather than pre-recruited, survey-ready respondents.

Address-based sampling (ABS) vendors

ABS vendors provide samples drawn from the U.S. Postal Service's Delivery Sequence File, which covers nearly all residential addresses in the United States. ABS has become the dominant frame for probability-based survey recruitment and is used by all three probability-based panels evaluated in Pew Research Center's benchmarking study.4

Blended or multi-source samples

Increasingly, surveys blend sample from multiple sources to reach target sample sizes or specific subgroups. A survey might draw its core sample from a probability-based panel and supplement it with opt-in respondents to reach a sufficient number of low-incidence subgroups. When this occurs, AAPOR's disclosure standards require that the source and sample size of each component be reported, and that weighting techniques used to combine the sources, such as propensity-score matching, be described.6

8. Related Elements

External Sample Provider and Proportion of Sample Provided connect most directly to Sampling Procedure and Sampling Frame, since the provider's recruitment and panel-management practices determine what kind of frame is available and how respondents are selected from it. They also connect to Use of Breakout Routers and Quality Control, both of which describe mechanisms that operate within or alongside the sample provider's infrastructure.

These elements are closely related to Survey Organization. The conducting organization's relationship with its sample suppliers, whether it maintains an in-house panel or contracts externally, shapes the quality-control chain. AAPOR's Transparency Initiative pledge explicitly addresses this relationship: if the survey organization did not collect the research data itself, it is obligated to obtain the required disclosure information from its fieldwork subcontractor, and this obligation extends to sample providers.10

The AAPOR Transparency Initiative checklist items most directly connected to provider identity and proportions are sample recruitment method, sample sizes by frame, and weighting procedures. Together, these items create a disclosure chain that links the sample's origin to its statistical treatment.1

9. Further Reading

Baker, R. et al. (2010). AAPOR Report on Online Panels. Public Opinion Quarterly, 74(4), 711-781. aapor.org/wp-content/uploads/2022/11/nfq048.pdf

AAPOR Task Force on Data Quality Metrics for Online Samples (2023). Full report. aapor.org/wp-content/uploads/2023/02/Task-Force-Report-FINAL.pdf

Pew Research Center. (2023). Comparing Accuracy of Two Types of Online Survey Samples. pewresearch.org/methods/2023/09/07/comparing-two-types-of-online-survey-samples/

Craig, B. M. et al. (2013). Comparison of US Panel Vendors for Online Surveys. Journal of Medical Internet Research, 15(11), e260. PMC 3869084.

AAPOR Code of Professional Ethics and Practices, April 2021, Section III. aapor.org/standards-and-ethics/disclosure-standards/

References

1. AAPOR Transparency Initiative Disclosure Checklist, Items 5 and 8. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. AAPOR Code of Professional Ethics and Practices, Section III.A (multiple-samples clause). aapor.org/standards-and-ethics/disclosure-standards/

3. Baker, R. et al. (2010). AAPOR Report on Online Panels. Public Opinion Quarterly, 74(4), 711-781.

4. Pew Research Center. (2023). Comparing Two Types of Online Survey Samples. pewresearch.org/methods/2023/09/07/comparing-two-types-of-online-survey-samples/

5. Craig, B. M. et al. (2013). Comparison of US Panel Vendors for Online Surveys. Journal of Medical Internet Research, 15(11), e260.

6. AAPOR Task Force on Data Quality Metrics for Online Samples (2023). aapor.org/wp-content/uploads/2023/02/Task-Force-Report-FINAL.pdf

7. AAPOR Task Force on Data Quality Metrics, Blended Sampling section.

8. East Carolina University, TI Disclosure Elements example. surveyresearch.ecu.edu/wp-content/pv-uploads/sites/315/2018/06/TI-Disclosure-Elements.pdf

9. AAPOR Code of Professional Ethics and Practices, Section III.A, Items 5 and 8.

10. AAPOR Transparency Initiative Pledge, Section D. aapor.org/wp-content/uploads/2022/11/TI-Attachment-B.pdf

1. Plain-Language Definition

The breakoff rate is the percentage of respondents who start a survey but do not finish it. A respondent who answers the first several pages of questions and then closes the browser window, hangs up the phone, or otherwise stops participating before reaching the end has 'broken off.' This is distinct from “unit nonresponse,” where a person never starts the survey at all, and from “item nonresponse,” where a person skips individual questions but continues to the end.1,2

AAPOR's Standard Definitions distinguish between a 'partial complete' (a respondent who answers enough questions to be usable) and a 'breakoff,' which is classified as a type of refusal. The line between the two is defined by the researcher before data collection and depends on the survey's objectives and which questions are considered essential.3

Breakoff rates can be substantial. Meta-analyses of web surveys report median breakoff rates ranging from 16% to 34%. A major telephone panel study (the Panel Study of Income Dynamics) found that nearly 23% of all completed interviews required at least one session restart because the respondent had temporarily broken off.2,4

2. Why It Matters for Interpreting Results

Breakoff matters because the people who quit a survey partway through are often systematically different from those who finish it. Research shows that breakoff is predicted by respondent characteristics such as education level, age, and income. For example, respondents with less education are more likely to break off, as are those who find individual questions cognitively demanding. This means that excluding non-completers from the final dataset can introduce bias, making the remaining sample less representative of the target population.2

Breakoff is not random across the questionnaire either. It tends to spike at specific points: pages with many questions, open-ended response formats, questions requiring mental calculations or long recall periods, and section introduction pages that signal the start of a new block of content. Progress indicators that suggest the survey is longer than expected also increase breakoff risk.2

Even among respondents who do complete the survey, data quality tends to degrade as the questionnaire progresses. Questions asked later in a survey receive shorter response times, briefer open-ended answers, and less variability in grid-style responses, all of which suggest declining respondent effort. This means that breakoff patterns affect not only who is in the final sample but also the quality of responses from those who remain.5

Importantly, the relationship between response rates and actual nonresponse bias is weak. A survey with a low breakoff rate is not guaranteed to be unbiased, and a survey with a moderate breakoff rate is not necessarily flawed. What matters is whether the people who break off differ from completers on the variables being measured.6

3. What Effective Disclosure Looks Like

Effective disclosure reports both the number of people who started the survey and the number who completed it, allowing users to calculate the breakoff rate themselves. AAPOR's Transparency Initiative requires disclosure of sample sizes and, for probability surveys, disposition summaries showing how each sampled case was resolved (completed, partially completed, refused, not contacted, etc.). Good practice also specifies the threshold used to classify a case as a 'partial complete' versus a 'breakoff.'7,3

While less common, some well-documented surveys go further by reporting where in the questionnaire breakoffs were concentrated. Breakoff doesn't happen evenly across a questionnaire; it clusters at specific points rather than building up steadily as respondents move through the survey. ² Certain question types drove much higher dropout: open-ended questions (which require typing out an answer), long or multi-sentence questions, and pages announcing the start of a new section all showed substantially elevated breakoff. Grids of similar questions and unfamiliar response formats, such as slider scales, had the same effect. This matters for interpreting results: when breakoffs spike at a sensitive or demanding section, the respondents who remain may differ systematically from those who started, since people aren't leaving at random. They're leaving in response to specific features of the survey.

4. What Incomplete Disclosure Looks Like

The most common form of incomplete disclosure is simply not reporting breakoff rates at all. Researchers have noted that breakoff rates are 'not usually reported' in published survey results, making it a commonly omitted element. Without this information, users cannot assess how many respondents started but did not finish the survey, or whether the final sample may be biased by selective attrition.2,8

Another gap occurs when surveys report only the final sample size without clarifying whether partial completes are included or excluded. Some organizations count partially completed interviews in their reported N, while others exclude them entirely. Without knowing which convention was used, users may be comparing sample sizes that mean different things across studies.3

5. Strengths When Well-Disclosed

A well-disclosed breakoff rate provides useful diagnostic information about the survey. High completion rates suggest the questionnaire was well-designed and not overly burdensome. Breakoff point data can identify specific questions or sections where problems emerged, enabling targeted improvements in future survey waves.2,5

When breakoff rates are reported alongside respondent characteristics, users can assess whether completers are systematically different from non-completers. This is the foundation for evaluating whether the final results may be biased by who dropped out. Disclosure also helps distinguish between survey design problems (such as a confusing question format that causes breakoffs at a specific point) and broader population-level nonresponse issues.2

6. Limitations to Keep in Mind

A low breakoff rate does not guarantee unbiased results. If the people who break off are similar to those who complete the survey on the variables of interest, even a moderate breakoff rate may produce little bias. Conversely, even a small number of breakoffs can introduce meaningful bias if those who quit are systematically different on key measures.6

Breakoff patterns differ across survey modes. In web surveys, breakoff is typically permanent: once a respondent closes the browser, they rarely return. In telephone surveys, breakoffs are often temporary, with the respondent asking to be called back and the interviewer eventually completing the interview in a second session. This means that the same breakoff rate can have different implications depending on the mode of data collection.4

Breakoff rates are influenced by many factors simultaneously, including survey length, question difficulty, topic sensitivity, incentive structure, progress indicator design, and respondent characteristics. This makes it difficult to attribute a high breakoff rate to any single cause without additional analysis.2,5

7. Variations You May Encounter

Introduction breakoff

The respondent quits at the welcome or introduction page before answering any substantive questions. This is conceptually closer to unit nonresponse than to questionnaire breakoff, since the respondent has not yet engaged with any survey content.2

Questionnaire breakoff

The respondent answers at least one substantive question but quits before completing the survey. This is the most common use of the term 'breakoff' in survey research and is the type most likely to introduce bias in the final dataset.2

Temporary breakoff

The respondent stops the interview but later returns to complete it, either on their own (in a web survey with a saved-progress feature) or because an interviewer calls back (in a telephone survey). In the Panel Study of Income Dynamics, about two-thirds of interviews with a breakoff had exactly one breakoff event, and the vast majority were eventually completed through interviewer follow-up.4

Terminal breakoff

The respondent permanently abandons the survey and never returns. In web surveys, this is the default outcome of a breakoff, since there is no interviewer to call back. Terminal breakoffs are the primary concern for data quality because these respondents contribute no further data.8

8. Related Elements

Breakoff Rate is most closely related to Response Rate \, since breakoffs are a component of overall nonresponse. It also connects to Question Wording and  Data Collection Mode because longer surveys produce higher breakoff rates and different modes have different breakoff dynamics. Understanding the breakoff rate alongside the incentive structure is also useful, since incentives are one of the primary tools for keeping respondents engaged through the full survey.2,5,7

References

1. Peytchev, A. (2009). Survey Breakoff. Public Opinion Quarterly, 73(1), 74–97.

2. Peytchev, A. (2009). Survey Breakoff. Public Opinion Quarterly, 73(1), 74–97. (empirical results, pp. 85–93)

3. AAPOR (2023). Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th ed.

4. Sakshaug, J.W. & Crawford, S.D. (2013). Survey Breakoffs in a Computer-Assisted Telephone Interview. PMC.

5. Galesic, M. & Bosnjak, M. (2009). Effects of Questionnaire Length on Participation and Indicators of Response Quality. POQ, 73(2), 349–360.

6. Groves, R.M. (2006). Nonresponse Rates and Nonresponse Bias in Household Surveys. POQ, 70(5), 646–675.

7. AAPOR Transparency Initiative Disclosure Elements (revised April 2021). aapor.org

8. Do Question Topic and Placement Shape Survey Breakoff Rates? Survey Methods: Insights from the Field (2024).

1. Plain-Language Definition

Survey routers are online systems that intercept people while they are browsing the web, through banner ads, pop-ups, or other online placements (sometimes called "river sampling"), and screens them with a short set of demographic and behavioral questions. Based on their answers, the system then directs ('routes') them to whichever open survey they qualify for. A chain is a related practice in which a respondent who finishes one survey is immediately redirected to begin another, often without being told they are entering a different study.1,2

Both routers and chains differ fundamentally from traditional panel-based surveys, where a respondent is recruited into a panel, profiled over time, and then specifically invited to take a particular survey. In router-based sampling, the respondent may have no prior relationship with the research organization, may not know which study they are entering, and may be assigned to a survey based on whatever quotas happen to be open at that moment rather than on any predetermined sampling plan.1,2

AAPOR's Transparency Initiative requires that publicly released survey results describe the method used to generate and recruit the sample, including whether respondents came from a probability-based frame or from non-probability methods. Use of routers, chains, and river sampling fall squarely in the non-probability category and should be described as such.3

2. Why It Matters for Interpreting Results

Knowing whether a survey used routers or chains is important because these methods introduce several layers of unknown selection into the sample. First, there is exclusion bias: only people who happen to visit certain websites and click on certain ads are exposed to the survey invitation, effectively excluding much of any target population. Second, there is volunteer bias: only a small fraction of those exposed to the invitation choose to click through. Third, the routing process itself can introduce bias by assigning respondents to whichever survey needs to fill a demographic quota, rather than using a predetermined sampling design.1,4

Because router-sourced respondents often have no prior relationship with the survey organization, there is limited profiling information available to assess who they are or how representative they might be. Unlike panelists who complete detailed profile surveys upon joining, river-sampled respondents are screened with only a few questions before being routed. This makes it difficult to evaluate or adjust for potential biases in the sample.1,2

Research comparing surveys from multiple online panels and river samples has consistently found differences in results. The Advertising Research Foundation found a 40% respondent duplication rate across 17 different panel companies, and respondents recruited through different methods can produce different estimates even when answering the same questions.2

3. What Effective Disclosure Looks Like

Effective disclosure states explicitly whether the sample was sourced through a router, river sampling, or chain, and distinguishes this from direct panel invitations. It should describe where the respondents were recruited from (which types of websites or ad networks), how respondents were screened and routed, and whether any respondents were chained from a prior survey. The name of the sample provider or router operator should be identified.3,5

In many cases, the specific websites or placements behind a router-sourced sample are not fully knowable, since the advertising is often placed through an automated platform such as Google Ads, with the individual sites determined by an algorithm rather than chosen directly by the survey organization. Where that is the case, a realistic disclosure may only be able to describe broad categories, for example “ads placed through an automated ad network on websites with traffic above a stated threshold, targeted at users the platform's algorithm classified as likely news consumers.” Even a description at this level of detail is more specific than what most disclosures provide.

AAPOR's Transparency Initiative checklist calls for disclosing whether the sample comes from a probability-based methodology or non-probability methods, the name of the supplier, and a description of how participants were contacted, recruited, or intercepted. For router-based samples, this means describing the routing mechanism and the screening criteria used to qualify respondents for the specific study.3

Well-documented surveys also report what percentage of the final sample came from router or river sources versus direct panel invitations, since many studies blend samples from multiple sources. When multiple sources are used, the methodology should describe how the sources were combined and whether the data were analyzed separately or pooled.1,5

4. What Incomplete Disclosure Looks Like

A common gap is describing the sample as coming from 'an online panel' when the respondents were actually recruited through a router or river sample and never belonged to a panel in any meaningful sense. If an actual panel is used, the identity of that panel should be disclosed to prevent such confusion. Mislabeling river sampling with routers as panel sampling obscures the recruitment mechanism and prevents users from evaluating the sampling approach.1

Another form of incomplete disclosure is failing to mention that respondents were chained from a prior survey. Again, chaining means that a respondent who just finished answering questions about one topic is immediately redirected to a new survey on a different topic, potentially introducing fatigue, satisficing, or carryover effects that would not be present in a freshly recruited respondent.1,2

Omitting the sample provider's name is also a concern. Without knowing which company operated the router, users cannot evaluate the provider's track record, the quality of its traffic sources, or whether its respondents have been used in comparable studies.3

5. Strengths When Well-Disclosed

Clear disclosure of router or chain use enables users to appropriately calibrate their confidence in the results. Router-based samples can be useful for some research purposes, particularly when the goal is to reach specific subpopulations quickly, test advertising concepts, or conduct experiments where internal validity matters more than population representativeness.1

Transparency about the sample source also allows data users to apply appropriate analytical methods. If a survey discloses that it used river sampling, analysts know statistically grounding inferences about the target population cannot be made and can consider whether additional adjustment methods might be warranted.1,4

When router use is well-disclosed, it supports comparability across studies. Researchers evaluating a body of evidence can distinguish between polls that used probability-based recruitment and those that used router-based convenience samples.1

6. Limitations to Keep in Mind

Router-based samples lack a defined sampling frame, which means there is no way to calculate selection probabilities or traditional response rates. AAPOR's Standard Definitions note that outcome rates designed for probability-based surveys are not appropriate for non-probability samples, including those recruited through routers. Instead, participation rates (the proportion of those invited who completed the survey) should be reported, though even these can be difficult to determine for river samples where the denominator is often unknown.4,6

Post-survey weighting and adjustment techniques can reduce but generally cannot eliminate the biases inherent in router-based samples. Research evaluating the effectiveness of these adjustments consistently finds that they offer only a partial remedy, and there is no statistical theory that guarantees the resulting estimates will be unbiased.1

The quality of router-sourced respondents can vary dramatically depending on the traffic sources used, the screening criteria applied, and whether respondents were chained from other surveys. Two studies that both describe their sample as 'river-sourced' may have drawn from very different respondent pools.1,2

Respondent duplication is a particular concern when samples are drawn from multiple panels or router sources. The same person may be a member of several panels and may encounter the same survey through multiple channels, creating the risk of being counted more than once.2

7. Variations You May Encounter

Direct panel invitation

A panelist is specifically selected from a pre-recruited panel and invited to complete a particular survey. This is the traditional online panel survey model and provides the most control over who is invited and who responds.2

River sampling (real-time intercept)

Respondents are recruited in real time from general web traffic through banner ads, pop-ups, or other online placements. They are typically not pre-profiled and are screened on the spot before being directed to a survey.1,2

Router-based assignment

A system that screens respondents and assigns them to whichever open survey they qualify for, based on their demographic profile and which quotas need to be filled. The respondent typically does not choose which survey to take.1

Chaining (post-survey redirect)

After completing one survey, the respondent is immediately redirected to begin a second survey, often without a clear indication that they are entering a different study. This raises concerns about respondent fatigue and carryover effects.1,2

Blended samples

Many commercial surveys combine respondents from multiple sources, including direct panel invitations, router-recruited respondents, and sometimes chained respondents, in a single dataset. The methodology should describe how these sources were combined.1

8. Related Elements

Routers and Chains is most closely related to External Sample Provider (Element 11), since the sample provider determines whether panel, router, or blended recruitment is used. It also connects to Sample Design (Element 5) and Data Collection Mode (Element 6), since router-based recruitment represents a fundamentally different sampling approach than probability-based designs. Understanding whether routers were used also helps contextualize the Incentive Structure (Elements 13 and 14), as river-sampled respondents may receive different incentives than panelists, and the Response Rate (Element 28), since traditional response rate calculations do not apply to non-probability router samples.1,3

References

1. Baker, R. et al. (2013). Report of the AAPOR Task Force on Non-Probability Sampling. JSSAM, 1, 90–143.

2. Baker, R. et al. (2010). AAPOR Report on Online Panels. Public Opinion Quarterly, 74(4), 711–781.

3. AAPOR Transparency Initiative Disclosure Elements (revised April 2021). aapor.org

4. AAPOR (2023). Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th ed.

5. ECU Center for Survey Research. Checklist for AAPOR Transparency Initiative.

6. AAPOR (2022). Best Practices for Survey Research. aapor.org/standards-and-ethics/best-practices/

1. Plain-Language Definition

The noncovered population is the group of people who belong to the target population but have no chance of being selected into the survey because the sampling frame does not include them. The estimated size of this group tells you how many people are systematically left out. For example, a telephone survey that dials only landline numbers has no way of reaching people who live in cell-phone-only households. Those households are the noncovered population for that survey.1

Knowing the estimated size of the noncovered population matters because it directly affects how well the survey results represent the full population of interest. If the noncovered group is small and similar to the covered group on the topics being measured, the risk of bias is low. If the noncovered group is large or systematically different, the survey results may not generalize to the full population. AAPOR's Transparency Initiative requires that organizations disclose any segment of the target population not covered by the survey design.2

2. Why It Matters for Interpreting Results

Coverage error is one of the four major sources of survey error, alongside sampling error, nonresponse error, and measurement error. Unlike nonresponse, where at least some effort was made to reach the missing people, noncoverage means certain groups were never given a chance to participate.3

The practical importance of noncoverage depends on how different the excluded group is from the included group on the variables being measured. In a telephone survey conducted only on landlines, young adults, renters, and lower-income households are disproportionately excluded because they are more likely to live in cell-phone-only households. Research has documented that these excluded groups hold systematically different views on political, social, and health topics.4,5

Online surveys face a related but distinct coverage problem. People without internet access or without the digital literacy to participate in web-based research are excluded. Although internet understanding has increased substantially, the remaining non-internet population is not randomly distributed. It is concentrated among older adults, lower-income households, rural residents, and people with less formal education.6

Users who ignore this element risk assuming that survey results apply to the full target population when they actually apply only to the portion of the population reachable through the survey's chosen method.1

3. What Effective Disclosure Looks Like

Effective disclosure provides a numerical estimate (or at least a range) of the share of the target population not covered by the survey design. It identifies which groups are excluded and explains why. For example, a report might state that the survey covered approximately 97% of U.S. households through an address-based sampling frame, with the remaining 3% consisting primarily of households not on the USPS Delivery Sequence File, including some rural households and people experiencing homelessness.1,2

Well-documented records also describe whether any steps were taken to reduce the coverage gap. For example, NORC's AmeriSpeak panel supplements its address-based frame with in-person listing of rural households not recorded on postal files, and offers a telephone mode for panelists without internet access.7

For telephone surveys, effective disclosure specifies whether the sample included cell phones, landlines, or both, and cites an external estimate of the cell-phone-only rate for the target population.4

4. What Incomplete Disclosure Looks Like

Incomplete disclosure describes the sampling frame without acknowledging who is left out. A report that says "the survey used an address-based sample" without noting that people experiencing homelessness, those in institutional settings, and some rural households are excluded leaves users unable to assess coverage.1

Another common gap is describing the survey as "nationally representative" without specifying the frame or its coverage rate. This claim may be technically unsupported if significant portions of the population had no chance of being selected.2

For online surveys, claiming representativeness without noting the exclusion of non-internet households is a significant disclosure failure. Even if the non-internet share of the population is small, it is not randomly distributed, and its exclusion introduces systematic coverage bias.6

5. Strengths When Well-Disclosed

When the noncovered population is well-documented, users can make informed judgments about whether the survey results are likely to apply to the groups they care about. For policy research, this is especially important because the groups most likely to be excluded from surveys (lower-income populations, people with disabilities, people without stable housing) are often the groups most affected by the policies being studied.1,3

Quantifying the noncovered population also allows users to assess how much adjustment would be needed to extend the results to the full target population. If the noncovered group is 2% and unlikely to differ sharply from the covered group, the risk is minimal. If it is 15% and concentrated in a demographic that holds distinct views, the risk is substantial.4,5

Clear coverage disclosure supports comparisons across surveys. Users can evaluate whether differences in results between two surveys might be attributable to differences in their frames rather than genuine differences in public opinion.2

6. Limitations to Keep in Mind

Estimating the noncovered population requires external data sources, such as census figures or administrative records, which may themselves have coverage limitations. The precision of the estimate depends on the quality of those external sources.3

Coverage rates change over time. The share of U.S. households that are cell-phone-only has increased from under 5% in 2004 to over 70% by the mid-2020s. Internet penetration has followed a similar upward trajectory. A coverage rate reported for one survey may not apply to a survey conducted even a few years later.4,6

Noncoverage is only one source of survey error. A survey with excellent coverage can still produce biased results due to nonresponse, measurement error, or weighting failures. The noncoverage estimate is most useful when considered alongside other methodological disclosure elements.3

7. Variations That Exist

The specific noncovered populations vary by survey mode and frame. Landline-only telephone surveys exclude cell-phone-only households. Address-based samples exclude people without a fixed residential address. Online surveys exclude people without internet access. Surveys conducted only in English exclude non-English speakers.1,4

Some surveys attempt dual-frame or multi-mode designs to reduce noncoverage. For example, dual-frame random-digit-dial surveys sample from both landline and cell phone frames. Mixed-mode surveys offer web and telephone options. These approaches reduce but do not eliminate coverage gaps.7

Coverage bias has been documented across several specific populations. Landline-only surveys significantly underrepresent young adults, renters, racial and ethnic minorities, and lower-income households. Online-only surveys underrepresent older adults without internet access and rural populations. The direction and magnitude of coverage bias varies by survey topic.4,5,6

8. Related Elements

Noncovered population connects most directly to Sampling Frame, since the frame determines who can and cannot be reached. It also relates to Sample Design, since design choices such as dual-frame sampling are intended to reduce coverage gaps. Weighting is relevant because weighting adjustments sometimes attempt to compensate for known coverage deficiencies, though such adjustments rely on assumptions about the noncovered group that may or may not hold.1,2

References

1. AAPOR Code of Professional Ethics and Practices, April 2021, Section III.A. aapor.org/standards-and-ethics/disclosure-standards/

2. AAPOR Transparency Initiative. aapor.org/standards-and-ethics/transparency-initiative/

3. Groves, R.M. et al. (2009). Survey Methodology, 2nd edition. Wiley.

4. Blumberg, S.J. & Luke, J.V. Wireless Substitution: Early Release of Estimates from the National Health Interview Survey. cdc.gov/nchs/nhis

5. Keeter, S. et al. (2007). What's Missing from National Landline RDD Surveys? Public Opinion Quarterly, 71(5), 772-792.

6. Couper, M.P. (2000). Web Surveys: A Review of Issues and Approaches. Public Opinion Quarterly, 64(4), 464-494.

7. NORC AmeriSpeak Technical Overview. no

1. Plain-Language Definition

Survey incentives are rewards offered to potential respondents to encourage them to participate. These can take many forms: cash enclosed with a mailed questionnaire, a gift card sent after completing a web survey, entry into a prize drawing, points accumulated in an online panel system, or non-monetary tokens such as pens or charitable donations made on the respondent's behalf. Incentives vary along two key dimensions: form (monetary versus non-monetary) and timing (prepaid, meaning sent before or with the survey request, versus conditional, meaning promised upon completion).1,2

The “use of incentives” element asks whether any incentive was used at all. The “what incentive was provided” element asks what specific incentive was provided, including its type, amount, and delivery method. Together, these two elements describe the incentive structure of the survey. AAPOR's Transparency Initiative requires that publicly released survey findings disclose details about compensation and incentives provided to respondents, including the method of delivery.3

2. Why It Matters for Interpreting Results

Incentives matter because they influence who participates in a survey, which in turn affects what the results look like. Meta-analytic research has found that incentives increase response rates, with prepaid monetary incentives tending to produce the largest effects. This evidence draws heavily on studies conducted before the shift to online and mixed-mode data collection, so the size of the effect in current survey environments is less certain.1,2

But higher response rates are not the only consideration. Incentives can also change the composition of the sample. Research shows that incentivized respondents may differ demographically from those who participate without incentives: they tend to be younger, more racially diverse, and less engaged with the survey topic. One study found that a $5 incentive attracted respondents who were less interested in the survey subject and more likely to select 'don't know' on background questions, while non-incentivized respondents read more of the material and expressed stronger attitudes about the topic.4,5

This does not necessarily mean the non-incentivized sample is more representative of the general population. Topic interest itself predicts who participates in a survey at all, so a sample skewed toward people who care most about a subject can also diverge from the population, since most people are not deeply engaged with any given topic. Disclosure of incentive use is useful precisely because it helps a reader weigh this trade-off rather than assume one sample is simply better than the other.6

This matters because if incentives bring different kinds of people into the sample, the survey's overall results may shift depending on whether and how incentives were used. A user who does not know the incentive structure cannot assess whether the sample might overrepresent or underrepresent certain groups.1,5

At the same time, most research finds that incentives do not degrade the quality of individual responses. Studies examining item nonresponse, straight lining, response time, and length of open-ended answers generally find no significant differences between incentivized and non-incentivized respondents.4,5

3. What Effective Disclosure Looks Like

Effective disclosure states clearly whether incentives were used and provides enough detail for a reader to evaluate their potential effects. According to AAPOR's Transparency Initiative, this includes the type of incentive (cash, gift card, points, lottery entry), the amount or value, whether it was prepaid or conditional on completion, and the method of delivery (debit card, gift card, cash). For surveys drawn from online panels, effective disclosure also describes the panel's incentive structure, including how points are earned and what they convert to in real-world value.3

In some cases, the incentive is not uniform across respondents. Organizations may offer a larger incentive to reach groups that are harder to recruit or may let respondents choose among incentive options. Effective disclosure covers this by describing the range rather than a single figure, for example 'gift cards worth between $5 and $30' or 'respondent choice of points or gift cards with values from $1 to $5.7

Well-documented surveys typically include this information in the methodology section, written in plain terms: for example, 'Respondents received a $5 electronic gift card sent via email within 10 days of completion' or 'Panel members earned 200 points (approximately $2 equivalent) for completing this survey.'3,5

4. What Incomplete Disclosure Looks Like

Common gaps include stating that 'respondents were compensated' without specifying the type or amount, or omitting incentive information entirely. For online panel surveys, a frequent problem is that respondents earn 'points' in a rewards system, but the conversion rate to real monetary value is rarely disclosed to the survey consumer. Without knowing the conversion rate, a reader who sees that panel members were offered 'points' cannot evaluate whether this represents a trivial or substantial incentive.8

Another gap occurs when surveys fail to distinguish between prepaid and conditional incentives, even though research consistently shows these operate through different psychological mechanisms and produce different effects on who responds. The same dollar amount can also represent a high or low incentive depending on the length and burden of the survey, so disclosure without context about survey length leaves users unable to calibrate the incentive's likely effect.1,2

5. Strengths When Well-Disclosed

When incentive use is clearly documented, users can evaluate potential selection effects and assess who the incentive may have drawn into or kept out of the sample. Meta-analytic evidence spanning decades confirms that incentives reliably increase response rates across virtually all survey modes, including mail, telephone, face-to-face, and web surveys, and in both cross-sectional and longitudinal designs. Prepaid monetary incentives are consistently the most effective.1,2

Specific disclosure also allows researchers to compare incentive structures across studies. If two polls on the same topic produce different results, knowing that one used a $20 prepaid cash incentive and the other relied on panel points may help explain potential differences (e.g., in sample composition across surveys).1,2

Incentives can also improve the representativeness of a sample when they bring in hard-to-reach populations who would otherwise not participate. Several studies show that monetary incentives are especially effective at recruiting low-income, minority, and younger respondents into surveys where they would otherwise be underrepresented.1

6. Limitations to Keep in Mind

Monetary incentives may also change the composition of the sample by drawing in respondents who would not otherwise have participated. Leverage-salience theory predicts that incentives are particularly effective at recruiting people who lack other reasons to participate, such as interest in the topic.1 As a result, an incentivized sample may include a somewhat different mix of respondents by topic interest than a non-incentivized one, a separate consideration from the topic-driven skew discussed above.

Lottery or prize-draw incentives may attract different types of respondents than guaranteed payments, and research generally shows that lotteries are less effective at boosting response rates than even small, guaranteed incentives.1

For online panel surveys, the incentive structure creates a pool of effectively paid survey-takers. Panelists who complete many surveys for points or cash are not a random cross-section of the population, and their experience with surveys may affect how they respond. This is important context for interpreting any panel-based survey data.8

The effect of incentives can also vary over time. While there is currently no strong evidence that offering incentives conditions the public to expect payment for survey participation, the possibility remains a concern among researchers as incentive use becomes more widespread.1

7. Variations You May Encounter

No incentive

Some surveys rely entirely on altruistic motivation, topic interest, or civic obligation. Government statistical surveys, for example, often do not offer monetary incentives but may carry a legal mandate to respond.1

Prepaid monetary incentives

Cash or gift cards sent with or before the survey invitation. Research consistently shows these produce the largest response rate increases because they invoke a norm of reciprocity: the respondent feels obligated to respond after receiving something of value.1,2

Conditional monetary incentives

Payment promised upon completion. These are less effective than prepaid incentives at increasing response rates, and some research finds they produce no statistically significant improvement over no incentive at all.2

Lottery or prize drawing

Entry into a drawing for a larger prize. Research generally finds these are no more effective than no incentive, because the expected value per respondent is very small and the reward is promised rather than prepaid.1

Panel points or rewards

Online panel members accumulate points redeemable for cash, gift cards, merchandise, or airline miles. The real-world value of these points varies widely across panels and is often opaque to survey consumers.8

Non-monetary incentives

Items such as pens, key rings, charitable donations, or a summary of results. These produce modest but significant response rate improvements when sent with the initial survey mailing.2

8. Related Elements

Incentives are most closely related to the External Sample Provider, since online panels have their own built-in incentive structures that interact with any study-specific incentive. They also connect directly to Response Rate, because incentives are one of the primary tools researchers use to increase participation. Understanding the incentive structure helps users interpret the response rate in context: a high response rate achieved through large incentives may reflect a different sample composition than a high response rate achieved through topic salience or repeated contact attempts.1,3

Because incentives can also draw professional or high-frequency survey-takers, incentive disclosure connects to Quality Control as well, since checks such as attention and logic checks, straight lining detection, and speeding checks are among the main tools used to identify and remove low-effort respondents drawn in mainly by the reward.

References

1. Singer, E. & Ye, C. (2013). The Use and Effects of Incentives in Surveys. The ANNALS of the AAPSS, 645(1), 112–141.

2. Church, A.H. (1993). Estimating the Effect of Incentives on Mail Survey Response Rates: A Meta-Analysis. Public Opinion Quarterly, 57(1), 62–79.

3. AAPOR Transparency Initiative Disclosure Elements (revised April 2021). aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

4. Incentive Impact on Data Quality, Sample Composition, and Respondents' Topic Interest. Survey Practice (2025).

5. AAPOR (2022). Best Practices for Survey Research. aapor.org/standards-and-ethics/best-practices/

6. Groves, R.M., Presser, S., & Dipko, S. (2004). The Role of Topic Interest in Survey Participation Decisions. Public Opinion Quarterly, 68(1), 2–31.

7. Pew Research Center. 2023. "Methodology: Diverse Workplaces Are Divided Over Diversity, Equity and Inclusion Efforts." Washington, DC: Pew Research Center. https://www.pewresearch.org/social-trends/2023/05/17/culture-of-work-dei-methodology/

8. Baker, R. et al. (2010). AAPOR Report on Online Panels. Public Opinion Quarterly, 74(4), 711–781.

1. Plain-Language Definition

Quality assurance and quality control are related but distinct parts of a survey's quality system. Quality assurance is process-oriented. It covers the inputs and procedures used to produce the data, such as how the sample was drawn, how the questionnaire was tested, and how interviewers or data collection systems were prepared. Quality control is product-oriented. It covers the checks applied to the data and research claims themselves, after collection, to confirm that responses are usable and that reported findings are supported by the data.

Quality control refers to the checks a survey team performs on responses to identify and handle low-effort or invalid answers or respondents before the data are analyzed. Common checks include watching for respondents who finish unusually fast (speeding), give the same answer down a long list of questions (straight lining), fail attention or logic checks, or appear to be automated (bots) or duplicate entries. Some in-person studies also re-contact a subset of respondents to confirm the interview took place.1

These checks exist because respondents do not always answer carefully or are not always who they claim to be. On the interviewer side, falsification is often traced to identifiable pressures. Interviewers asked to work in regions where they fear for their personal safety may fabricate part or all of an interview rather than conduct it as instructed, a longstanding practice known as curbstoning.13 On the respondent side, online surveys that offer cash or other incentives can draw bots and professional survey takers who complete surveys simply to collect payment. One research team fielding a public health survey saw their eligibility screener overwhelmed by suspicious submissions within an hour of launch and had to redesign the survey to keep fraudulent responses out.14 Quality control checks catch some of this after the fact, but these underlying pressures are also why quality assurance procedures, such as screening and identity verification, matter before data collection begins.

The survey methods literature describes low-effort answering as satisficing, giving a good-enough, low-effort answer rather than working through a question fully.3 Quality control should include the set of procedures designed to detect that behavior, and current reviews organize the specific indicators used to flag it.4

2. Why It Matters for Interpreting Results

Quality assurance and quality control together affect how much of a dataset reflects careful, genuine responses. Low-effort answering degrades measurement. One widely cited study flagged roughly 10 to 12 percent of respondents as careless in a lengthy survey.6 The underlying reason is well established. When a question is demanding, some respondents’ shortcut the mental work rather than fully engage with the survey.5

Whether and how a survey screens for this shapes what its numbers can support. AAPOR's quality framework places quality control within measurement error, since careless responses that are never identified remain in the results and can affect estimates.9 Knowing what procedures were used at each stage, from sample sourcing through post-collection checks, helps a reader judge how much error was likely addressed.

3. What Effective Disclosure Looks Like

Quality assurance disclosure

Quality assurance procedures differ by research method and mode. For probability-based surveys, disclosure includes how the sample frame was sourced and how the sample was drawn.1 It also includes how the population parameters used in weighting were sourced and applied, and how respondents were contacted, selected, or screened for eligibility.1 Where interviewers are used, it includes their training and supervision protocols.2 It also includes interview validation procedures such as re-contacting respondents to confirm the interview took place or to verify their identity.1

This includes disclosure of efforts to identify infiltration of nonprobability panels by AI-generated or bot-driven responses, since incentive payments can draw both automated bots and ineligible respondents attempting to pass as qualified participants.15

In both cases, quality assurance also includes verification of appropriate questionnaire design through cognitive testing, as well as disclosure of scripts and instructions given to interviewers and respondents.1,12

Quality control disclosure

Effective quality control disclosure names the specific checks that were used rather than referring generally to cleaning. Disclosure states the particular procedures used, such as attention and logic checks, the straight lining rule, the time threshold for speeding, duplicate prevention, and any re-contacts, along with any software used.1 Methodological guidance recommends a multiple-hurdle approach that specifies which indices and thresholds were applied, so a reader can see which checks were run and at what cutoffs.7

Quality control disclosure also includes dataset checks for attention, speeding, and straight lining; sample performance measures such as proximity to benchmarks and design effects; and appropriate use of significance testing.10

Because quality control connects directly to the share of respondents removed, thorough disclosure also reports that removal rate and notes whether any automated or AI tools were part of the process..

4. What Incomplete Disclosure Looks Like

A statement that quality assurance and quality control were performed, with no further detail, conveys little on its own, because practices vary widely and the phrase does not indicate what was actually done.1 AAPOR's disclosure elements can serve as a standard here. Descriptions vaguer than that level of specificity, such as data were cleaned, leave out the information a reader needs.1

The research literature explains why the detail matters. Studies define and apply careless-response criteria inconsistently, so without named methods the cleaning cannot be reproduced.6 Different indices catch different problems, so a bare reference to quality control does not reveal which kinds of error were addressed and which were not.7 The same is true on the quality assurance side. A report that states a sample was screened for bots, with no description of the method used, leaves a reader unable to judge how much confidence to place in the sample's composition.

5. Strengths When Well-Disclosed

Documented quality assurance and quality control procedures let a reader see how the data were produced and screened. Careless responding takes more than one form, since random and non-random patterns call for different indices, so combining several quality control checks tends to catch more than any single one.6 Re-interview validation, in which a subset of respondents is re-contacted to confirm the interview occurred, is a particularly direct check, and AAPOR's disclosure elements treat such re-contacts as a documented practice.1

On the quality assurance side, disclosed interviewer training and supervision protocols let a reader assess how in-person or telephone data collection was managed. Disclosed panel vetting procedures let a reader assess how an online sample was protected from bot and fabricated-identity infiltration before data collection began.

More broadly, current best-practice reviews present well-documented and pre-specified quality procedures as an indicator of methodological care, because they make the process transparent and repeatable rather than left to unstated judgment.4

6. Limitations to Keep in Mind

Quality assurance and quality control choices involve judgment, and those choices have consequences for who is included in or removed from a study. There are no universal cutoffs for quality control checks, so where a threshold is set directly affects who gets flagged and how many false positives occur.7 A blunt speeding cutoff illustrates the trade-off. Speeding is concentrated among younger and less-educated respondents, so an aggressive threshold can remove them at higher rates. The same cutoff can also catch legitimate fast responders, such as people already familiar with a topic.8

Automated and AI-assisted screening add another consideration, since algorithmic decisions may be applied consistently but are not always transparent or easy to audit.11 The absence of quality control disclosure leaves a reader unable to tell whether the data were screened at all. The absence of quality assurance disclosure leaves a reader unable to tell how participants were recruited and verified in the first place. AAPOR's guidance on online samples also notes that cleaning can be taken further than necessary and that removal practices are not always comparable across firms, which use different rules.10

7. Types of Quality Assurance and Quality Control Procedures You May Encounter

Quality assurance procedures

Interviewer training and supervision. Where interviewers are used, describing training protocols and supervision practices.2

Interview validation. Re-contacting a sample of respondents to confirm an interview occurred as reported.1

Panel infiltration safeguards. For nonprobability surveys, documenting efforts to identify bots, AI agents, or other false actors attempting to enter or complete surveys within a panel.16

Additionally, central to quality assurance are the many choices involved in the survey design: how the sample frame was sourced, how the sample was drawn, how population parameters used in weighting were sourced and applied, and household and/or respondent screening and selection. See individual disclosure elements for more information about how these affect data quality.

Quality control procedures

Speeding checks. Flagging responses completed faster than a reading-speed-based threshold, question by question, as an indicator of low engagement.8

Straight lining detection. Identifying identical answers given straight down a battery of items.1

Attention and logic checks. Embedded questions or internal-consistency tests used to confirm engagement.1

Duplicate response prevention. Measures used to prevent respondents from completing the survey more than once.1

Re-interview validation. Re-contacting respondents to confirm the interview occurred or to verify their identity.1

AI-assisted flagging. Automated tools that score or flag responses for review.11

Sample performance checks. Comparing the sample against known benchmarks and reviewing design effects from weighting or clustering.10

Significance testing. Applying statistical tests, where appropriate, before reporting a difference as meaningful.

Fact-checking of research claims. Having analysts other than those who produced a finding review it against the underlying data before publication.

No quality control performed. Some studies report no checks, which is itself information for the reader.

Historical disclosure

Quality assurance and quality control both predate the internet, though the emphasis has shifted over time. Self-administered web surveys shifted quality control toward speeding, straight lining, attention checks, and bot or duplicate detection.1 Quality assurance in this environment has also come to include panel vetting and identity verification, with AI-assisted flagging emerging as a newer layer on both sides.10,11 When comparing older telephone-era and newer online studies, expect the vocabulary of quality assurance and quality control to differ.

8. Related Elements

Quality control connects most directly to the share of respondents removed due to checks, which is essentially its numeric output. Quality assurance connects most directly to the sampling procedure and the mode of data collection, since it documents how the sample was built and how interviewers or systems were prepared to collect it. Both relate to the breakoff rate as well. AAPOR's disclosure elements pair data processing with exclusions and any imputation, and AAPOR's quality framework places quality control between measurement error on one side and nonresponse and breakoff on the other.9

References

1. AAPOR Transparency Initiative, Disclosure Elements (revised April 2021), covering measurement tools and scripts (Item 3), sample frame and recruitment method (Item 5), and how data were weighted (Item 9) and processed and checked for quality (Item 10). aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. AAPOR Terms and Conditions for Transparency Certification (revised April 2021), Section B, Item 2 (interviewer and coder training). aapor.org/wp-content/uploads/2022/11/TI-Attachment-A.pdf

3. Anand, S. (2008). Satisficing. In Encyclopedia of Survey Research Methods (Vol. 0, pp. 798–799). Sage Publications. methods.sagepub.com

4. Ward, M.K. & Meade, A.W. (2023). Dealing with Careless Responding in Survey Data: Prevention, Identification, and Recommended Best Practices. Annual Review of Psychology, 74, 577–596.

5. Krosnick, J.A. (1991). Response Strategies for Coping with the Cognitive Demands of Attitude Measures in Surveys. Applied Cognitive Psychology, 5, 213–236.

6. Meade, A.W. & Craig, S.B. (2012). Identifying Careless Responses in Survey Data. Psychological Methods, 17(3), 437–455.

7. Curran, P.G. (2016). Methods for the Detection of Carelessly Invalid Responses in Survey Data. Journal of Experimental Social Psychology, 66, 4–19.

8. Zhang, C. & Conrad, F. (2014). Speeding in Web Surveys: The Tendency to Answer Very Fast and Its Association with Straightlining. Survey Research Methods, 8(2), 127–135.

9. AAPOR (2016). Evaluating Survey Quality in Today's Complex Environment. aapor.org

10. AAPOR (2023). Data Quality Metrics for Online Samples: Task Force Report. aapor.org

11. AAPOR (2026). Responsible AI Integration in Survey Research, §3.2.2.

12.

12. Beatty, P.C. & Willis, G.B. (2007). Research Synthesis: The Practice of Cognitive Interviewing. Public Opinion Quarterly, 71(2), 287–311.

13.AAPOR Data Falsification Task Force. Falsification in Surveys. American Association for Public Opinion Research. https://www.aapor.org/wp-content/uploads/2022/11/AAPOR_Data_Falsification_Task_Force_Report-updated.pdf

14. Wang, J., Calderon, G., Hager, E.R., Edwards, L.V., Berry, A.A., Liu, Y., Dinh, J., Summers, A.C., Connor, K.A., Collins, M.E., Prichett, L., Marshall, B.R., & Johnson, S.B. (2023). Identifying and preventing fraudulent responses in online public health surveys: Lessons learned during the COVID-19 pandemic. PLOS Global Public Health, 3(8), e0001452. https://doi.org/10.1371/journal.pgph.0001452

15. Caven, I., Yang, Z., & Okrainec, K. (2025). It's Raining Bots: How Easier Access to Internet Surveys Has Created the Perfect Storm. BMJ Open Quality, 14, e003208.

1. Plain-Language Definition

This element is the share of respondents dropped from a survey before analysis because they did not pass quality-control checks. It is, in effect, the numeric result of the quality-control process: once a survey applies its rules for speeding, straight lining, failed attention checks, duplicates, or automated responses, the removal rate summarizes how many cases those rules excluded.1 In practice, detection and removal work as a pipeline: identify questionable cases, remove those that meet the removal criteria, and report the percentage of total cases that were removed due to those standards.2

In AAPOR's disclosure scheme, the figure sits within the item on data processing and exclusions, next to the sample and weighting information.1 The Roper Center's Transparency Project provides a specific place in each study disclosure scorecard where a data-processing or removal figure would appear.8

2. Why It Matters for Interpreting Results

The removal rate tells a reader how much the achieved sample was reshaped during cleaning. Because respondents are not removed at random, dropping them changes the composition of who remains, and the removal rate indicates how much of that reshaping took place.4 A very low or very high figure can each be informative: as a rough point of reference, one widely cited study flagged around 10–12% of respondents as careless in a long survey, so a rate far from that range may be worth a closer look.3

The figure also connects to who ends up in the final dataset. AAPOR's quality framework notes that exclusions change the composition of the sample, which in turn bears on measurement and on who is represented in the results.7 Read together with the rules that produced it, the removal rate helps a reader understand how much cleaning occurred and what it may have done to the resulting sample.

3. What Effective Disclosure Looks Like

Effective disclosure reports both the removal rate and the checks that produced it. AAPOR's disclosure elements describe stating what was excluded, why, and how much: the rule together with the count or percentage and noting any imputation or replacement of removed cases.1 Methodological guidance adds that reporting the specific indices and thresholds, and how many cases each was removed, makes the total traceable to particular checks rather than presented as a single unexplained number.2

A fully populated data-processing field, such as those in iPoll transparency scores, offers a concrete model of how removal information reads when it is disclosed clearly.8 The common thread is that the percentage and the method are reported together, so the figure can be understood in context.

4. What Incomplete Disclosure Looks Like

Incomplete disclosure gives one half of the picture without the other: a percentage with no description of the checks behind it, or a description of checks with no percentage, are each of limited use.1 Either on its own is difficult to interpret, since the rate and the rule are meaningful mainly together. When removal criteria are undocumented, the resulting rate is difficult to verify, particularly when the underlying data are not available for review.3

There is also the matter of when the removal rules were set. Current best-practice reviews note that exclusion rules decided or adjusted after seeing the data introduce what researchers call "researcher degrees of freedom," because the rules can be tuned in ways that affect the result.4 

5. Strengths When Well-Disclosed

A disclosed removal rate tied to named checks tells a reader that the data were actively screened and gives a sense of whether cleaning was light or extensive. It also makes the process auditable and repeatable, so the cleaning can be examined rather than taken on trust.1 Evidence that removing careless cases can recover genuine relationships in the data supports the value of screening, and a transparent rate signals that the data went through that step.2

Pairing the removal figure with the specific detection methods that produced it is what makes the rate interpretable, since the number gains meaning from the rules that generated it.3

6. Limitations to Keep in Mind

The central insight for this element is that removal is not neutral with respect to who stays in the sample. Because dropping cases is non-random, a higher removal rate can shift the composition of the achieved sample even as it removes noise, which can affect how far the results generalize.4 Particular rules can systematically drop particular groups, resulting in a less representative sample: speeding-based removal, for instance, tends to affect younger and less-educated respondents more, so a speeding rule can quietly degrade the sample’s quality.5

Removal rates are also difficult to compare across surveys, because firms use different rules and thresholds, and cleaning can be taken further than necessary.6 Automated or AI-assisted removals add a further consideration, since such removals may not be disclosed or explained in the same way as manual checks.9 For all these reasons, the rate is most informative when read alongside the rule that produced it rather than on its own.

7. Reasons Respondents Are Removed That You May Encounter

Different triggers, and different ways of combining them, produce different removal rates, which is part of why the rate is only interpretable alongside the rule.1

Attention-check failure. Removing respondents who miss embedded checks meant to confirm engagement.1

Speeding. Dropping respondents who finish faster than a set time threshold.5

Straight lining. Excluding respondents who give identical answers down a battery of items.1

Duplicate or device flags. Removing repeated entries identified by IP address or device.1

Low-quality open-ends. Excluding gibberish or off-topic open-text responses.1

Bot or AI-generated-response detection. Screening out entries that appear automated.1

Single versus multiple-hurdle rules. A removal can rest on one check or on several combined, and removal before analysis differs from flagging cases for later sensitivity checks.2

Historical context: A removal rate from quality-control checks is largely a feature of online and self-administered surveys. In interviewer-administered surveys, comparable exclusions came from validation or verification failures rather than automated speeding or straight lining flags, so older records may not carry a percentage of this kind at all.  This is worth noting mainly when comparing across eras.

In-person surveys can detect submission of falsified surveys by interviewers by sending supervisors into the field to recontact respondents and confirm that the interview took place. If fraud is detected, generally all of an interviewer’s work is removed from the study. Replacement surveys may be conducted. A thorough quality control report will include information on the quality control procedures and resulting removals, but generally will not calculate removal as a percentage, particularly if additional interviews were conducted in replacement.

8. Related Elements

This element is the numeric output of quality control, so the two are best read together.1 It also connects to the breakoff rate, since partial responses are handled at the cleaning stage, and to the sampling procedure. Because removal reshapes who remains, it feeds back into the representativeness and weighting elements: what a sample looks like after cleaning affects how it is adjusted and how well it represents its universe.7 AAPOR's disclosure elements tie data processing and exclusions to weighting and to the checks that generated the removals.1

References

1. AAPOR Transparency Initiative: Disclosure Elements (revised April 2021), Item 10. aapor.org/wp-content/uploads/2023/01/TI-Attachment-C.pdf

2. Curran, P.G. (2016). Journal of Experimental Social Psychology, 66, 4–19.

3. Meade, A.W. & Craig, S.B. (2012). Psychological Methods, 17(3), 437–455.

4. Ward, M.K. & Meade, A.W. (2023). Annual Review of Psychology, 74, 577–596. annualreviews.org

5. Zhang, C. & Conrad, F. (2014). Speeding in Web Surveys. Survey Research Methods, 8(2), 127–135.

6. AAPOR (2023). Data Quality Metrics for Online Samples: Task Force Report. aapor.org

7. AAPOR (2016). Evaluating Survey Quality in Today's Complex Environment. aapor.org

8. Roper Center iPoll study record (Axios/Ipsos American Health Index, Wave 1, #31120137). ropercenter.cornell.edu

9. AAPOR (2026). Responsible AI Integration in Survey Research, §3.2.2.

1. Plain-Language Definition

The use of Artificial Intelligence (AI) in survey research means any use of LLMs, chatbots, AI agents, or other AI technology in the process of survey design, data collection, or data analysis. 

Survey researchers have experimented with many possible applications of AI in the research process, including but not limited to questionnaire design, questionnaire translation, interviewing, response coding, and data analysis.,,,,, AAPOR’s 2026 Task Force Report on Responsible AI Integration in Survey Research identifies a large range of potential survey tasks that might be conducted by AI in part or whole, and describes the state of research on each.  The AAPOR Report is the most comprehensive overview of both the potential and risks of AI use in survey research at the current moment.

At the most extreme end of the spectrum of uses, AI has been used to model “synthetic respondents” who can be queried with new questions. These models are also referred to as “digital twins” or “silicon respondents.” AAPOR’s 2026 Code of Conduct prohibits the use of the term “participants” in describing these models. Roper Center does not currently archive data created using synthetic respondents.

In addition to the use of AI by survey researchers, respondents could potentially utilize AI in answering surveys. The most problematic possibility for survey researchers would be bots posing as respondents to complete surveys, either to access monetary incentives or to influence the results of a survey. Protection against AI use by respondents is considered under Quality Control. 

2. Why It Matters for Interpreting Results 

Disclosure of the use of AI is the most recent addition to transparency standards in the field of public opinion polling. AI technology is rapidly progressing, so at this time there is no clear consensus on how these tools might be effectively used in survey research, nor what particular risks AI poses for polling. However, some well-documented issues with AI in general can be assumed to potentially affect AI in surveys. 

AI can introduce bias, replicating and amplifying the bias in training data or introducing new algorithmic bias., AI can “hallucinate” false information, and as a result, an increasing number of academic research papers include non-existent citations. A team of U.S. academics has documented over 6000 instances of “harms or near harms realized in the real world by the deployment of artificial intelligence systems.”

AAPOR’s Code of Conduct requires disclosure of use of AI for sampling, data collection or data processing. These areas are particularly important for disclosure because quality issues cannot be identified purely on output, compared to question wording or translations in which the AI outputs can be judged on their own merits. 

3. What Effective Disclosure Looks Like

Effective disclosure includes information on the exact tasks that were conducted by AI and the validation or quality control measures researchers implemented, if any.  Because the major AI systems can give very different results, the system used should be identified, including the version number and the date the AI was used.

This level of disclosure is as yet relatively rare outside of academic papers on experimental designs, in part because the larger industry has not generally adopted many potential AI uses, particularly those related to sampling and weighting. For one of the most common implementations, use of AI to code open-ended responses, this Washington Post methodology provides both model and quality review procedures: “For Q2/Q3, open-ended responses were combined and coded by BTInsights, an AI open-end coding software that sorted responses into similar categories. Each response was then reviewed by a Washington Post polling team member to ensure it was accurately categorized, if necessary, recategorized. Category names were also edited after a review of all codes within each group.” 

Unfortunately, even the most thorough documentation of use of AI faces a fundamental disclosure stumbling block: commonly used AI systems are ever-changing and often opaque, leading to what some fear could be a “reproducibility crisis.”

4. What Incomplete Disclosure Looks Like

Most commonly, incomplete disclosure is invisible, with no information about AI use given at all. Ideally, organizations that do not use AI should specifically state that AI was not used, though in practice this is currently rare. Statements that mention that AI was used, with no detail about which AI model, what inputs were given, and what oversight was provided on the output, are insufficient to allow users to understand the possible effects of AI on the final survey results.

5. Strengths When Well-Disclosed

Well-disclosed information about the use of AI models allows users to identify possible problems based on the stage of the survey process affected. Full disclosure – including sharing of data files, prompts, code, and models used, with version information - improve the chances of replicating results, to the benefit of the research community at a critical juncture in the evaluation and adoption of AI. 

6. Limitations to Keep in Mind

Any specific guidance around disclosure of AI use in survey research is likely to become outdated quickly, given the speed with which AI technologies are developing and changing, the explosion of new AI models, the opacity of the most widely used commercial models, and the current state of widespread experimentation. Furthermore, without extensive testing and the development of accepted methods of evaluation for AI uses, even thorough disclosure of an AI process does not provide a poll consumer with sufficient information about how to judge the potential quality issues involved.  

Rothschild et al considered the range of issues facing the field in their 2025 article “Successfully Navigating the Disruption AI will Bring to Survey Research.” After calling for cataloging of use case, evaluation of models and tools, and development of evaluation frameworks and benchmarks, the authors lay out the importance of transparency and the challenges of achieving it:

Researchers must clearly document how AI tools are used throughout the survey process. Without such transparency, it becomes nearly impossible to assess the validity, replicability, or limitations of AI-mediated research. To incentivize openness, journals, conferences, and professional associations should adopt and enforce rigorous standards that make AI use auditable and replicable. This is no small challenge. At the most recent (May 2025) meeting of the American Association of Public Opinion Research (AAPOR), several presenters described their prompt engineering methods as proprietary, effectively shielding key methodological details from scrutiny. This highlights a growing tension: innovation in AI often occurs at the intersection of academia and industry, where commercial interests may conflict with scientific norms of openness and reproducibility. Many journals and associations need stronger policies requiring disclosure of commercial affiliations and funding. At the same time, collaborations and partnerships between survey statisticians, computer scientists, and behavioral researchers are critical. Ultimately, the responsible integration of AI into survey research will depend not only on technical innovation, but on a shared commitment to transparency, accountability, and cross-disciplinary collaboration.

Until standards for evaluation of AI uses in survey research are developed and adopted, disclosure alone is insufficient. But without thorough disclosure, such standards cannot be constructed.

7. Related Elements

Depending on the stage of the process in which AI is utilized, related elements might include Sampling Frame, Sampling Procedure, Sample Provider, Claims of Representativeness, Weighting, Data Collection Mode, or Question Wording. All uses of AI are related to Quality Control, as an essential element of disclosure for AI use is the review and validation process of AI outputs. 

8. Further Reading

Rothschild, D. M., Marla, J., Amaya, A., Barari, S., Buskirk, T., Cobb, C., & Webb, B. (2026). Responsible AI integration in survey research. American Association for Public Opinion Research.

Rothschild, D. M., Buskirk, T. D., Eckman, S., Hillygus, D. S., Kreuter, F., & Lazer, D. (2025). Successfully navigating the disruption AI will bring to survey research. The Survey Statistician, 92(1), 30-44.

 

  1. Xiao, Z., Zhou, M. X., Liao, Q. V., Mark, G., Chi, C., Chen, W., & Yang, H. (2020). Tell me about yourself: Using an AI-powered chatbot to conduct conversational surveys with open-ended questions. ACM Transactions on Computer-Human Interaction (TOCHI), 27(3), 1-37.
  2. DiGiuseppe, M. R., & Flynn, M. E. (2026). Scaling open-ended survey responses using llm-paired comparisons. Public Opinion Quarterly, 90(3), 630-656.
  3. Kunst, J. R., & Bierwiaczonek, K. (2023). Utilizing AI questionnaire translations in cross-cultural and intercultural research: Insights and recommendations. International Journal of Intercultural Relations, 97, 101888.
  4. Krägeloh, C. U., Alyami, M. M., & Medvedev, O. N. (2023). AI in questionnaire creation: Guidelines illustrated in AI acceptability instrument development. In International Handbook of Behavioral Health Assessment (pp. 1-23). Cham: Springer International Publishing.
  5. Kenett, R. S. (2026). Artificial Intelligence Perspectives in Survey Data Analysis. International Statistical Review.
  6. Yan, T. (2025). How NORC Is Using AI to Enhance the Research Process. NORC at the University of Chicago. https://www.norc.org/research/library/how-norc-using-ai-enhance-research-process.html
  7. Rothschild, David M., et al. "Responsible AI integration in survey research." American Association for Public Opinion Research (2026).
  8. González-Bustamante, B., Verelst, N., & Cisternas, C. (2025). Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case. arXiv preprint arXiv:2509.09871.
  9. American Association for Public Opinion Research (2026). The Code of Professional Ethics and Practices. https://aapor.org/wp-content/uploads/2026/05/New-Code-of-Ethics-for-member-approval.pdf
  10. Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), e2518075122.
  11. Ferrara, E. (2023). Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies. Sci, 6(1), 3.
  12. Hall, P., & Ellis, D. (2023). A systematic review of socio-technical gender bias in AI algorithms. Online Information Review, 47(7), 1264-1279.
  13. Bove, T. (2026). AI hallucinations are infiltrating expert work—and entering the permanent body of knowledge. Fortune.Com, N.PAG.
  14. AI Incidence Database. https://incidentdatabase.ai/
  15. Ball, P. (2023). Is AI leading to a reproducibility crisis in science?. Nature, 624(7990), 22-25.
  16. Rothschild, D. M., Buskirk, T. D., Eckman, S., Hillygus, D. S., Kreuter, F., & Lazer, D. (2025). Successfully navigating the disruption AI will bring to survey research. The Survey Statistician, 92(1), 30-44.