Diagnostics for Respondent-driven Sampling

Krista J Gile, Lisa G Johnston, Matthew J Salganik, Krista J Gile, Lisa G Johnston, Matthew J Salganik

Abstract

Respondent-driven sampling (RDS) is a widely used method for sampling from hard-to-reach human populations, especially populations at higher risk for HIV. Data are collected through peer-referral over social networks. RDS has proven practical for data collection in many difficult settings and is widely used. Inference from RDS data requires many strong assumptions because the sampling design is partially beyond the control of the researcher and partially unobserved. We introduce diagnostic tools for most of these assumptions and apply them in 12 high risk populations. These diagnostics empower researchers to better understand their data and encourage future statistical research on RDS.

Keywords: HIV/AIDS; diagnostics; exploratory data analysis; hard-to-reach populations; link-tracing sampling; non-ignorable design; respondent-driven sampling; social networks; survey sampling.

Figures

Fig. 1
Fig. 1
Recruitment Trees Plot from sample of men who have sex with men in Higuey. Shading indicates self-identify as “heterosexual.”
Fig. 2
Fig. 2
Sample sizes from the 12 studies. In total, 3,866 people participated, of which 1,677 (43%) completed a follow-up survey.
Fig. 3
Fig. 3
Percent of respondents reporting 0, 1-3, or 4+ failed recruitment attempts. In 6 sites, at least 25% of respondents reported at least one failed recruitment attempt.
Fig. 4
Fig. 4
Histogram of difference between Successive Sampling and Volz-Heckathorn estimators, over many traits. Successive Sampling estimates based on a “worst case” small approximated population size, based on the lower bound of the Highest Posterior Density interval generated by the population size estimation method in Handcock et al. (2012).
Fig. 5. Convergence Plots showing p̂ 1…
Fig. 5. Convergence Plots showing 1, 2, . . . , p̂n
The headers and footers plot the sample observations with and without the trait. The white line shows the estimate based on the complete sample (p̂n).
Fig. 6
Fig. 6
Convergence Plot and Bottleneck Plot for the proportion of MSM in Santo Domingo that have had sex with a women in the last six month. The Convergence Plot masks important differences between the seeds that are revealed by the Bottleneck Plot. In both plots, the white line shows the estimate based on the complete sample (p̂n).
Fig. 7
Fig. 7
Three diagnostic plots for estimates of MSM in Higuey that self-identify as heterosexual. The Convergence Plot (a) shows that data collected late in the sample differs from data collected early in the sample. The Bottleneck Plot (b) shows that the chains explored different subgroups suggesting a problem with bottlenecks. The All Points Plots (c) shows that the self-identified heterosexuals (represented by up-ticks in the plot) were unusual in that they both arrived in the sample late and arrived from a small number of chains, a fact that is difficult to infer from the previous two plots.
Fig. 8
Fig. 8
Disease prevalence estimates from 12 studies for 4 diseases using Question G at enrollment (solid circle) and follow-up (hollow circle). The plot includes only people who participated in both the initial and follow-up survey (see Fig. 2 for sample sizes).
Fig. 9. Recruitment Effectiveness Plot
Fig. 9. Recruitment Effectiveness Plot
Average recruits for HIV+ and HIV− respondents by site. The ratio is provided under the bars. Differential recruitment effectiveness can lead to bias in the RDS estimates under some conditions (Tomas and Gile, 2011).
Fig. 10. Recruitment Bias Plot
Fig. 10. Recruitment Bias Plot
Percent of drug users employed, by location and question.
Fig. 11. Motivation-Outcome Plot
Fig. 11. Motivation-Outcome Plot
Odds ratios of having HIV given HIV test as motivation for study participation. Ratios greater than 1 indicate those participating for HIV test results more likely to have HIV. For reference, nominal 95% intervals are based on the inversion of Fisher's exact test (these would be confidence intervals if the data were independent identically distributed).

Source: PubMed

Upcoming Clinical Trials

Subscribe