The Hinkin and Tracey (1999) content-rating procedure asks judges to evaluate how well each item corresponds to each construct definition under consideration. The typical design is fully crossed within judges: the same judge rates an item against the intended definition and against one or more orbiting definitions.
That design provides two complementary kinds of evidence:
contentvalidR keeps those questions separate rather than
reducing the study to a single coefficient.
rating_dat <- expand.grid(
item = c("A1", "A2", "A3", "B1"),
rater = 1:24,
construct = c("A", "B", "C")
)
rating_dat$target_construct <- ifelse(rating_dat$item == "B1", "B", "A")
rating_dat$rating <- ifelse(
rating_dat$construct == rating_dat$target_construct,
pmin(5, pmax(1, round(rnorm(nrow(rating_dat), 4.4, .6)))),
pmin(5, pmax(1, round(rnorm(nrow(rating_dat), 2.2, .8))))
)Each item-judge combination appears once for every construct definition. A duplicated item-rater-construct row is treated as a data error.
Following Colquitt et al. (2019), the Hinkin-Tracey correspondence index is
\[ HTC = \frac{\bar{x}_{target}}{a}, \]
where \(a\) is the number of
response anchors when ratings use a 1-to-\(a\) scale. contentvalidR can
also accept an equally spaced integer scale such as 0-to-4; it shifts
that scale internally to the equivalent 1-to-5 anchor metric before
computing HTC.
htc(rating_dat, scale_min = 1, scale_max = 5)
#> item target n_target target_mean anchors htc
#> 1 A1 A 24 4.375000 5 0.8750000
#> 2 A2 A 24 4.458333 5 0.8916667
#> 3 A3 A 24 4.416667 5 0.8833333
#> 4 B1 B 24 4.583333 5 0.9166667Higher HTC means stronger correspondence with the intended definition.
HTD compares intended-definition ratings with orbiting-definition ratings:
\[ HTD = \frac{\text{average}(x_{target} - x_{orbiting})}{a - 1}. \]
It ranges from -1 to 1. Positive values favor the intended definition; negative values indicate that orbiting definitions are rated more highly on average.
htd(rating_dat, scale_min = 1, scale_max = 5)
#> item target n_complete n_pairs target_mean_complete strongest_competitor
#> 1 A1 A 24 48 4.375000 B
#> 2 A2 A 24 48 4.458333 B
#> 3 A3 A 24 48 4.416667 C
#> 4 B1 B 24 48 4.583333 C
#> competitor_mean anchors htd
#> 1 2.375000 5 0.5572917
#> 2 2.000000 5 0.6302083
#> 3 2.250000 5 0.5468750
#> 4 2.166667 5 0.6145833The item-level table also identifies the strongest orbiting competitor. That is often more useful for revision than merely knowing that distinctiveness is weak.
The same judges provide multiple construct ratings, so those
observations are not independent. anova_content() uses a
one-way repeated-measures ANOVA for the standard fully crossed design
and follows it with planned paired comparisons of the intended
definition against every orbiting definition.
aov_out <- anova_content(rating_dat, design = "within")
aov_out
#> item target design n_raters n_complete n_constructs target_mean
#> 1 A1 A within 24 24 3 4.375000
#> 2 A2 A within 24 24 3 4.458333
#> 3 A3 A within 24 24 3 4.416667
#> 4 B1 B within 24 24 3 4.583333
#> strongest_competitor competitor_mean F df1 df2 p
#> 1 B 2.375000 72.64064 2 46 5.820538e-15
#> 2 B 2.000000 105.82309 2 46 6.164319e-18
#> 3 C 2.250000 82.24514 2 46 6.443079e-16
#> 4 C 2.166667 108.28649 2 46 3.987307e-18
#> epsilon_gg df1_gg df2_gg p_gg p_screen partial_eta2
#> 1 0.9923166 1.984633 45.64657 7.290445e-15 7.290445e-15 0.7595165
#> 2 0.8090412 1.618082 37.21590 6.097704e-15 6.097704e-15 0.8214606
#> 3 0.9992587 1.998517 45.96590 6.595204e-16 6.595204e-16 0.7814626
#> 4 0.9262517 1.852503 42.60758 5.895899e-17 5.895899e-17 0.8248106
#> min_mean_diff max_contrast_p contrast_pass posthoc_pass
#> 1 2.000000 8.344091e-10 TRUE TRUE
#> 2 2.458333 1.372752e-10 TRUE TRUE
#> 3 2.166667 5.918580e-11 TRUE TRUE
#> 4 2.416667 1.786074e-12 TRUE TRUE
attr(aov_out, "contrasts")
#> item design target competitor n mean_target mean_competitor mean_diff
#> 1 A1 within A B 24 4.375000 2.375000 2.000000
#> 2 A1 within A C 24 4.375000 1.916667 2.458333
#> 3 A2 within A B 24 4.458333 2.000000 2.458333
#> 4 A2 within A C 24 4.458333 1.875000 2.583333
#> 5 A3 within A B 24 4.416667 2.208333 2.208333
#> 6 A3 within A C 24 4.416667 2.250000 2.166667
#> 7 B1 within B A 24 4.583333 2.083333 2.500000
#> 8 B1 within B C 24 4.583333 2.166667 2.416667
#> t df p p_adj dz pass
#> 1 9.591663 23 8.344091e-10 8.344091e-10 1.957890 TRUE
#> 2 11.336315 23 3.409941e-11 3.409941e-11 2.314016 TRUE
#> 3 10.552406 23 1.372752e-10 1.372752e-10 2.154001 TRUE
#> 4 17.643975 23 3.644893e-15 3.644893e-15 3.601561 TRUE
#> 5 11.072214 23 5.409806e-11 5.409806e-11 2.260106 TRUE
#> 6 11.021286 23 5.918580e-11 5.918580e-11 2.249711 TRUE
#> 7 13.133926 23 1.786074e-12 1.786074e-12 2.680951 TRUE
#> 8 14.269216 23 3.241706e-13 3.241706e-13 2.912692 TRUEThe omnibus F test asks whether the item’s mean ratings differ somewhere across definitions. The planned contrasts ask the more direct content-validity question: is the target mean higher than each orbiting mean?
With more than two construct definitions, the conventional
repeated-measures F test assumes sphericity. contentvalidR
reports that historical omnibus test but does not hide the assumption.
The planned target-versus-orbiting comparisons are therefore important
diagnostic evidence rather than decorative post-hoc tests.
fit <- rating_validity(
rating_dat,
scale_min = 1,
scale_max = 5
)
fit
#> contentvalidR construct-rating analysis
#> ---------------------------------------
#> Items: 4 | Raters: 24 | Target scales: 2 | Constructs: 3
#> Design: within-judge ratings | Scale: 1 to 5
#> Item inference: one-way repeated-measures ANOVA (Greenhouse-Geisser corrected omnibus p) plus planned paired target-versus-orbiting contrasts
#> Planned-contrast adjustment: none
#> Judges: naive
#>
#> 4 item(s) meet the full item-level screening criterion; 0 item(s) are flagged for review.
#>
#> Item-level evidence:
#> item target n_complete strongest_competitor htc htd p_value max_contrast_p
#> A1 A 24 B 0.875 0.557 0 0
#> A2 A 24 B 0.892 0.630 0 0
#> A3 A 24 C 0.883 0.547 0 0
#> B1 B 24 C 0.917 0.615 0 0
#> recommendation
#> Retain
#> Retain
#> Retain
#> Retain
#>
#> Target-scale Colquitt benchmark summary:
#> target n_items n_htc n_htd mean_htc htc_strength mean_htd htd_strength
#> A 3 3 3 0.883 Strong 0.578 Very Strong
#> B 1 1 1 0.917 Very Strong 0.615 Very Strong
#> benchmark_set
#> overall
#> overall
#>
#> Colquitt labels are empirical percentile norms for scale-level HTC/HTD averages, not universal cutoffs.
#> HTC is an average rating and HTD is a difference between ratings, so they sit on
#> different scales with different typical values. A high HTC can be labeled Weak in
#> the same analysis where a much smaller HTD is labeled Very Strong. Compare each
#> index against its own benchmark, never against the other index's number.
#>
#> What these columns mean
#> htc -- Hinkin-Tracey Correspondence. Average rating of the item against
#> its intended construct definition, expressed as a proportion of the
#> rating scale. (0 to 1; higher is stronger)
#> htd -- Hinkin-Tracey Distinctiveness. How far the intended construct's
#> average rating exceeds the best competing construct's, as a
#> proportion of the rating scale. It is a difference, so its typical
#> values are far smaller than HTC's. (usually a small positive number;
#> higher is stronger)
#>
#> What the status labels mean
#> Supported -- The evidence met the criteria set for this analysis.
#> Review -- Something here needs a closer look. This is not an instruction
#> to delete anything.
#> Insufficient data -- Too little usable data to reach a judgment.
#> Descriptive only -- Reported for description only; no decision rule was
#> applied.
#> Each workflow also uses its own wording in the recommendation column
#> (Retain, Strong support, Typical, Covered, and so on). Those words map
#> onto the shared statuses above.
#>
#> See `contentvalid_glossary()` for all terms, or set
#> `options(contentvalidR.show_key = FALSE)` to hide this key.
#>
#> 'Review' is not an automatic deletion decision. Consider construct definitions, item wording,
#> orbiting-construct choice, domain coverage, and qualitative judge feedback.
summary(fit)
#> Summary of construct-rating content-validity evidence
#> ---------------------------------------------------
#> Retain: 4 of 4 item(s)
#> Review: 0 of 4 item(s)
#>
#> Target-scale evidence:
#> target n_items n_htc n_htd n_retain n_review mean_htc htc_strength mean_htd
#> A 3 3 3 3 0 0.883 Strong 0.578
#> B 1 1 1 1 0 0.917 Very Strong 0.615
#> htd_strength overall_strength
#> Very Strong Strong
#> Very Strong Very Strong
#>
#> A: Strong normative standing on the weaker of definitional correspondence (HTC) and distinctiveness (HTD).
#> B: Very Strong normative standing on the weaker of definitional correspondence (HTC) and distinctiveness (HTD).
#>
#> All analyzed items met the item-level inferential screening criterion.
#>
#> Interpret these results alongside theory, domain coverage, and qualitative feedback.
#> The analysis does not by itself establish comprehensiveness or the full content-validity argument.The item-level recommendation has deliberately limited meaning:
Review is not an instruction to delete an item. Content
coverage can be harmed by mechanical item deletion.
Colquitt et al. (2019) created empirical norms from
scale-level averages of HTC and HTD across 112
published scales. rating_validity() therefore averages item
HTC/HTD within each target scale before assigning those descriptive
normative labels.
fit$scale_summary
#> target n_items n_htc n_htd n_retain n_review n_insufficient mean_htc
#> 1 A 3 3 3 3 0 0 0.8833333
#> 2 B 1 1 1 1 0 0 0.9166667
#> htc_strength mean_htd htd_strength overall_strength orbiting_r benchmark_set
#> 1 Strong 0.5781250 Very Strong Strong NA overall
#> 2 Very Strong 0.6145833 Very Strong Very Strong NA overall
#> evidence
#> 1 Strong normative standing on the weaker of definitional correspondence (HTC) and distinctiveness (HTD).
#> 2 Very Strong normative standing on the weaker of definitional correspondence (HTC) and distinctiveness (HTD).
colquitt_benchmarks("htc")
#> statistic benchmark_set benchmark_label interpretation
#> 1 htc overall Overall (not correlation-normed) Very Strong
#> 2 htc overall Overall (not correlation-normed) Strong
#> 3 htc overall Overall (not correlation-normed) Moderate
#> 4 htc overall Overall (not correlation-normed) Weak
#> 5 htc overall Overall (not correlation-normed) Lack of
#> percentile minimum
#> 1 80th-99th 0.91
#> 2 60th-79th 0.87
#> 3 40th-59th 0.84
#> 4 20th-39th 0.60
#> 5 0th-19th -Inf
colquitt_benchmarks("htd")
#> statistic benchmark_set benchmark_label interpretation
#> 1 htd overall Overall (not correlation-normed) Very Strong
#> 2 htd overall Overall (not correlation-normed) Strong
#> 3 htd overall Overall (not correlation-normed) Moderate
#> 4 htd overall Overall (not correlation-normed) Weak
#> 5 htd overall Overall (not correlation-normed) Lack of
#> percentile minimum
#> 1 80th-99th 0.35
#> 2 60th-79th 0.27
#> 3 40th-59th 0.18
#> 4 20th-39th 0.04
#> 5 0th-19th -InfThe overall bands are empirical percentile standing, not universal validity cutoffs. If the average correlation between a focal scale and its orbiting scales is known, correlation-conditional norms can be requested:
rating_validity(
rating_dat,
orbiting_r = c(A = .42, B = .55)
)$scale_summary
#> target n_items n_htc n_htd n_retain n_review n_insufficient mean_htc
#> 1 A 3 3 3 3 0 0 0.8833333
#> 2 B 1 1 1 1 0 0 0.9166667
#> htc_strength mean_htd htd_strength overall_strength orbiting_r benchmark_set
#> 1 Moderate 0.5781250 Very Strong Moderate 0.42 moderate
#> 2 Very Strong 0.6145833 Very Strong Very Strong 0.55 stronger
#> evidence
#> 1 Generally supportive normative standing, with at least one content-validity dimension in the moderate range; inspect weaker items and construct overlap before finalizing the scale.
#> 2 Very Strong normative standing on the weaker of definitional correspondence (HTC) and distinctiveness (HTD).A given level of distinctiveness can be more impressive when the focal and orbiting constructs are known to correlate strongly.
Colquitt et al.’s normative distributions were developed using naive judges representative of substantive target populations. Their paper cautions against applying those norms to expert panels. The package therefore separates calculation from norm applicability:
rating_validity(rating_dat, judge_type = "expert")$scale_summary
#> target n_items n_htc n_htd n_retain n_review n_insufficient mean_htc
#> 1 A 3 3 3 3 0 0 0.8833333
#> 2 B 1 1 1 1 0 0 0.9166667
#> htc_strength mean_htd htd_strength overall_strength orbiting_r benchmark_set
#> 1 <NA> 0.5781250 <NA> <NA> NA overall
#> 2 <NA> 0.6145833 <NA> <NA> NA overall
#> evidence
#> 1 HTC/HTD are reported descriptively; Colquitt et al. (2019) normative labels are suppressed for expert judges.
#> 2 HTC/HTD are reported descriptively; Colquitt et al. (2019) normative labels are suppressed for expert judges.HTC/HTD are still computed, but the Colquitt labels are suppressed.
For HTD and repeated-measures inference, a judge must have a usable rating for every construct definition presented for that item. Incomplete profiles are excluded itemwise and counted explicitly in the output. This preserves the paired design rather than quietly treating incomplete repeated observations as independent data.
The original one-index views remain available:
A correspondence-distinctiveness evidence map displays HTC and HTD together:
Target-scale averages are shown as diamonds and items needing review are labeled by default. As with the item-sort map, Colquitt norm regions are not drawn across individual items because those benchmarks were constructed from scale averages.
The target-versus-competitor gap plot makes the Hinkin-Tracey mean-rating logic more directly visible:
Filled points are intended-definition means, open points are the strongest orbiting-definition means, and the connecting segment is the observed content distinctiveness gap. A reversed segment immediately identifies an item whose strongest competitor outrates its intended definition. This is a graphical extension of the mean-rating tables used in the original procedure, not a new statistical cutoff.
A useful report should identify:
The quantitative analysis is evidence about definitional correspondence and distinctiveness. It does not by itself demonstrate that the item pool comprehensively samples the full construct domain.
Hinkin, T. R., & Tracey, J. B. (1999). An analysis of variance approach to content validation. Organizational Research Methods, 2(2), 175-186. https://doi.org/10.1177/109442819922004
Colquitt, J. A., Sabey, T. B., Rodell, J. B., & Hill, E. T. (2019). Content validation guidelines: Evaluation criteria for definitional correspondence and definitional distinctiveness. Journal of Applied Psychology, 104(10), 1243-1265. https://doi.org/10.1037/apl0000406