Expert-panel content validation is not one statistical task. A
relevance rating, an essential/not-essential judgment, and an
item-objective congruence judgment ask experts different questions and
therefore support different indices. expert_validity() uses
an explicit mode so those designs are not treated as
interchangeable.
The three modes are:
All three are quantitative complements to qualitative expert comments, construct coverage, comprehensibility review, and other parts of the content-validity argument.
Suppose six experts rate item relevance from 1 (not relevant) to 4 (highly relevant):
R <- matrix(
c(4,4,4,4,4,4,
4,4,4,3,4,4,
4,3,4,4,3,4,
3,3,4,3,2,3),
nrow = 6,
dimnames = list(NULL, paste0("Item", 1:4))
)
fit <- expert_validity(R, mode = "relevance", lo = 1, hi = 4, seed = 1)
fit
#> contentvalidR expert-panel analysis
#> -----------------------------------
#> Mode: relevance
#> Items: 4 | Experts/item: 6
#> Mean Aiken V: 0.875 | S-CVI/Ave: 0.958 | S-CVI/UA: 0.75
#> Strong support: 4 | Support: 0 | Review: 0
#> Panel agreement, Krippendorff's alpha (ordinal): 0.374 (95% interval -0.121
#> to 0.634). Identical rating pairs: 63.3%
#>
#> item N V ci_low ci_high I_CVI I_CVI_low I_CVI_high kappa_mod
#> Item1 6 1.000 0.824 1.000 1.000 0.610 1.00 1.000
#> Item2 6 0.944 0.742 0.990 1.000 0.610 1.00 1.000
#> Item3 6 0.889 0.672 0.969 1.000 0.610 1.00 1.000
#> Item4 6 0.667 0.437 0.837 0.833 0.436 0.97 0.816
#> recommendation
#> Strong support
#> Strong support
#> Strong support
#> Strong support
#>
#> ci_low and ci_high bound Aiken's V (Penfield-Giacobbi score interval);
#> I_CVI_low and I_CVI_high bound I-CVI.
#> 95% intervals for proportions: Wilson score (the default). Newcombe (1998)
#> compared seven methods and recommends score intervals over the Wald
#> interval. An interval reflects how few ratings an item received, not
#> whether the right judges were chosen.
#>
#> Panel agreement is one coefficient for the whole panel, whereas kappa_mod
#> describes each item. Alpha can be low when nearly every rating is the same
#> value, even on a panel that agrees closely, so read it beside the share of
#> identical rating pairs. A low alpha with many identical pairs is not by
#> itself evidence of a poor panel. Print `details$agreement` for the full
#> explanation and interval details.
#>
#> CVI thresholds shown by the workflow are common panel-size guidelines, not universal validity cutoffs.
#>
#> What these columns mean
#> V -- Aiken's V. Relevance index that rescales the experts' average rating
#> to run from 0 to 1 given the bounds of the rating scale used. (0 to
#> 1; higher is stronger)
#> I_CVI -- Item-level Content Validity Index. Proportion of experts who
#> rated the item as relevant, after applying the relevance cut. (0 to
#> 1; compared against a panel-size guideline)
#> I_CVI_low/I_CVI_high -- Interval for I-CVI. Lower and upper limits of an
#> interval around I-CVI. Expert panels are usually small, so these
#> intervals are often wide: a single I-CVI value can look more settled
#> than the number of experts behind it supports. (between 0 and 1; the
#> method and level are named in the output)
#> kappa_mod -- Modified kappa. I-CVI adjusted for the chance that experts
#> would have agreed even if rating at random. With small panels, chance
#> agreement is substantial, which is why the raw I-CVI alone can
#> overstate consensus. (0 to 1; higher is stronger)
#> agreement -- Panel-level agreement. One coefficient describing how
#> consistently the whole panel rated the item set: Krippendorff's alpha
#> by default, or Gwet's AC1 if chosen. It is separate from modified
#> kappa, which describes one item at a time. (1 is perfect agreement
#> and 0 is agreement no better than chance; it can be low on a
#> close-agreeing panel whose ratings cluster on one value)
#>
#> What the status labels mean
#> Supported -- The evidence met the criteria set for this analysis.
#> Review -- Something here needs a closer look. This is not an instruction
#> to delete anything.
#> Insufficient data -- Too little usable data to reach a judgment.
#> Descriptive only -- Reported for description only; no decision rule was
#> applied.
#> Each workflow also uses its own wording in the recommendation column
#> (Retain, Strong support, Typical, Covered, and so on). Those words map
#> onto the shared statuses above.
#>
#> See `contentvalid_glossary()` for all terms, or set
#> `options(contentvalidR.show_key = FALSE)` to hide this key.
#>
#> Use quantitative indices alongside expert comments, construct coverage, and comprehensibility review.
summary(fit)
#> Summary of expert-panel content-validity evidence
#> ---------------------------------------------
#> Mode: relevance
#> Supported: 4 | Review: 0
#> Panel agreement, Krippendorff's alpha (ordinal): 0.374 (95% interval -0.121
#> to 0.634). Identical rating pairs: 63.3%
#> No items were flagged by the workflow's quantitative review rules.
#>
#> These summaries support, but do not replace, qualitative content review.Aiken’s V rescales the bounded expert ratings to the 0-1 interval. The default confidence interval is the score interval proposed by Penfield and Giacobbi (2004), rather than a simulation-dependent bootstrap interval. Bootstrap intervals remain available through the low-level function:
aikens_v(R, lo = 1, hi = 4, ci = "bootstrap", B = 200, seed = 1)
#> item N n_missing V ci_low ci_high ci_method
#> 1 Item1 6 0 1.0000000 1.0000000 1.0000000 percentile bootstrap
#> 2 Item2 6 0 0.9444444 0.8333333 1.0000000 percentile bootstrap
#> 3 Item3 6 0 0.8888889 0.7763889 1.0000000 percentile bootstrap
#> 4 Item4 6 0 0.6666667 0.5000000 0.8333333 percentile bootstrapFor CVI, the workflow dichotomizes ratings at
relevance_cut. On a 1-4 scale the default is 3, so ratings
of 3 or 4 count as relevant. Declare a different threshold if the study
protocol used one.
The workflow reports common panel-size I-CVI guidelines (1.00 for panels of 3-5 experts and .78 for 6 or more) as review aids. They are not presented as universal proof that an item is or is not content valid. Modified kappa provides a chance-corrected complement to I-CVI.
At the scale level, S-CVI/Ave and S-CVI/UA are reported together. S-CVI/Ave is generally less brittle than universal agreement, but both should be interpreted alongside the distribution of item-level evidence.
I-CVI and modified kappa describe one item at a time. Relevance mode also reports how consistently the panel rated the whole item set, as one coefficient with a bootstrap interval:
fit$scale_summary[, c("agreement", "agreement_low", "agreement_high")]
#> agreement agreement_low agreement_high
#> 1 0.3743873 -0.1210084 0.6340909
fit$details$agreement
#> Panel-level agreement
#> Items rated by two or more raters: 4 Raters: 6
#> Krippendorff's alpha (ordinal): 0.374 95% interval: -0.121 to 0.634
#> Identical rating pairs: 63.3%
#>
#> Alpha compares the disagreement observed within items with the disagreement
#> expected if these same ratings were assigned to items at random: 1 means
#> perfect agreement and 0 means agreement no better than chance. Alpha falls
#> when ratings cluster on a few values, because little disagreement is then
#> expected by chance. A high share of identical rating pairs alongside a low
#> alpha reflects that clustering, which is common when nearly every item is
#> rated relevant, and is not by itself evidence of a poor panel.
#>
#> Krippendorff's alpha is the default because it handles ordinal ratings and
#> missing ratings (Zapf et al., 2016). It is a general reliability
#> coefficient; no publication applying it specifically to content-validity
#> panels was found.
#>
#> The interval resamples items with all of their ratings, following Zapf et
#> al. (2016), and varies slightly between runs unless `seed` is set. In 5 of
#> 1000 resamples the coefficient could not be computed, usually because every
#> resampled rating was identical; the interval uses the rest. With few items
#> this interval is imprecise and can be misleading.
#>
#> Panel agreement describes how consistently raters rated these items.
#> It does not show that the items are relevant or that the domain is covered.The default coefficient is Krippendorff’s alpha. It accepts any number of experts and missing ratings, and Zapf et al. (2016) recommend it when ratings are ordinal or incomplete, which describes most expert panels. It is a general reliability coefficient (Hayes & Krippendorff, 2007) rather than one developed for content validity; no publication applying it specifically to content-validity panels was found.
Choose the measurement level that matches the rating scale. Relevance
ratings are treated as ordinal by default.
agreement_level = "interval" treats the distances between
scale points as equal, and "nominal" treats every
disagreement as equally serious:
expert_validity(R, mode = "relevance", lo = 1, hi = 4,
agreement_level = "interval", agreement_B = 0)$scale_summary$agreement
#> [1] 0.3715847Alpha compares the disagreement within items with the disagreement expected if the same ratings were scattered across items at random. When a panel rates nearly every item 4, very little disagreement is expected by chance, so a few 3s pull alpha down even though most rating pairs are identical. Feinstein and Cicchetti (1990) described the same pattern for kappa. The output reports the share of identical rating pairs next to alpha so the two can be read together. A low alpha alongside a high share of identical pairs is not by itself evidence of a poor panel.
Gwet’s (2008) AC1 was designed to stay high in that situation, and it
is available with agreement = "ac1". It is never the
default. Vach and Gerke (2023) show that AC1 rises as ratings
concentrate in one category even when agreement does not change, and
that it can be above zero when experts rate independently. Its output
always repeats that critique. In relevance mode, AC1 is computed on the
relevant/not-relevant decision at relevance_cut:
ac1_fit <- expert_validity(R, mode = "relevance", lo = 1, hi = 4,
agreement = "ac1", agreement_B = 0)
ac1_fit$details$agreement
#> Panel-level agreement
#> Items rated by two or more raters: 4 Raters: 6
#> Gwet's AC1: 0.909
#> Identical rating pairs: 91.7%
#>
#> AC1 compares observed agreement with the agreement expected by chance,
#> estimated so that it stays high when nearly every rating falls in one
#> category (Gwet, 2008).
#>
#> Gwet's AC1 is available but is not the default. Vach and Gerke (2023) show
#> that it rises as ratings concentrate in one category even when agreement is
#> unchanged, that it can be non-zero when raters are independent, and that
#> benchmark labels developed for kappa, such as Landis and Koch's, must not
#> be applied to it.
#>
#> Panel agreement describes how consistently raters rated these items.
#> It does not show that the items are relevant or that the domain is covered.The interval resamples items with all of their ratings intact, the
procedure Zapf et al. (2016) evaluated; they found that Krippendorff’s
original bootstrap, which ignores dependence between raters, reached
only about 60% coverage. The interval varies slightly between runs, so
set seed to make it reproducible, or set
agreement_B = 0 to skip it. panel_agreement()
runs the same analysis on any rater-by-item matrix.
Lawshe’s task asks experts whether an item is essential. With twelve experts:
expert_validity(c(10, 8, 6), mode = "essentiality", N = 12)
#> contentvalidR expert-panel analysis
#> -----------------------------------
#> Mode: essentiality
#> Items: 3 | Experts/item: 12
#> Method: Lawshe CVR with exact binomial critical values
#>
#> item ne N cvr p_value critical_ne recommendation
#> Item1 10 12 0.667 0.019 10 Supported
#> Item2 8 12 0.333 0.194 10 Review
#> Item3 6 12 0.000 0.613 10 Review
#>
#> What these columns mean
#> cvr -- Lawshe's Content Validity Ratio. How far the panel leans toward
#> calling the item essential rather than merely useful. (-1 to 1; above
#> 0 means more than half the panel called it essential)
#>
#> What the status labels mean
#> Supported -- The evidence met the criteria set for this analysis.
#> Review -- Something here needs a closer look. This is not an instruction
#> to delete anything.
#> Insufficient data -- Too little usable data to reach a judgment.
#> Descriptive only -- Reported for description only; no decision rule was
#> applied.
#> Each workflow also uses its own wording in the recommendation column
#> (Retain, Strong support, Typical, Covered, and so on). Those words map
#> onto the shared statuses above.
#>
#> See `contentvalid_glossary()` for all terms, or set
#> `options(contentvalidR.show_key = FALSE)` to hide this key.
#>
#> Use quantitative indices alongside expert comments, construct coverage, and comprehensibility review.cvr() derives the smallest essential count whose
one-sided binomial upper-tail probability is no greater than
alpha. This makes the panel-size dependency explicit and
follows the exact-probability logic revisited by Ayre and Scally
(2014).
Judge-by-item binary data can be supplied directly:
E <- cbind(
Item1 = c(1,1,1,1,1,1,1,1),
Item2 = c(1,1,1,1,1,0,0,0)
)
expert_validity(E, mode = "essentiality")
#> contentvalidR expert-panel analysis
#> -----------------------------------
#> Mode: essentiality
#> Items: 2 | Experts/item: 8
#> Method: Lawshe CVR with exact binomial critical values
#>
#> item ne N cvr p_value critical_ne recommendation
#> Item1 8 8 1.00 0.004 7 Supported
#> Item2 5 8 0.25 0.363 7 Review
#>
#> What these columns mean
#> cvr -- Lawshe's Content Validity Ratio. How far the panel leans toward
#> calling the item essential rather than merely useful. (-1 to 1; above
#> 0 means more than half the panel called it essential)
#>
#> What the status labels mean
#> Supported -- The evidence met the criteria set for this analysis.
#> Review -- Something here needs a closer look. This is not an instruction
#> to delete anything.
#> Insufficient data -- Too little usable data to reach a judgment.
#> Descriptive only -- Reported for description only; no decision rule was
#> applied.
#> Each workflow also uses its own wording in the recommendation column
#> (Retain, Strong support, Typical, Covered, and so on). Those words map
#> onto the shared statuses above.
#>
#> See `contentvalid_glossary()` for all terms, or set
#> `options(contentvalidR.show_key = FALSE)` to hide this key.
#>
#> Use quantitative indices alongside expert comments, construct coverage, and comprehensibility review.A failure to clear the exact criterion is labeled
Review, not automatic deletion. Expert rationales and
domain coverage matter when deciding whether an item should be
rewritten, retained for breadth, or removed.
IOC uses expert ratings of -1, 0, and +1 for item-objective congruence. A target mapping lets the workflow compare intended and competing objectives:
d <- expand.grid(
item = c("I1", "I2"),
judge = 1:4,
objective = c("A", "B")
)
d$target_objective <- ifelse(d$item == "I1", "A", "B")
d$score <- ifelse(d$objective == d$target_objective, 1, -1)
expert_validity(d, mode = "congruence")
#> contentvalidR expert-panel analysis
#> -----------------------------------
#> Mode: congruence
#> Items: 2 | Experts/cell: 4 | Objectives: 2
#> Method: Rovinelli-Hambleton item-objective congruence
#>
#> item target target_ioc strongest_competitor competitor_ioc margin
#> I1 A 1 B -1 2
#> I2 B 1 A -1 2
#> recommendation
#> Target favored
#> Target favored
#> interpretation
#> The intended objective has the highest IOC; use the margin and expert comments to judge practical distinctiveness.
#> The intended objective has the highest IOC; use the margin and expert comments to judge practical distinctiveness.
#> status
#> Supported
#> Supported
#>
#> What these columns mean
#> ioc -- Item-Objective Congruence. How consistently experts linked the
#> item to the objective it was written for rather than to another
#> objective. (-1 to 1; higher is stronger)
#>
#> What the status labels mean
#> Supported -- The evidence met the criteria set for this analysis.
#> Review -- Something here needs a closer look. This is not an instruction
#> to delete anything.
#> Insufficient data -- Too little usable data to reach a judgment.
#> Descriptive only -- Reported for description only; no decision rule was
#> applied.
#> Each workflow also uses its own wording in the recommendation column
#> (Retain, Strong support, Typical, Covered, and so on). Those words map
#> onto the shared statuses above.
#>
#> See `contentvalid_glossary()` for all terms, or set
#> `options(contentvalidR.show_key = FALSE)` to hide this key.
#>
#> Use quantitative indices alongside expert comments, construct coverage, and comprehensibility review.The workflow reports target IOC, the strongest competitor, and their margin. This is a diagnostic comparison, not a manufactured significance test. If no target mapping is provided, all IOC cells are returned descriptively.
Missing data are never silently ignored by default. Set
na.rm = TRUE only when itemwise/cellwise deletion matches
the study protocol. Effective expert counts and missing counts are then
reported so downstream interpretation uses the actual panel size.
Each mode uses a plot matched to the expert task rather than forcing unlike indices into one generic chart.
Relevance mode displays Aiken’s V with its score interval and overlays I-CVI as a separate marker. Essentiality mode displays observed CVR against the exact panel-specific critical CVR. Congruence mode connects target IOC to the strongest competitor so the alignment margin is visually explicit. These displays are diagnostic summaries; they do not create new validity thresholds.
A concise methods/results description should identify:
A content-validity coefficient is evidence about a defined expert task. It is not, by itself, a complete validity argument.
Aiken, L. R. (1980). Content validity and reliability of single items or questionnaires. Educational and Psychological Measurement, 40(4), 955-959. https://doi.org/10.1177/001316448004000419
Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563-575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x
Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe’s content validity ratio: Revisiting the original methods of calculation. Measurement and Evaluation in Counseling and Development, 47(1), 79-86. https://doi.org/10.1177/0748175613513808
Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543-549.
Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29-48. https://doi.org/10.1348/000711006X126600
Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77-89. https://doi.org/10.1080/19312450709336664
Krippendorff, K. (2011). Computing Krippendorff’s alpha-reliability. Annenberg School for Communication, University of Pennsylvania.
Penfield, R. D., & Giacobbi, P. R., Jr. (2004). Applying a score confidence interval to Aiken’s item content-relevance index. Measurement in Physical Education and Exercise Science, 8(4), 213-225. https://doi.org/10.1207/S15327841MPEE0804_3
Polit, D. F., Beck, C. T., & Owen, S. V. (2007). Is the CVI an acceptable indicator of content validity? Appraisal and recommendations. Research in Nursing & Health, 30(4), 459-467. https://doi.org/10.1002/nur.20199
Rovinelli, R. J., & Hambleton, R. K. (1977). On the use of content specialists in the assessment of criterion-referenced test item validity. Dutch Journal of Educational Research, 2, 49-60.
Turner, R. C., & Carlson, L. (2003). Indexes of item-objective congruence for multidimensional items. International Journal of Testing, 3(2), 163-171. https://doi.org/10.1207/S15327574IJT0302_5
Vach, W., & Gerke, O. (2023). Gwet’s AC1 is not a substitute for Cohen’s kappa: A comparison of basic properties. MethodsX, 10, 102212.
Zapf, A., Castell, S., Morawietz, L., & Karch, A. (2016). Measuring inter-rater reliability for nominal data: Which coefficients and confidence intervals are appropriate? BMC Medical Research Methodology, 16, 93. https://doi.org/10.1186/s12874-016-0200-9