---
title: "How to Read contentvalidR Output"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{How to Read contentvalidR Output}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
set.seed(1)
```

```{r setup}
library(contentvalidR)
```

## Who this is for

This vignette is for readers meeting these methods for the first time: graduate
students developing a scale, and researchers who need to report content-validity
evidence without first reading five primary sources.

It walks through what each part of the output means, what it does **not** mean,
and how to describe the result in a manuscript. The other vignettes show how to
*run* each workflow; this one shows how to *read* what comes back.

If you only remember one thing: **every statistic here is evidence for a
judgment you make, not a judgment the package makes for you.**

## The shared vocabulary

Every flagship workflow returns an object with the same five parts.

| Component | What it holds |
|---|---|
| `results` | One row per unit of analysis, with a status and a plain-language interpretation |
| `scale_summary` | The same evidence aggregated to the scale or panel level |
| `settings` | Every threshold and option that produced this result |
| `design` | What data you actually had: counts, missingness, effective judge numbers |
| `details` | Supporting tables and component analyses |

The unit of analysis in `results` differs by workflow, and this catches people
out. `sort_validity()`, `rating_validity()`, and `expert_validity()` return **one
row per item**. `judge_validity()` returns **one row per judge**.
`domain_validity()` returns **one row per blueprint cell**.

Every `results` table carries a `status` column drawn from a shared vocabulary:

```{r}
attr(contentvalid_glossary(), "statuses")
```

`Review` is the one people misread. It does not mean "delete this." It means
something here rewards a closer look, and you are the one who looks.

## Reading an item-sort result

```{r}
sorts <- read.csv(
  system.file("extdata", "sort_example.csv", package = "contentvalidR"),
  stringsAsFactors = FALSE
)
fit_sort <- sort_validity(sorts)
fit_sort
```

Working across the item table:

- **`n_target` / `n`** — how many judges put the item where you intended, out of
  how many judged it. Always read these before the coefficients: an impressive
  `psa` computed from four judges is not impressive.
- **`psa`** — the proportion of judges who assigned the item to its intended
  construct. It answers "did judges recognize what this item is about?"
- **`competitor`** — the construct judges picked most often *instead*. This is
  the single most useful diagnostic column in the table, because it tells you
  *where* a weak item drifted, which points at the fix.
- **`csv`** — how much more often the item went to its target than to that
  competitor. An item can have a decent `psa` and a poor `csv` if one rival
  construct keeps attracting it. That pattern means your two construct
  definitions overlap, which is a definitional problem rather than a wording
  problem.
- **`p_value`** — the Howard-Melloy exact test against chance assignment.

The scale-level table adds Colquitt strength labels. These are **percentile
positions relative to scales published in the literature**, not absolute
judgments, and they are the most frequently misreported part of this output.

## The benchmark trap

Here is the misreading to guard against, using construct ratings:

```{r}
ratings <- read.csv(
  system.file("extdata", "rating_example.csv", package = "contentvalidR"),
  stringsAsFactors = FALSE
)
fit_rating <- rating_validity(ratings)
fit_rating$scale_summary[, c("target", "mean_htc", "htc_strength",
                             "mean_htd", "htd_strength")]
```

An HTC of about 0.83 is labeled `Weak`, while an HTD of about 0.44 in the same
row is labeled `Very Strong`. Read as raw numbers this looks backwards, and it
is a common source of confusion.

The two indices are not on the same scale and their numbers are not comparable:

- **HTC is an average rating** expressed as a proportion of the scale. Across
  published scales the Moderate band starts at 0.84 and Very Strong at 0.91, so
  0.83 genuinely sits low *against that distribution*.
- **HTD is a difference** between the target rating and the best competitor's,
  so it is much smaller by construction. Its Moderate band starts at 0.18 and
  Very Strong at 0.35, so 0.44 genuinely sits high.

Compare each index against its own benchmark. Never compare an HTC number with
an HTD number, and never treat a benchmark label as a cutoff for retention.

```{r}
colquitt_benchmarks("htc")
colquitt_benchmarks("htd")
```

## Reading expert-panel evidence

```{r}
relevance <- read.csv(
  system.file("extdata", "expert_relevance_example.csv", package = "contentvalidR"),
  stringsAsFactors = FALSE
)
panel <- as.matrix(relevance[, setdiff(names(relevance), "expert")])
fit_expert <- expert_validity(panel, mode = "relevance", lo = 1, hi = 4)
fit_expert$results[, c("item", "N", "V", "I_CVI", "I_CVI_low", "I_CVI_high",
                       "kappa_mod", "recommendation")]
```

`I_CVI` is the proportion of experts calling the item relevant. `kappa_mod`
corrects that for chance agreement, and the gap between them matters most with
small panels, where several experts can agree by luck alone. Aiken's `V` uses
the whole rating scale rather than a relevance cut, so it distinguishes items
that `I_CVI` rates identically.

Note that the I-CVI criterion depends on panel size: 1.00 for three to five
experts, 0.78 for six or more. An item can therefore change status simply
because a judge was added or dropped.

### Reading the intervals

`I_CVI_low` and `I_CVI_high` bound the I-CVI. With eight experts, an item that
every expert rated relevant still has a lower limit well below 1: a panel that
size cannot rule out a noticeably lower relevance rate. Read the interval before
treating an I-CVI as settled, and report it alongside the point estimate.

The interval method is a choice. The Wilson score interval is the default,
following Newcombe (1998), and alternatives are available when a study needs to
match earlier work:

```{r}
exact <- expert_validity(panel, mode = "relevance", lo = 1, hi = 4,
                         proportion_ci = "exact")
exact$results[, c("item", "I_CVI", "I_CVI_low", "I_CVI_high")]
```

The exact (Clopper-Pearson) interval is conservative, so its limits sit further
apart. The printed output always names the method that produced the interval.

### Reading panel agreement

Relevance mode also reports one agreement coefficient for the whole panel, in
`scale_summary` as `agreement`, `agreement_low`, and `agreement_high`. The full
result, with its explanation, is in `details$agreement`, and
`panel_agreement()` produces the same result directly:

```{r}
panel_agreement(panel, seed = 1)
```

Read two numbers together: the coefficient and the share of identical rating
pairs. Krippendorff's alpha, the default, can be low on a panel that agrees
closely when nearly every rating is the same value, because chance then predicts
very little disagreement. A low alpha next to a high share of identical pairs
reflects that clustering. Agreement describes the raters, not the items: a
consistent panel can still consistently rate an item as irrelevant.

## Reading judge heterogeneity

`judge_validity()` asks a different question: do your conclusions depend on
*these particular judges*?

```{r}
judge_ratings <- rbind(
  c(4, 4, 4, 3, 2, 2), c(4, 4, 3, 4, 2, 1), c(4, 3, 4, 4, 1, 2),
  c(3, 4, 4, 4, 2, 2), c(4, 4, 4, 4, 2, 1), c(4, 3, 4, 3, 1, 2),
  c(4, 4, 3, 4, 2, 2), c(2, 2, 2, 2, 1, 1)
)
dimnames(judge_ratings) <- list(paste0("Judge", 1:8), paste0("Item", 1:6))
fit_judge <- judge_validity(judge_ratings, lo = 1, hi = 4)
fit_judge$results[, c("judge", "mean_rating", "severity_raw",
                      "differentiation", "n_items_flipped", "recommendation")]
```

`severity_raw` is signed so that **positive means harsher**: the judge rates
lower than the panel. `differentiation` near 1 means the judge used the scale
about as widely as everyone else. `n_items_flipped` is the influence
diagnostic: how many items would change review status if this judge were
removed.

A judge flagged here is not a judge to delete. A dissenting expert may be the
one reading the construct definition correctly. The flag tells you a conclusion
rests on one person's ratings, which is worth knowing before you write it up.

The panel-level dependability coefficient answers a planning question:

```{r}
summary(fit_judge)$gtheory$judges_needed
```

That table reports how many judges a target coefficient would need. Where it
returns `NA`, no realistic panel reaches that target, which usually means the
judges barely distinguished the items.

## Reading domain coverage

```{r}
assignments <- data.frame(
  item = paste0("I", 1:7),
  construct = c("Autonomy", "Autonomy", "Autonomy", "Autonomy",
                "Competence", "Competence", "Relatedness"),
  stringsAsFactors = FALSE
)
fit_domain <- domain_validity(
  assignments,
  cell_col = "construct",
  domain = c("Autonomy", "Competence", "Relatedness", "Belonging")
)
fit_domain$results[, c("cell", "n_items", "share", "recommendation")]
```

The `Belonging` row is the point of this analysis. No item addresses that cell,
and **no item-level index could ever have told you**, because relevance
statistics can only describe items that exist. A perfect I-CVI on every item you
wrote says nothing about the facet you forgot.

This only works because the full cell list was supplied through `domain`. Omit
it and the analysis can only describe the cells that already contain items, and
the output says so rather than implying complete coverage.

## Common misreadings

**"Review means delete."** It does not. It means look closer. Deleting every
flagged item optimizes a statistic at the cost of the content domain, which is
the opposite of content validity.

**"Strong means good."** Benchmark labels are percentile positions against
published scales. `Strong` means typical of published work, not that the item
is fit for your purpose.

**"One index is enough."** No single coefficient establishes content validity.
These are components of an argument that also rests on construct definitions,
domain coverage, cognitive interviewing, and expert comment.

**"The numbers are comparable."** HTC against HTD, Psa against Csv, or one index
against the benchmark belonging to a different index: these comparisons are not
meaningful.

**"A high coefficient means the domain is covered."** Relevance and coverage are
different questions. Only `domain_validity()` addresses coverage, and only when
you give it the full blueprint.

**"More judges always helps."** More judges narrows intervals, but it can also
change a criterion: the I-CVI guideline shifts at six experts. Check
`judges_needed` rather than assuming.

## Writing it up

A defensible sentence reports the estimate, its precision, the decision rule you
applied, and what you concluded, without implying the statistic made the
decision:

> Six items were sorted by 20 naive judges. Four items met the Howard-Melloy
> exact target-assignment criterion (*p* < .05). Items B2 and C2 did not
> (*psa* = .65 and .70), and in both cases the competing construct was the
> adjacent scale. Rather than removing them, we revised their wording to
> sharpen the distinction from that scale and re-sorted.

Report what you cannot establish as well:

> These analyses address definitional correspondence and distinctiveness. They
> do not establish that the item set covers the construct domain; coverage was
> assessed separately against the blueprint.

`settings` and `design` hold everything a reader needs to reproduce the decision
rules, so report them rather than only the coefficients:

```{r}
fit_sort$settings[c("p0", "alpha", "judge_type")]
```

## Turning off the inline key

Printed output includes a key defining each column. Once the terms are familiar:

```{r, eval = FALSE}
options(contentvalidR.show_key = FALSE)
```

The definitions remain available at any time:

```{r, eval = FALSE}
contentvalid_glossary("expert-panel")
```
