---
title: "Working with Survey Weights"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Working with Survey Weights}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r setup}
library(mariposa)
library(dplyr)
data(survey_data)
```

## Overview

Survey weights correct for sampling biases so your results represent the target population, not just your sample. Common reasons for weighting include oversampling of small groups, differential non-response, and post-stratification to known demographics.

In mariposa, adding weights is as simple as `weights = sampling_weight` in any function call. This guide covers the dedicated `w_*` functions, the effective sample size concept, and best practices.

## The Impact of Weights

```{r}
# Without weights (describes the sample)
unweighted <- survey_data %>%
  summarise(
    mean_age = mean(age, na.rm = TRUE),
    mean_income = mean(income, na.rm = TRUE)
  )

# With weights (represents the population)
age_w <- w_mean(survey_data, age, weights = sampling_weight)
income_w <- w_mean(survey_data, income, weights = sampling_weight)

comparison <- data.frame(
  Variable = c("Age", "Income"),
  Unweighted = c(unweighted$mean_age, unweighted$mean_income),
  Weighted = c(age_w$results$weighted_mean, income_w$results$weighted_mean)
)
print(comparison)
```

The difference between weighted and unweighted results reflects how biased your sample is. Larger differences mean weights are doing more important work.

## Weighted Statistics Functions

mariposa provides 11 `w_*` functions for individual weighted statistics.

### Central Tendency

```{r}
# Weighted mean
w_mean(survey_data, income, weights = sampling_weight)

# Weighted median
w_median(survey_data, income, weights = sampling_weight)

# Weighted mode
w_modus(survey_data, education, weights = sampling_weight)
```

### Dispersion

```{r}
# Weighted standard deviation
w_sd(survey_data, income, weights = sampling_weight)

# Weighted variance
w_var(survey_data, income, weights = sampling_weight)

# Weighted interquartile range
w_iqr(survey_data, income, weights = sampling_weight)
```

### Distribution Shape

```{r}
# Weighted skewness
w_skew(survey_data, income, weights = sampling_weight)

# Weighted kurtosis
w_kurtosis(survey_data, income, weights = sampling_weight)
```

### Precision and Range

```{r}
# Weighted standard error
w_se(survey_data, income, weights = sampling_weight)

# Weighted quantiles
w_quantile(survey_data, income,
           probs = c(0.25, 0.5, 0.75),
           weights = sampling_weight)

# Weighted range
w_range(survey_data, income, weights = sampling_weight)
```

### Multiple Variables

All `w_*` functions accept multiple variables:

```{r}
w_mean(survey_data, age, income, life_satisfaction,
       weights = sampling_weight)
```

## Effective Sample Size

Weighting reduces statistical precision. The *effective sample size* ($n_{eff}$) tells you how much information your weighted sample actually carries:

$$n_{eff} = \frac{\left(\sum w_i\right)^2}{\sum w_i^2}$$

```{r}
actual_n <- nrow(survey_data)
age_w <- w_mean(survey_data, age, weights = sampling_weight)
effective_n <- age_w$results$effective_n

cat("Actual sample size:", actual_n, "\n")
cat("Effective sample size:", round(effective_n), "\n")
cat("Design effect:", round(actual_n / effective_n, 2), "\n")
```

A design effect of 1.0 means weights have no impact on precision. Values above 1 mean larger standard errors than an unweighted sample of the same size.

## Weights in Analysis Functions

Every analysis function in mariposa supports the `weights` argument:

```{r}
# Descriptive statistics
survey_data %>%
  describe(income, life_satisfaction, weights = sampling_weight)
```

```{r}
# Frequency tables
survey_data %>%
  frequency(education, weights = sampling_weight)
```

```{r}
# t-test
survey_data %>%
  t_test(income, group = gender, weights = sampling_weight)
```

```{r}
# ANOVA
survey_data %>%
  oneway_anova(life_satisfaction, group = education,
               weights = sampling_weight)
```

```{r}
# Chi-square
survey_data %>%
  chi_square(education, employment, weights = sampling_weight)
```

```{r}
# Correlation
survey_data %>%
  pearson_cor(age, income, weights = sampling_weight)
```

```{r}
# Regression
survey_data %>%
  linear_regression(life_satisfaction ~ age + income,
                    weights = sampling_weight)
```

## Grouped Weighted Analysis

All functions work seamlessly with `group_by()`:

```{r}
survey_data %>%
  group_by(region) %>%
  describe(age, income, life_satisfaction,
           weights = sampling_weight)
```

```{r}
survey_data %>%
  group_by(region) %>%
  t_test(income, group = gender, weights = sampling_weight)
```

## Diagnosing Weight Issues

### Checking Weight Distribution

```{r}
weight_stats <- survey_data %>%
  summarise(
    min = min(sampling_weight, na.rm = TRUE),
    max = max(sampling_weight, na.rm = TRUE),
    mean = mean(sampling_weight, na.rm = TRUE),
    sd = sd(sampling_weight, na.rm = TRUE)
  )
print(weight_stats)
```

A rule of thumb: weights should range from about 0.2 to 5. Extreme values (below 0.1 or above 10) may indicate problems with the weighting scheme.

### Missing Weights

```{r}
missing <- sum(is.na(survey_data$sampling_weight))
total <- nrow(survey_data)
cat("Missing weights:", missing, "/", total,
    "(", round(missing / total * 100, 1), "%)\n")
```

### Weight Trimming

If weights are extreme, trim them to reduce variance at the cost of a small increase in bias:

```{r}
trimmed <- survey_data %>%
  mutate(
    weight_trimmed = case_when(
      sampling_weight > 5 ~ 5,
      sampling_weight < 0.2 ~ 0.2,
      TRUE ~ sampling_weight
    )
  )

original <- w_mean(survey_data, income, weights = sampling_weight)
trimmed_result <- w_mean(trimmed, income, weights = weight_trimmed)

cat("Original weighted mean:", round(original$results$weighted_mean, 1), "\n")
cat("Trimmed weighted mean:", round(trimmed_result$results$weighted_mean, 1), "\n")
```

## Complete Example

```{r}
# 1. Inspect weight properties
survey_data %>%
  summarise(
    n = n(),
    mean_weight = mean(sampling_weight, na.rm = TRUE),
    sd_weight = sd(sampling_weight, na.rm = TRUE),
    min_weight = min(sampling_weight, na.rm = TRUE),
    max_weight = max(sampling_weight, na.rm = TRUE)
  )

# 2. Compare weighted vs unweighted
vars <- c("age", "income", "life_satisfaction")
for (var in vars) {
  unw <- mean(survey_data[[var]], na.rm = TRUE)
  w_result <- w_mean(survey_data, !!sym(var), weights = sampling_weight)
  w <- w_result$results$weighted_mean
  cat(sprintf("%s: unweighted = %.2f, weighted = %.2f (diff = %.1f%%)\n",
              var, unw, w, (w - unw) / unw * 100))
}

# 3. Weighted group comparison
survey_data %>%
  group_by(region) %>%
  describe(income, weights = sampling_weight)

# 4. Weighted hypothesis test
survey_data %>%
  t_test(income, group = gender, weights = sampling_weight)
```

## Best Practices

1. **Always use weights when they are provided.** They correct for sampling biases and make your results generalizable.

2. **Report both weighted and unweighted sample sizes.** The actual *n* tells readers about data collection; the effective *n* tells them about statistical precision.

3. **Check for extreme weights.** Large variation inflates standard errors and reduces statistical power. Consider trimming if necessary.

4. **Document the weight source.** Explain how weights were calculated and what they adjust for (e.g., age, gender, region post-stratification).

5. **Report the effective sample size.** This tells readers how much precision was lost due to weighting.

## Summary

1. Survey weights make results **representative of the population**, not just the sample
2. The 11 `w_*` functions provide individual weighted statistics
3. **All analysis functions** support weights via `weights =`
4. The **effective sample size** measures how much precision weighting costs
5. **Check and document** weight distributions, extreme values, and missing weights

## Next Steps

- Get started with the package --- see `vignette("introduction")`
- Explore your data --- see `vignette("descriptive-statistics")`
- Test hypotheses --- see `vignette("hypothesis-testing")`
- Import weighted data from SPSS --- see `vignette("data-io")`
