---
title: "Introduction to mariposa"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introduction to mariposa}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r setup}
library(mariposa)
library(dplyr)
```

## What is mariposa?

mariposa (*Marburg Initiative for Political and Social Analysis*) is a comprehensive R package for professional survey data analysis. It covers the entire workflow --- from importing SPSS, Stata, SAS, and Excel files through label management, recoding, and standardization to statistical analysis with survey weights and publication-ready output.

Every statistical result is validated against SPSS v29, so researchers migrating from SPSS can trust their numbers.

### Key Features

- **76 functions** across 15 categories
- **Full data pipeline**: import → labels → transformation → analysis → export
- **Survey weights** built into every function
- **Tidyverse integration**: pipes (`%>%`), `group_by()`, tidyselect
- **Two-level output**: compact `print()` and detailed `summary()`
- **SPSS-validated**: 4,986+ tests ensure results match SPSS v29

## The Example Dataset

mariposa includes `survey_data`, a synthetic survey of 2,500 respondents with demographics, attitudes, and a sampling weight:

```{r}
data(survey_data)
glimpse(survey_data)
```

All examples in this guide use this dataset.

## Five-Minute Tour

Here is a complete analysis workflow showing what mariposa can do:

### 1. Explore the Data

```{r}
# Find variables related to "trust"
find_var(survey_data, "trust")
```

```{r}
# Descriptive statistics with survey weights
survey_data %>%
  describe(age, income, life_satisfaction, weights = sampling_weight)
```

```{r}
# Frequency table
survey_data %>%
  frequency(education, weights = sampling_weight)
```

### 2. Transform Variables

```{r}
# Create age groups
survey_data <- rec(survey_data, age,
  rules = "18:29=1 [Young]; 30:49=2 [Middle]; 50:99=3 [Older]",
  suffix = "_group", as_factor = TRUE)

# Build a trust scale
survey_data <- survey_data %>%
  mutate(m_trust = row_means(., trust_government, trust_media, trust_science,
                             min_valid = 2))
```

### 3. Compare Groups

```{r}
# t-test with survey weights
survey_data %>%
  t_test(life_satisfaction, group = gender, weights = sampling_weight)
```

```{r}
# ANOVA across education levels
result <- survey_data %>%
  oneway_anova(life_satisfaction, group = education, weights = sampling_weight)
result
```

Every result has a detailed view with `summary()`:

```{r}
summary(result, descriptives = FALSE)
```

### 4. Post-Hoc Analysis

```{r}
# Which education groups differ?
tukey_test(result)
```

### 5. Measure Relationships

```{r}
survey_data %>%
  pearson_cor(age, income, life_satisfaction, weights = sampling_weight)
```

### 6. Build Models

```{r}
survey_data %>%
  linear_regression(life_satisfaction ~ age + income + m_trust,
                    weights = sampling_weight)
```

## Compact vs. Detailed Output

Every analysis function in mariposa provides two output levels:

- **`print()`** (default): A compact one-line summary with the key statistic
- **`summary()`**: Full SPSS-style output with all details

You can toggle individual sections in the detailed output:

```{r}
result <- survey_data %>%
  t_test(life_satisfaction, group = gender, weights = sampling_weight)

# Compact
result

# Detailed
summary(result)

# Detailed, skip effect sizes
summary(result, effect_sizes = FALSE)
```

## Grouped Analysis

All functions support `dplyr::group_by()` for subgroup analysis:

```{r}
survey_data %>%
  group_by(region) %>%
  describe(income, life_satisfaction, weights = sampling_weight)
```

```{r}
survey_data %>%
  group_by(region) %>%
  t_test(income, group = gender, weights = sampling_weight)
```

## Quick Reference

### Data Import & Export

| Function | Purpose |
|----------|---------|
| `read_spss()`, `read_por()` | Import SPSS files with tagged NA support |
| `read_stata()` | Import Stata files |
| `read_sas()`, `read_xpt()` | Import SAS files |
| `read_xlsx()` | Import Excel files with label reconstruction |
| `write_spss()` | Export to SPSS with label/missing roundtripping |
| `write_stata()` | Export to Stata |
| `write_xpt()` | Export to SAS transport format |
| `write_xlsx()` | Export to Excel (data, codebook, frequencies) |

### Label Management

| Function | Purpose |
|----------|---------|
| `var_label()` | Get/set variable labels |
| `val_labels()` | Get/set value labels |
| `find_var()` | Search variables by name or label |
| `to_label()` | Labelled → factor |
| `to_character()` | Labelled → character |
| `to_numeric()` | Factor/labelled → numeric |
| `to_labelled()` | Factor/character → labelled |
| `set_na()` | Declare values as missing |
| `unlabel()` | Strip all label metadata |
| `copy_labels()` | Restore labels after dplyr operations |
| `drop_labels()` | Remove unused value labels |

### Data Transformation

| Function | Purpose |
|----------|---------|
| `rec()` | Recode with string syntax (ranges, reverse, median split) |
| `to_dummy()` | One-hot encoding / dummy variables |
| `std()` | Z-standardization (sd, 2sd, mad, gmd methods) |
| `center()` | Mean-centering (grand-mean, group-mean) |
| `row_means()` | Row-wise means with min_valid threshold |
| `row_sums()` | Row-wise sums |
| `row_count()` | Count specific values per row |
| `pomps()` | Percent of Maximum Possible Scores (0--100) |

### Descriptive Statistics

| Function | Purpose |
|----------|---------|
| `codebook()` | Interactive HTML data dictionary |
| `describe()` | Numeric summaries (mean, sd, median, range, skewness) |
| `frequency()` | Frequency tables with valid/cumulative percent |
| `crosstab()` | Cross-tabulations with row/column/cell percentages |

### Hypothesis Testing

| Function | Purpose |
|----------|---------|
| `t_test()` | Independent and one-sample t-tests |
| `oneway_anova()` | One-way ANOVA |
| `factorial_anova()` | Multi-factor ANOVA with Type III SS |
| `ancova()` | ANCOVA with estimated marginal means |
| `mann_whitney()` | Mann-Whitney U test |
| `kruskal_wallis()` | Kruskal-Wallis H test |
| `wilcoxon_test()` | Wilcoxon signed-rank test |
| `friedman_test()` | Friedman test |
| `binomial_test()` | Exact binomial test |
| `chi_square()` | Chi-square test of independence |
| `fisher_test()` | Fisher's exact test |
| `chisq_gof()` | Chi-square goodness-of-fit |
| `mcnemar_test()` | McNemar's test for paired proportions |

### Post-Hoc & Effect Sizes

| Function | Purpose |
|----------|---------|
| `tukey_test()` | Tukey HSD pairwise comparisons |
| `scheffe_test()` | Scheffe pairwise comparisons |
| `levene_test()` | Test for homogeneity of variances |
| `dunn_test()` | Dunn's post-hoc for Kruskal-Wallis |
| `pairwise_wilcoxon()` | Pairwise Wilcoxon for Friedman |
| `phi()` | Phi coefficient |
| `cramers_v()` | Cramer's V |
| `goodman_gamma()` | Goodman-Kruskal gamma |

### Scale Analysis

| Function | Purpose |
|----------|---------|
| `reliability()` | Cronbach's Alpha with item statistics |
| `efa()` | Exploratory Factor Analysis (PCA/ML, Varimax/Oblimin/Promax) |

### Regression

| Function | Purpose |
|----------|---------|
| `linear_regression()` | Linear regression with SPSS-style output |
| `logistic_regression()` | Logistic regression with odds ratios |

### Weighted Statistics

| Function | Purpose |
|----------|---------|
| `w_mean()`, `w_median()`, `w_sd()`, `w_var()` | Central tendency and spread |
| `w_se()`, `w_quantile()`, `w_iqr()`, `w_range()` | Precision and distribution |
| `w_skew()`, `w_kurtosis()`, `w_modus()` | Shape and mode |

## Guides

Explore the full documentation:

- **Data Management**
  - `vignette("data-io")` --- Importing and exporting data
  - `vignette("labels-and-missing-values")` --- Working with labels and missing values
  - `vignette("data-transformation")` --- Recoding, standardization, and row operations
- **Core Analysis**
  - `vignette("descriptive-statistics")` --- Summaries, frequencies, and cross-tabulations
  - `vignette("hypothesis-testing")` --- Comparing groups and testing hypotheses
  - `vignette("correlation-analysis")` --- Measuring relationships between variables
- **Advanced Topics**
  - `vignette("scale-analysis")` --- Reliability, factor analysis, and scale construction
  - `vignette("regression-analysis")` --- Linear and logistic regression
  - `vignette("survey-weights")` --- Working with weighted data
