| Type: | Package |
| Title: | Conformal Local Influence Screening for Bounded-Response Regression |
| Version: | 0.3.6 |
| Description: | Provides fast, statistically calibrated influence diagnostics for zero-or-one inflated beta (BIc) regression models with variable dispersion. The core idea is to use the conformal normal curvature of Poon and Poon (1999) <doi:10.1111/1467-9868.00162> as a non-conformity score within a split-conformal testing procedure, yielding per-observation conformal p-values whose Benjamini-Hochberg adjustment controls the false discovery rate at a user-specified level (Bates and others, 2023) <doi:10.1214/22-AOS2244>. Unlike classical local influence diagnostics, which rely on visual inspection of index plots and do not scale beyond a few hundred observations, 'clis' provides a finite-sample error guarantee and runs in linear time per observation after a single model fit. Methods for four perturbation schemes, block decomposition of influence into the inflation-probability and conditional-mean/precision components, penalised additive (semiparametric) submodels, and a full suite of diagnostic plots are included. |
| License: | GPL-3 |
| URL: | https://github.com/Raydonal/clis, https://raydonal.github.io/clis/ |
| BugReports: | https://github.com/Raydonal/clis/issues |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1.0) |
| Imports: | gamlss (≥ 5.4.0), gamlss.dist (≥ 6.0.0), stats, graphics, grDevices, utils |
| Suggests: | testthat (≥ 3.0.0), knitr, rmarkdown, ggplot2, betareg, gamlss.data |
| VignetteBuilder: | knitr |
| Config/testthat/edition: | 3 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-19 00:32:12 UTC; raydonal |
| Author: | Raydonal Ospina |
| Maintainer: | Raydonal Ospina <raydonal@de.ufpe.br> |
| Config/roxygen2/version: | 8.1.0 |
| Repository: | CRAN |
| Date/Publication: | 2026-09-29 14:00:17 UTC |
clis: Conformal Local Influence Screening for Bounded-Response Regression
Description
Provides fast, statistically calibrated influence diagnostics for zero-or-one inflated beta (BIc) regression models with variable dispersion. The core idea is to use the conformal normal curvature of Poon and Poon (1999) doi:10.1111/1467-9868.00162 as a non-conformity score within a split-conformal testing procedure, yielding per-observation conformal p-values whose Benjamini-Hochberg adjustment controls the false discovery rate at a user-specified level (Bates and others, 2023) doi:10.1214/22-AOS2244. Unlike classical local influence diagnostics, which rely on visual inspection of index plots and do not scale beyond a few hundred observations, 'clis' provides a finite-sample error guarantee and runs in linear time per observation after a single model fit. Methods for four perturbation schemes, block decomposition of influence into the inflation-probability and conditional-mean/precision components, penalised additive (semiparametric) submodels, and a full suite of diagnostic plots are included.
Author(s)
Maintainer: Raydonal Ospina raydonal@de.ufpe.br (ORCID)
Authors:
Raydonal Ospina raydonal@de.ufpe.br (ORCID)
See Also
Useful links:
Report bugs at https://github.com/Raydonal/clis/issues
Observed or Fisher information matrix for a BIc model
Description
Computes the block-structured information matrix of a fitted zero-or-one inflated beta regression with variable dispersion. By construction the matrix is block diagonal between the inflation parameters and the mean/precision parameters (information orthogonality).
Usage
bic_info(object, use_fisher = TRUE, penalty = NULL)
Arguments
object |
A fitted |
use_fisher |
Logical; if |
penalty |
Optional penalty matrix |
Value
A list with the information matrix (info), its inverse
(info_inv), the parameter-block indices (idx), the weight
sequences (weights), and (if a penalty was supplied) the effective
degrees of freedom (edf) and their block split (edf_blocks).
References
Ospina, R. and Ferrari, S. L. P. (2012). A general class of zero-or-one inflated beta regression models. Computational Statistics & Data Analysis, 56(6), 1609-1623.
Link function and its derivatives for BIc submodels
Description
Constructs a list with the link function, its inverse, and the first and second derivatives of the inverse link, used throughout the influence computations.
Usage
bic_link(link)
Arguments
link |
Character string naming the link. One of |
Value
A list with components name, linkfun, linkinv, mu.eta
(first derivative of the inverse link), and mu.eta2 (second derivative).
Examples
lk <- bic_link("logit")
lk$mu.eta(0.3) # d mu / d eta at mu = 0.3
Extract the penalty matrix from a fitted additive BIc model
Description
Assembles the block-diagonal penalty matrix S(lambda) from the smooth
terms of a fitted gamlss model whose submodels use penalised additive
terms (for example pb() P-spline terms). The result is suitable as the
penalty argument of bic_info().
Usage
bic_penalty(object)
Arguments
object |
A fitted |
Details
The function reads the smoothing structure stored by gamlss
for each parameter (mu, sigma, nu) and places the corresponding
penalty contributions on the diagonal blocks. Terms without a penalty
contribute a zero block, so a model with a mix of linear and smooth
terms is handled transparently. By construction the returned matrix is
block diagonal across the three coefficient groups, so it preserves the
separability required by the semiparametric theory.
Value
A block-diagonal penalty matrix of dimension
(M+m+p) x (M+m+p), block diagonal across the inflation, mean, and
precision coefficient groups.
See Also
Randomized quantile residuals for BEZI/BEOI models
Description
Computes randomized quantile residuals for a fitted zero- or one-inflated
beta model. For the continuous observations the beta CDF is scaled by the
non-inflation probability; for the boundary observations the residual is
randomized uniformly over the probability atom, which is the standard
construction for a mixed discrete-continuous distribution. This is provided
because the generic residuals() method does not always handle the
inflation atom correctly, which can make every residual share the same sign.
Usage
bic_quantile_residuals(object, seed = NULL)
Arguments
object |
A fitted |
seed |
Optional integer seed for the randomization at the boundary. |
Value
A numeric vector of randomized quantile residuals.
Conformal Local Influence Screening
Description
Performs scalable, FDR-controlled influence screening for a fitted
zero-or-one inflated beta regression model. The conformal normal curvature
score of each observation is used as a non-conformity score within a
split-conformal procedure: the data are partitioned into a calibration set
(presumed clean) and a screening set; conformal p-values are computed for
the screening set; and Benjamini-Hochberg adjustment declares an
influential subset with false discovery rate controlled at level alpha.
Usage
clis_screen(
object,
scheme = "caseweights",
p = 1L,
alpha = 0.1,
calib_frac = 0.5,
calib_idx = NULL,
use_fisher = TRUE,
r_max = 4L,
score = c("B_Et", "m_r"),
penalised = FALSE,
penalty = NULL,
seed = NULL
)
Arguments
object |
A fitted |
scheme |
Character; the perturbation scheme. One of |
p |
Integer; covariate index for the covariate-perturbation schemes. |
alpha |
Numeric in (0, 1); target false discovery rate. Default 0.1. |
calib_frac |
Numeric in (0, 1); fraction of observations used for
conformal calibration when |
calib_idx |
Optional integer vector of observation indices to use as
the calibration set. The conformal false discovery rate guarantee
requires the calibration set to be (nearly) free of influential points;
when a trusted clean subset is known, supply it here. If |
use_fisher |
Logical; use the Fisher information (default |
r_max |
Integer; order for aggregate contributions used as the score. |
score |
Character; which CNC quantity to use as the non-conformity
score: |
penalised |
Logical; if |
penalty |
Optional penalty matrix, of the dimension of the full coefficient vector, added to the information before inversion. Use this when the smooth basis is supplied explicitly as design columns. |
seed |
Optional integer for reproducible calibration splitting. |
Details
Classical local influence diagnostics rank observations by a curvature
measure and rely on visual inspection of an index plot against an ad-hoc
threshold. This neither scales to large n nor provides any control of
the type-I error rate. CLIS addresses both limitations: the conformal
wrapper supplies a finite-sample FDR guarantee, and the per-observation
score is computed in linear time after a single model fit.
The calibration set is assumed to be predominantly free of influential observations. Because influential points are typically rare, a random split satisfies this approximately; for adversarial settings, a robust pre-filter can be applied before calibration.
Value
An object of class clis with components:
- influential
Integer indices (into the screening set) declared influential at FDR level
alpha.- influential_global
Integer indices into the original data.
- pvalues
Conformal p-values for the screening set.
- padj
Benjamini-Hochberg adjusted p-values.
- scores
Non-conformity scores for all observations.
- calib_idx, screen_idx
Calibration and screening indices.
- alpha, scheme
Inputs echoed back.
- decomp
Block decomposition (case-weights scheme only).
Examples
if (requireNamespace("gamlss", quietly = TRUE) &&
requireNamespace("betareg", quietly = TRUE)) {
data("ReadingSkills", package = "betareg")
ReadingSkills$dys <- as.numeric(ReadingSkills$dyslexia == "yes")
fit <- gamlss::gamlss(accuracy1 ~ dys * iq,
sigma.formula = ~ dys, nu.formula = ~ iq,
family = gamlss.dist::BEOI, data = ReadingSkills,
control = gamlss::gamlss.control(trace = FALSE))
res <- clis_screen(fit, alpha = 0.1, seed = 1)
print(res)
}
Block decomposition of conformal normal curvature
Description
Decomposes the per-observation CNC scores into a contribution from the inflation-probability submodel and a contribution from the conditional-mean/precision submodel, exploiting the information orthogonality of the BIc model. The two contributions sum exactly to the total, with no cross terms.
Usage
cnc_block_decomp(delta_out, info_out)
Arguments
delta_out |
A perturbation-matrix list (e.g. from
|
info_out |
An information list from |
Value
A list with the discrete contribution B_gamma, the continuous
contribution B_betadelta, their sum B_total, and the discrete
fraction ratio_gamma.
Conformal normal curvature matrix
Description
Computes the Frobenius-normalised conformal normal curvature (CNC) matrix of Poon and Poon (1999) for a given perturbation matrix and information matrix inverse.
Usage
cnc_matrix(Delta, info_inv)
Arguments
Delta |
A perturbation matrix of dimension |
info_inv |
The inverse information matrix from |
Value
A list with the n x n CNC matrix B, its eigenvalues
lambda (normalised so that the sum of squares is one), eigenvectors
vectors, and the Frobenius norm normF of the unnormalised curvature.
References
Poon, W.-Y. and Poon, Y. S. (1999). Conformal normal curvature and assessment of local influence. Journal of the Royal Statistical Society: Series B, 61(1), 51-61.
Per-observation conformal normal curvature scores
Description
Computes the basic-perturbation CNC scores B_{E_t} and the aggregate
contribution measures m[r]_t from a fitted CNC matrix.
Usage
cnc_scores(cnc, r_max = 4L)
Arguments
cnc |
A list returned by |
r_max |
Integer; the maximum order of |
Value
A list with per-observation scores B_Et, the matrix of aggregate
contributions m_r (one column per r), the eigenvalue thresholds
threshold_r, the aggregate-contribution thresholds threshold_mt, the
number of r-influential eigenvectors k_r, and the cutoff b2 for
B_Et.
Linear-time per-observation CNC scores
Description
Computes the basic-perturbation CNC scores B_{E_t} without ever
forming the n \times n curvature matrix. The score is the scaled
diagonal of F_0 = \Delta^\top \mathcal{I}^{-1}\Delta, and both the
diagonal and the Frobenius norm \|F_0\|_F are obtained from
quantities of dimension (M+m+p), so the cost is linear in n and
the memory footprint is negligible. This is the scalable path used for
large samples, where the dense n\times n eigenproblem of
cnc_matrix() is infeasible.
Usage
cnc_scores_linear(Delta, info_inv)
Arguments
Delta |
A perturbation matrix of dimension |
info_inv |
The inverse information matrix from |
Value
A list with per-observation scores B_Et, the cutoff b2, the
Frobenius norm normF, and n.
Conformal p-values from non-conformity scores
Description
Given calibration scores (assumed to come from non-influential, "clean" observations) and test scores, computes marginal conformal p-values in the sense of Bates and others (2023). Larger scores indicate stronger evidence of influence, so the p-value for a test point is the calibrated rank of its score among the calibration scores.
Usage
conformal_pvalues(cal_scores, test_scores)
Arguments
cal_scores |
Numeric vector of calibration non-conformity scores. |
test_scores |
Numeric vector of test non-conformity scores. |
Details
For a test score s and calibration scores
c_1, \ldots, c_n, the conformal p-value is
p = \frac{1 + \#\{i : c_i \ge s\}}{n + 1}.
These p-values are marginally valid (super-uniform under the null that the test point is exchangeable with the calibration set) and, by the positive-dependence result of Bates and others (2023), permit Benjamini-Hochberg FDR control.
Value
A numeric vector of conformal p-values, one per test score.
References
Bates, S., Candes, E., Lei, L., Romano, Y. and Sesia, M. (2023). Testing for outliers with conformal p-values. The Annals of Statistics, 51(1), 149-178.
Examples
set.seed(1)
cal <- rnorm(100)
test <- c(rnorm(8), 4, 5) # last two are outliers
conformal_pvalues(cal, test)
Perturbation matrix for the case-weights scheme
Description
Perturbation matrix for the case-weights scheme
Usage
delta_caseweights(object)
Arguments
object |
A fitted BEZI/BEOI |
Value
A list with the full Delta matrix and its gamma, beta,
delta blocks.
Perturbation matrix for the discrete-covariate scheme
Description
Perturbation matrix for the discrete-covariate scheme
Usage
delta_disccovar(object, p = 1L)
Arguments
object |
A fitted BEZI/BEOI |
p |
Index of the perturbed covariate in the inflation design matrix. |
Value
A list with the full Delta and its blocks.
Perturbation matrix for the mean-covariate scheme
Description
Perturbation matrix for the mean-covariate scheme
Usage
delta_meancovar(object, p = 1L)
Arguments
object |
A fitted BEZI/BEOI |
p |
Index of the perturbed covariate in the mean design matrix. |
Value
A list with the full Delta and its blocks.
Perturbation matrix for the precision-covariate scheme
Description
Perturbation matrix for the precision-covariate scheme
Usage
delta_preccovar(object, p = 1L)
Arguments
object |
A fitted BEZI/BEOI |
p |
Index of the perturbed covariate in the precision design matrix. |
Value
A list with the full Delta and its blocks.
Simulated envelope for quantile residuals
Description
Simulated envelope for quantile residuals
Usage
envelope_bic(object, B = 200L, level = 0.95, seed = NULL)
Arguments
object |
A fitted BEZI/BEOI |
B |
Integer; number of simulated samples. |
level |
Numeric; envelope coverage (default 0.95). |
seed |
Optional integer seed. |
Value
Invisibly returns the observed residuals and envelope bands.
National DTP3 vaccination coverage (2022)
Description
Reads country-level proportions of diphtheria-tetanus-pertussis (DTP3)
vaccine coverage for 2022, with development covariates, from the CSV
shipped in the package's extdata directory. The response places a point
mass at one (countries achieving complete coverage), making it a natural
application of one-inflated beta (BEOI) regression.
Usage
load_vaccination()
Value
A data frame with one row per country and the columns
- iso3c
ISO 3166-1 alpha-3 country code.
- dtp3
DTP3 coverage proportion in
[0,1](response).- ln_gdp
Natural log of GDP per capita (PPP, constant USD).
- urb
Urbanisation rate (proportion of urban population).
- ln_pop
Natural log of total population.
- hdi
Human Development Index in
[0,1].
Source
WHO/UNICEF Joint Monitoring Programme (coverage); World Bank (GDP, urbanisation, population); UNDP Human Development Report (HDI).
Examples
vaccination <- load_vaccination()
summary(vaccination$dtp3)
mean(vaccination$dtp3 == 1) # fraction at the boundary
Plot a CLIS screening result
Description
Produces a two-panel diagnostic plot for a clis_screen() result: a plot
of conformal p-values (with the Benjamini-Hochberg rejection boundary) and
a plot of the non-conformity scores with the declared influential
observations highlighted.
Usage
plot_clis(x, label_top = 8L, ...)
Arguments
x |
A |
label_top |
Integer; how many top influential points to label. |
... |
Currently ignored. |
Value
Invisibly returns x.
Plot the classical CNC diagnostic panels
Description
Plot the classical CNC diagnostic panels
Usage
plot_cnc_panels(
cnc,
scores,
decomp = NULL,
r_plot = 3L,
label_top = 5L,
suffix = ""
)
Arguments
cnc |
A list from |
scores |
A list from |
decomp |
Optional list from |
r_plot |
Integer; which aggregate-contribution order to plot. |
label_top |
Integer; number of points to label. |
suffix |
Character; appended to panel titles. |
Value
Invisibly returns scores.
Classical local-influence index plot
Description
Reproduces the standard influence display of the beta-regression
diagnostics literature: the per-observation conformal normal curvature
B_{E_t} (or the aggregate contribution m[r]_t) plotted against
the observation index, with the reference cutoff drawn and the most
influential observations labelled. This is the plot practitioners of local
influence expect to see; the conformal screening in clis_screen() and
plot_clis() complements it with an error-controlled declaration.
Usage
plot_influence(
object,
scheme = "caseweights",
p = 1L,
measure = c("B_Et", "m_r"),
r_plot = 3L,
use_fisher = TRUE,
penalised = FALSE,
label_top = 5L,
labels = NULL
)
Arguments
object |
A fitted BEZI/BEOI |
scheme |
Perturbation scheme: "caseweights" (default), "disccovar", "meancovar", or "preccovar". |
p |
Covariate index for the covariate-perturbation schemes. |
measure |
Which quantity to plot: "B_Et" (default) or "m_r". |
r_plot |
Aggregate-contribution order when |
use_fisher |
Logical; use the Fisher information (default |
penalised |
Logical; use the penalised information (default |
label_top |
Integer; number of most-influential points to label. |
labels |
Optional character vector of observation labels (for example country codes); defaults to the observation index. |
Value
Invisibly, a list with the plotted scores and the cutoff.
Diagnostic residual plots for a BIc model
Description
Diagnostic residual plots for a BIc model
Usage
plot_residuals(object, label_top = 3L)
Arguments
object |
A fitted BEZI/BEOI |
label_top |
Integer; number of extreme points to label. |
Value
Invisibly returns a list of residual vectors.