washr turns a cleaned dataset into a documented R data package that follows the FAIR principles, with a website, a citation file and machine readable metadata. This page walks through the workflow in the order you run it. Each step is one function, and every function is safe to run again after you have edited the files it wrote.
The publishing guide covers the same steps with more explanation and the parts that happen outside R, such as creating the GitHub repository and the Zenodo release. The chapters of the guide follow the same order as the sections below.
From CRAN:
Or the development version from GitHub:
Create an R package for the dataset and open it as a project. The package name is the name of the dataset, in lower case, without dots or underscores.
Then put it under version control and on GitHub, for example with
usethis::use_git() and usethis::use_github().
The guide describes this step in detail.
setup_ci() adds the GitHub Actions workflow that runs
R CMD check on every push and pull request, on macOS,
Windows and three versions of R on Linux. The openwashdata review
standard requires it, so run it now and the package meets that part of
the standard from its first commit.
Run setup_rawdata() from the package root. It creates
data-raw/ and a processing script,
data-raw/data_processing.R, with the steps the script
should follow.
Copy the raw files into data-raw/. In the processing
script, read them, clean them into one or more tidy data frames, and end
with usethis::use_data() so each data frame lands in
data/ as an .rda file. The script also exports
each data frame as CSV and XLSX into inst/extdata/, which
is where the website and the metadata find the downloadable files.
setup_dictionary() reads every data object in
data/ and writes data-raw/dictionary.csv with
one row per variable: the file it belongs to, its name, its type, and an
empty description.
Open the CSV and write a description for every variable. The dictionary is the one place where variables are described; the roxygen documentation and the metadata are generated from it.
setup_roxygen() writes one file per data object under
R/, with a title and description placeholder and the
variable table from the dictionary.
Open each file, replace the title and the description, then run
devtools::document() so the help pages exist. When you
change a description in the dictionary later, run
setup_roxygen() again: it regenerates only the variable
table and keeps the title, the description and anything else you wrote
in the file.
update_description() sets the fields the openwashdata
standard expects: the language, the date, the repository URL, the bug
report URL, and the CC BY 4.0 license when the package has none yet.
Existing values are kept and merged.
Then open DESCRIPTION and check the title, the
description and the authors. Authors are written with
person() and carry their ORCID in the comment field. Three
more facts have no other home and go into DESCRIPTION by
hand:
X-schema.org-keywords: sanitation, faecal sludge, Kampala
X-schema.org-spatialCoverage: Kampala, Uganda
X-schema.org-temporalCoverage: 2022-03-01/2022-09-30
The keywords are a comma separated list. The spatial coverage is a place name. The temporal coverage is a start and an end date separated by a slash.
update_metadata() derives a schema.org description of
the dataset from DESCRIPTION, the dictionary, the citation
file and the files in inst/extdata, and writes it as
JSON-LD into pkgdown/templates/in-header.html. The website
build embeds it in the head of every page, where dataset search engines
read it.
The function ends with the fields it could not fill and where to fill them. Nothing is edited by hand: change the source and run it again. This function is experimental; its output shape may still change.
setup_readme() writes README.Rmd from the
openwashdata template, with the installation instructions, the download
table, and a section for the first data object. With
has_example = TRUE it adds an Example section with a
scaffold for a first plot.
Write the text, add a section for each further data object, then render it:
setup_website() writes _pkgdown.yml from
the openwashdata template and builds the site into docs/.
The template sets the Pages URL, the analytics header, the funding
sidebar, and a reference index with one entry per data object.
Running it again keeps your _pkgdown.yml and rebuilds
the site. docs/ is committed and served by GitHub Pages;
when the package deploys through a pkgdown GitHub Actions workflow
instead, the function leaves docs/ ignored.
use_brand() installs the openwashdata brand, the fonts
and colors of the site, from the central brand repository and wires
_pkgdown.yml to it. Run it once after
setup_website(), and again whenever the brand changes.
update_citation() writes CITATION.cff and
inst/CITATION from DESCRIPTION. Run it before
the first release so the files exist, then again with the DOI once the
release is deposited on Zenodo; the guide describes the Zenodo
steps.
With a DOI the function also adds the DOI badge to the README and
rebuilds it and the site. A later run without the DOI keeps the DOI on
file. Run update_metadata() once more after that, so the
metadata carries the DOI as well.
Every function above reads what is there, merges its changes, and
writes only what changed. The two exceptions stop instead of
overwriting: setup_dictionary() when the dictionary exists,
and setup_readme() when README.Rmd exists
(pass force = TRUE to replace it). A full second run of the
other steps on an unchanged package changes no file.