Saving your code as custom functions and then distributing them in a package has real payoffs. It makes your code easier for others to use, it lets you reuse the same code across projects without copying and pasting, and it lets you shape how the people around you work. If you build a ggplot theme that follows good design principles and put it in a package, for example, you have given your colleagues an easy way to follow those principles too. This page walks the whole path, from a first function to an installable package, at overview altitude: enough of the pieces and the tool for each that you can direct an AI assistant through the details.
Creating your own functions
Hadley Wickham, who developed the {tidyverse}, recommends turning code into a function once you have copied it three times. Functions have three pieces: a name, a body, and arguments.
Writing a simple function
I will start with a relatively simple function. This one,
show_in_excel_penguins(), opens the penguins data in Microsoft Excel
(or whatever program your computer uses to open CSV files):
library(tidyverse)
library(fs)
penguins <- read_csv("https://data.rfortherestofus.com/penguins-2007.csv")
show_in_excel_penguins <- function() {
csv_file <- str_glue("{tempfile()}.csv")
write_csv(
x = penguins,
file = csv_file,
na = ""
)
file_show(path = csv_file)
}
This loads the {tidyverse} (to build a filename and save the CSV) and
{fs} (to open it). The assignment operator (<-) together with
function() tells R that show_in_excel_penguins is a function, not a
variable. The open curly brace begins the function body, which does
three things: it builds a temporary file path with str_glue() and
tempfile(), writes the penguins data there with write_csv() (setting
na = "" so missing values show as blanks rather than the text NA),
and opens that file with file_show(). From now on, running
show_in_excel_penguins() opens the penguins data in Excel.
Adding arguments
That function only ever opens one data frame, which is not very useful.
A more practical version would open any data. Items listed inside the
parentheses of the function definition are its arguments, and adding
a data argument does exactly that:
show_in_excel <- function(data) {
csv_file <- str_glue("{tempfile()}.csv")
write_csv(
x = data,
file = csv_file,
na = ""
)
file_show(path = csv_file)
}
The only changes are function(data) on the first line and x = data
inside write_csv(). Now I tell the function what data to use, and it
works with anything, including at the end of a pipeline:
covid_data |>
filter(state == "California") |>
show_in_excel()
Writing the function once means I never have to rewrite that code again, no matter how many times I want to view data in Excel.
A more realistic example: race and ethnicity data
Here is a function that simplifies real work. When you use the
{tidycensus} package to pull data from the US Census Bureau, the
variables have nonintuitive names that are hard to remember. If I
regularly want race and ethnicity data from the American Community
Survey, I can wrap that lookup in a get_acs_race_ethnicity() function
so I never have to remember the codes again:
library(tidycensus)
get_acs_race_ethnicity <- function() {
race_ethnicity_data <- get_acs(
geography = "state",
variables = c(
"White" = "B03002_003",
"Black/African American" = "B03002_004",
"American Indian/Alaska Native" = "B03002_005",
"Asian" = "B03002_006",
"Native Hawaiian/Pacific Islander" = "B03002_007",
"Other race" = "B03002_008",
"Multi-Race" = "B03002_009",
"Hispanic/Latino" = "B03002_012"
)
)
race_ethnicity_data
}
The body calls get_acs() to retrieve state-level population data, but
swaps the cryptic variable codes for human-readable names like White and
Black/African American, then returns the result. Running
get_acs_race_ethnicity() now gives back data with readable group
names.
I could improve this. Maybe I want the option to clean the variable
names into snake case with clean_names() from the {janitor} package,
but keep the original names available too. I add a
clean_variable_names argument, give it a default of FALSE, and
use an if statement to act on it:
get_acs_race_ethnicity <- function(clean_variable_names = FALSE) {
race_ethnicity_data <- get_acs(
geography = "state",
variables = c(
"White" = "B03002_003",
"Black/African American" = "B03002_004",
"American Indian/Alaska Native" = "B03002_005",
"Asian" = "B03002_006",
"Native Hawaiian/Pacific Islander" = "B03002_007",
"Other race" = "B03002_008",
"Multi-Race" = "B03002_009",
"Hispanic/Latino" = "B03002_012"
)
)
if (clean_variable_names == TRUE) {
race_ethnicity_data <- clean_names(race_ethnicity_data)
}
race_ethnicity_data
}
Because the default is FALSE, calling the function unchanged behaves
as before. Calling get_acs_race_ethnicity(clean_variable_names = TRUE)
returns the same data with GEOID and NAME lowercased to geoid and
name.
Using ... to pass arguments along
So far this function always retrieves state-level data, because
geography = "state" is hard-coded. What if I want county or census
tract data, or a specific year? I could add an argument for each, but
get_acs() has many arguments, and repeating them all would be tedious.
The ... (“dots”) syntax is the efficient option. Putting ... in
my function and passing it straight to get_acs() forwards along any
argument the caller supplies:
get_acs_race_ethnicity <- function(clean_variable_names = FALSE, ...) {
race_ethnicity_data <- get_acs(
...,
variables = c(
"White" = "B03002_003",
"Black/African American" = "B03002_004",
"American Indian/Alaska Native" = "B03002_005",
"Asian" = "B03002_006",
"Native Hawaiian/Pacific Islander" = "B03002_007",
"Other race" = "B03002_008",
"Multi-Race" = "B03002_009",
"Hispanic/Latino" = "B03002_012"
)
)
if (clean_variable_names == TRUE) {
race_ethnicity_data <- clean_names(race_ethnicity_data)
}
race_ethnicity_data
}
Now get_acs_race_ethnicity(geography = "county") gets county data, and
get_acs_race_ethnicity(geography = "county", geometry = TRUE) returns
geospatial data alongside the demographics, all without my listing those
arguments myself. The dots let a wrapper stay as capable as the function
underneath it while keeping my own code concise.
Creating a package
Packages bundle your functions so you can use them across projects. If
you find yourself copying functions from one project to another, or
pasting a functions.R file into each new project, that is a strong
sign you should make a package. Running functions from a loose
functions.R file works in your own environment, but it may not work on
someone else’s computer: they might not have the packages your code
depends on, or they might not know how your arguments work. A package
solves both, because it declares its dependencies and carries its own
documentation.
Starting the package
The {usethis} and {devtools} packages automate nearly every step of building a package, so install them first:
install.packages("usethis")
install.packages("devtools")
The most reliable way to scaffold a package, and the one I use in
Positron, is a single console command. usethis::create_package()
builds the standard package structure (the R/ folder, a DESCRIPTION
file, and the roxygen setup) at the path you give it, and the last part
of the path becomes the package name:
usethis::create_package("~/Documents/dk")
Run interactively, it opens the new package as its own workspace in
Positron. (If you prefer a GUI, Positron’s Workspaces: New Folder from
Template command, reachable from the Command Palette or the New
menu, will create and open an R project folder for you, but the package
skeleton itself still comes from create_package().)
Adding functions with use_r()
Every function in a package lives in its own file in the R folder.
Create a file for your function with use_r(), passing a name that
hints at what the file holds:
usethis::use_r("acs")
This creates R/acs.R. Note the package::function() form, which lets
you call a function without loading its package. Open the new file and
paste in your get_acs_race_ethnicity() function.
Checking the package with devtools::check()
You need to make a couple of changes for the function to work inside a
package, and the easiest way to find out what is to let R tell you. Run
devtools::check(), which performs an R CMD check to confirm others
can install your package. On the {dk} package it prints a long report
ending with:
Undefined global functions or variables:
clean_names get_acs
0 errors ✔ | 2 warnings ✖ | 1 note ✖
Read that summary from the bottom up. Errors mean others cannot install your package and matter most, while warnings and notes may still cause problems. The goal is zero of all three.
Adding dependency packages
The note above appears because the function uses get_acs() and
clean_names() without saying where they come from. Those live in
{tidycensus} and {janitor}, and your package has to guarantee they get
installed for anyone who installs {dk}. Declare each with
use_package():
usethis::use_package("tidycensus")
usethis::use_package("janitor")
Each call adds the package to the Imports field of the DESCRIPTION
file, the file that holds your package’s metadata. Now anyone who
installs {dk} gets {tidycensus} and {janitor} too.
Referring to functions correctly
Running those commands also prints a reminder: refer to functions with
package::fun(). Inside a package you cannot assume library() has
been called, so you name both the package and the function to be sure
the right one is used (occasionally two packages share a function name,
and this removes the ambiguity). That is what clears the note. Update
the function so get_acs() becomes tidycensus::get_acs() and
clean_names() becomes janitor::clean_names():
get_acs_race_ethnicity <- function(clean_variable_names = FALSE, ...) {
race_ethnicity_data <- tidycensus::get_acs(
...,
variables = c(
"White" = "B03002_003",
"Black/African American" = "B03002_004",
"American Indian/Alaska Native" = "B03002_005",
"Asian" = "B03002_006",
"Native Hawaiian/Pacific Islander" = "B03002_007",
"Other race" = "B03002_008",
"Multi-Race" = "B03002_009",
"Hispanic/Latino" = "B03002_012"
)
)
if (clean_variable_names == TRUE) {
race_ethnicity_data <- janitor::clean_names(race_ethnicity_data)
}
race_ethnicity_data
}
Run devtools::check() again and the note is gone, leaving two warnings
to handle.
Creating documentation with roxygen2
One warning is about missing documentation. A big benefit of a package
is that users can type ?get_acs_race_ethnicity() and get help, the
same way they can for any built-in function. That help comes from
Roxygen, powered by the {roxygen2} package. You write a comment block
directly above the function where every line starts with #' (many
editors can insert a skeleton for you to fill in):
#' Access race and ethnicity data from the American Community Survey
#'
#' @param clean_variable_names Should variable names be cleaned (i.e. snake case)?
#' @param ... Other arguments passed to tidycensus::get_acs()
#'
#' @return A tibble with five variables: GEOID, NAME, variable, estimate, and moe
#' @export
get_acs_race_ethnicity <- function(clean_variable_names = FALSE, ...) {
# ... function body ...
}
@param documents each argument, @return says what comes back, and
@export makes the function available to users of your package
(internal helper functions leave @export off). Then run
devtools::document(), which generates a read-only .Rd help file in
the man directory and a NAMESPACE file listing your exported
functions.
Adding a license and metadata
Run devtools::check() once more and the documentation warning is gone,
but one remains: your package has no license. If you plan to share it, a
license tells people what they may do with your code (see
https://choosealicense.com to choose one). I will use the permissive
MIT license:
usethis::use_mit_license()
This sets the License field in DESCRIPTION and writes the license
files for you. While you are in DESCRIPTION, fill in the title,
author, maintainer, and description so people know what the package is
for. Run devtools::check() a final time and you should see:
0 errors ✔ | 0 warnings ✔ | 0 notes ✔
That is exactly what you want.
Adding more functions
To add another function, repeat the same loop:
- Create a new
.Rfile withusethis::use_r(), or add to an existing one. - Write the function, using
package::function()to refer to functions from other packages. - Declare any new dependencies with
usethis::use_package(). - Document the function with Roxygen.
- Run
devtools::check()to confirm everything is in order.
A package can hold a single function or as many as you like.
Installing the package
To use the package yourself, run devtools::install(), and it is
available in any project like any other package. To share it with
others, the usual route is GitHub. I pushed {dk} to
https://github.com/dgkeyes/dk; with the {remotes} package installed,
anyone can install it with:
remotes::install_github("dgkeyes/dk")
Where to go deeper
- Hadley Wickham and Jenny Bryan, R Packages (2nd ed.): https://r-pkgs.org
- Jenny Bryan, Happy Git and GitHub for the useR, for putting your package on GitHub: https://happygitwithr.com