3.4 Code Speed and Benchmarking
3.4 Code Speed and Benchmarking
Learning objectives
By the end of this chapter, you can:
- measure credible package-level comparisons with
bench::mark()(median,itr/sec,mem_alloc). - rank and explain three common sources of slowness: growing objects in loops, row-wise calculations, and repeated file I/O.
- explain copy-on-modify through a live
tracemem()demonstration, revisiting chapter 1.3. - compare data.table and dplyr for a specific situation rather than choose a camp.
- locate real hotspots in package functions using profvis.
- evaluate parallelism’s benefits and overhead for a particular task.
Prerequisite check (≤5 minutes)
Complete chapter 1.3, covering copy-on-modify and a first benchmark, before continuing.
- You run
y <- x, then calltracemem(y), then runy[[1]] <- 9. Which step prints a memory-copy message? - Guess the speed ratio between
rowwise() |> mutate(z = max(a, b))andmutate(z = pmax(a, b)). Write down your estimate and compare it in §2.
1. Measure before optimizing: bench in a package context
Chapter 1.3 developed intuition. Here the focus is package development: evidence should withstand questions from colleagues and reviewers. Put benchmarks in PRs and vignettes, not just chat screenshots. bench::mark() (https://bench.r-lib.org) remains our measuring instrument:
library(bench)
scores <- runif(1e5, min = 0, max = 100)
mark(
cut_default = as.character(cut(scores, c(0, 60, 70, 80, 90, 100),
labels = c("F", "D", "C", "B", "A"),
right = FALSE, include.lowest = TRUE)),
nested_ifelse = ifelse(scores >= 90, "A", ifelse(scores >= 80, "B",
ifelse(scores >= 70, "C", ifelse(scores >= 60, "D", "F")))),
iterations = 50
)Read median (typical time), itr/sec (throughput), and mem_alloc (memory). Three rules: ① Focus on relative ratios; absolute times vary by machine. ② The default check = TRUE verifies equivalent results. Comparing how quickly two different answers are calculated is an expensive benchmarking mistake. ③ Change one variable at a time.
Much allegedly slow R code is another language’s approach translated literally. Start by thinking in R: vectors and whole columns. Most remaining problems worth your effort fall into the three categories below.
2. Three common sources of slowness
| Rank | Source | Mechanism | Short remedy |
|---|---|---|---|
| 1 | Growing objects in loops: out <- c(out, x) |
Allocate a longer vector and copy each time, roughly n²/2 total work | Preallocate or accumulate vectorially |
| 2 | Row-wise calculations | Pay R interpreter overhead for every row | Ask for a whole-column operation: pmax, rowSums, ifelse |
| 3 | Repeated file I/O | Millisecond disk access is roughly 10⁵ times slower than memory; often combined with repeated rbind | Read once, read in batches, or scan lazily |
The standard comparison for source 2 also checks your prerequisite prediction:
library(dplyr)
df <- tibble(a = runif(1e5), b = runif(1e5))
mark(
rowwise_max = df |> rowwise() |> mutate(z = max(a, b)) |> ungroup(),
vectorized = df |> mutate(z = pmax(a, b)),
iterations = 20
) # Typical difference: 50–150 timesA common instance of source 3, with two alternatives:
# Slow: read one file and rbind at a time (sources 3 and 1 together)
out <- NULL
for (f in files) {
d <- utils::read.csv(f)
out <- rbind(out, d)
}
# Faster: batch reading, or an Arrow lazy scan for data larger than memory
out <- vroom::vroom(files)
# out <- arrow::open_dataset("data/scores_dir/")Move I/O out of loops whenever possible. When that is impossible, as with logging, accumulate records and write batches rather than one line at a time.
read.csv() returns a data.frame and vroom() a tibble, so mark() may report unequal results. Only when you have verified different containers, same contents should you use check = FALSE. Explain the reason in your report, as in chapter 1.3.
3. tracemem: a one-minute copy-on-modify refresher
Chapter 1.3 provides the full explanation. Replay the key observation: modifying a named object can involve copying it.
df <- data.frame(x = 1:5)
tracemem(df)
df$y <- 2 # A copy message when modifying the named object
df$z <- df$x * 2 # Another modification
untracemem(df)Two implications for packages: ① Receiving a large object as an argument does not itself copy it. ② Repeatedly modifying a named large object in a loop can produce costly copying, the underlying concern in source 1. For modification by reference, data.table’s := is a prominent alternative, introduced next.
4. data.table versus dplyr: choose honestly
library(dplyr); library(data.table)
flights <- nycflights13::flights
# dplyr: filter, group, summarize; compose verbs step by step
flights |> filter(month == 1) |>
group_by(carrier) |> summarise(delay = mean(dep_delay, na.rm = TRUE))
# data.table: rows where month == 1; calculate delay for each carrier
as.data.table(flights)[month == 1,
.(delay = mean(dep_delay, na.rm = TRUE)),
by = carrier]| Dimension | dplyr | data.table |
|---|---|---|
| Mental model | Composable pipeline verbs | One DT[i, j, by] expression |
| Speed/memory | Sufficient for everyday in-memory analysis | Fast grouping and large-table work; := reduces copying |
| Ecosystem | Integrates with tidyverse | Self-contained, no dependencies |
| Learning curve | Gentle | Steeper, but concise once familiar |
Three practical considerations: ① What the team knows often matters more than raw speed; people must read the code. ② With truly large data, I/O and query design (1.3 §6) often dominate verb overhead. ③ Choose package dependencies carefully: do not import both libraries for one small function.
① Group 100 million rows with limited memory. ② A team of tidyverse users analyzes less than 1 GB. ③ A small package only cleans strings. Choose dplyr, data.table, or “either” for each, and justify it in one sentence using §4.
5. profvis: examine a package’s pulse
A benchmark asks how much faster A is than B. profvis (https://rstudio.github.io/profvis/) asks where the time goes. It applies directly to package functions:
library(profvis)
profvis({
dat <- vroom::vroom("scores_big.csv")
grade <- scorekit::grade_letter(dat$score)
table(grade)
})Read the display in three ways: ① Wide flame-graph bars identify hotspots to investigate. ② Drill into per-line time and memory in the Data view. ③ Run in an interactive session; knitting does not open the viewer. Include a flame-graph screenshot in performance PRs so a fivefold slowdown is visible before merging.
The full cycle remains: find the hotspot with profvis → change that section → compare before and after with mark(). Changing code without measurement is simply a different form of guessing.
6. A brief introduction to parallelism
Consider parallel work when independent tasks take seconds, such as parsing 500 files. future chooses the execution strategy; furrr supplies parallel map operations:
library(furrr)
plan(multisession, workers = 4)
results <- future_map(files, ~heavy_parse(.x)) # Preserve result orderThree cautions: ① Process startup and data transfer cost time; millisecond tasks may become slower. ② Vectorized code already running efficiently in C may gain little. ③ Manage random seeds; future provides dedicated mechanisms. As a beginner’s rule of thumb, if a single-call median is below one second, look for vectorization before parallelism.
Rerun the §1 and §2 benchmarks: cut versus nested ifelse, and rowwise versus pmax. Record median, itr/sec, and mem_alloc in a table. Repeat ten minutes later or on another machine. In two lines, distinguish stable relative ratios from variable absolute times. Was your prerequisite prediction correct?
Generate 50 small CSVs of about 2000 rows each. Compare looping read.csv() plus rbind(), vroom::vroom(files), and arrow::open_dataset(). Handle check as discussed in §2 and state how results were aligned. Use fewer iterations, such as 5. Which is fastest, and which uses least memory?
Round 1 (AI off): Choose your package’s heaviest function, or the class repository’s profile-me.R, which contains all three slow patterns. Find hotspots with profvis → diagnose using §2 → rewrite the largest bottleneck → compare before and after with mark(), including mem_alloc. Round 2 (AI allowed): Give Posit Assistant both versions and a profiling summary. Ask only: “Which hidden copies, row-wise operations, or repeated I/O remain?” Record and fix one issue that you confirm by measurement.
Capstone
Task: “Performance dossier 1.0.” Choose real slow code, or use the class repository’s slow-registry.R, containing all three patterns. Submit a Quarto audit: flame graph with hotspots marked → diagnosis using the ranking → optimize only the most expensive item → before/after benchmark table with median and mem_alloc → selection rationale. Why use or avoid data.table? Why is parallelism worthwhile or not? Cite §4 and §6 criteria.
| Dimension | Meets expectations | Good | Excellent |
|---|---|---|---|
| Measurement | Before/after benchmark | Median, fixed iterations, repeated check | Machine/workload differences and applicability stated |
| Diagnosis | Identifies a bottleneck | profvis evidence and classified cause | Rules out a plausible but insignificant suspect |
| Verification | Measured improvement | Explains copying/interpreter/I/O costs | Finds other instances of the same pattern |
| Honest selection | Includes rationale | Uses §4/§6 criteria rather than slogans | Discusses readability, team, dependencies, and why to stop |
SOURCES · Attribution
| Section | Material | Use |
|---|---|---|
| Structure, slow-pattern ranking, selection guidance, exercises, capstone, rubric, and division of topics with 1.3 | This project | Original |
bench::mark() and output interpretation |
Official bench documentation, https://bench.r-lib.org | Reference |
| profvis | Official documentation, https://rstudio.github.io/profvis/ | Reference |
| data.table / dplyr comparison | Official Getting started documentation and vignettes | Reference |
This chapter is published under CC-BY-SA 4.0.