A little announcement. A lot to explore.

Happy learning!

Explore lessons, including SAS & R programs and practice data.

Search lessons

Type at least 3 characters to see matching lessons.

Open SAS ↗Open R ↗
← SASnRWorking with missing values (NA) in R

Introduction

  • Missing values are pervasive in clinical data — a visit not done, a lab not drawn, a code reported as “UNK” that should be NA. This lesson is the R-only deep dive on handling them end to end.
  • Seven blocks, working from foundations to real-world patterns: (1) what NA is and how to check for it, (2) how NA behaves in arithmetic and logical ops, (3) how to assign NA — one cell, a logical mask, a sentinel sweep across many columns, (4) how to replace NA with a value, a fallback chain, or a carried-forward observation, (5) how to count NAs per column, per row, and across a chosen subset of variables, (6) how to drop rows with NAs, and (7) how NA is introduced when datasets are joined together.
  • We use one tiny clinical-flavoured dataset throughout — demo with subject id, demographics, and baseline/visit scores — with NAs scattered through several columns so every scenario has something to work on.
  • Dataset-first style. In the tidyverse variant, almost every scenario operates on demo through a pipe and stores the result as a tibble — not as a bare vector or scalar. Pure language-behaviour demos that have no natural data-frame form (the typed NAs, the four short-circuit cases) are wrapped into small reference tibbles so the result of every scenario is still inspectable as a dataset.
  • Naming convention for this lesson. Every scenario stores its result in a dataset named scenario01, scenario02, … numbered sequentially across all seven blocks. After running each line, type the scenario name on its own to print and confirm the changed cells before moving on. Two language-behaviour demo lines (the filter(Sex == NA) trap in scenario05 and the no-na.rm summarise(mean_weight = mean(Weight)) in scenario06) are deliberately left unassigned — they exist to be seen interactively right before the correct version.
  • This is an R-exclusive lesson by design. The base R variant covers the same scenarios using base R’s vector and bracket-subscript idioms so you can see the tidyverse helpers side-by-side with what they replace.

Create sample dataset with missing values

  • Create the small clinical-flavoured demo dataset used by every later missing-value scenario.
  • Run this first so each block can be tested independently without rebuilding the whole lesson.

Block 1: NA basics

  • Start with the core missing-value checks: typed NAs, column flags, full-frame checks, and the common == NA trap.
  • Print each scenario and confirm which cells or rows are marked missing.

Block 2: NA in computation

  • Observe how missing values propagate through arithmetic and logical expressions.
  • The goal is to see when NA is sticky and when logical short-circuiting can decide the answer anyway.

Block 3: Assigning NA

  • Practice intentionally assigning missing values at one cell, by a logical condition, and across multiple columns.
  • Pay close attention to typed NA constants because strict functions care about the target column type.

Block 4: Replacing NA

  • Replace missing values with defaults, fallback columns, and carried-forward values.
  • These examples are common when preparing analysis-ready clinical datasets.

Block 5: Counting NA

  • Count missingness by column, by row, by a selected subset of variables, and across the whole frame.
  • Use these checks as quick data-quality summaries before modeling or reporting.

Block 6: Dropping rows with NA

  • Drop rows with missing values using broad and targeted rules.
  • Confirm the row counts after each step because complete-case filtering can remove far more data than expected.

Block 7: NA introduced across datasets

  • Show how joins introduce missing values when a subject has no matching row in the lookup dataset.
  • Compare the left-join plus is.na() pattern with the direct unmatched-record pattern.