A little announcement. A lot to explore.

Happy learning!

Explore lessons, including SAS & R programs and practice data.

Search lessons

Type at least 3 characters to see matching lessons.

Open SAS ↗Open R ↗
← SASnRAdvanced variable selection in tidyverse

Introduction

  • Lesson 1300 Subset variables introduced the basics of select() — bare names, minus signs, starts_with(), ends_with().
  • The tidy-select mini-language goes much further: type-based selection with where(), position helpers like everything() and last_col(), regex selection with matches(), lenient vs strict character-vector selection via any_of() / all_of(), and set operators (&, |, !) that combine all of the above.
  • This lesson walks twelve progressively richer scenarios on a single extended CLASS dataset (we added subj_id, three score columns, and an enrolled logical) so every helper has something to grab.
  • Naming convention for this lesson. Every scenario stores its result in a dataset named scenario01, scenario02, … When a single concept needs two or more related expressions (e.g. all_of vs any_of), we use suffixed names: scenario06a, scenario06b. Run the line, then type the scenario name on its own to print and inspect the result.
  • The final scenario is a teaser: the same tidy-select vocabulary you learn here is what powers across() inside mutate() and summarise() — one language, used everywhere a dplyr/tidyr verb takes a column-spec argument.
  • This lesson has no SAS counterpart by design — it is an R-only deep dive into tidy-select. A base R variant is included to show the contrast: most of what is one short helper in tidyverse is a grepl() + sapply() + bracket-subscript expression in base.

Create the selection practice dataset

  • Build the same extended CLASS dataset before comparing tidyselect and base R selection patterns.
  • The extra score and logical columns are intentional, so type-based and pattern-based selectors have something realistic to work on.

Select, drop, and reorder columns

  • Start with the everyday cases: keep a contiguous range, drop known variables, and move one important variable to the front.
  • Run each scenario and confirm the output columns are in the expected order.

Use right-edge, numbered, and named-list selectors

  • These examples cover selectors that depend on column position, repeated numbered names, or a character vector of requested variables.
  • Compare strict and lenient name matching carefully because that is the key behavioral difference in this block.

Select by pattern and column type

  • Move from explicitly named columns to rule-based selection using regular expressions and type predicates.
  • These are the patterns that scale when incoming datasets have many columns.

Combine selectors, rename, and reuse selection inside operations

  • Finish by combining selection rules, renaming as part of selection, and reusing the same selection language inside column-wise operations.
  • This is where tidyselect becomes more than a column picker: it becomes a reusable grammar for transformations.