Choosing and combining backbones

This vignette shows how to pick a taxonomic backbone for a name list, how to chain several backbones so that each name is resolved by the source best placed to judge it, and how to check afterwards which source resolved which name. taxify matches names against locally stored Darwin Core backbone databases. 19 backbones are available, each compiled from a different authoritative source. The backbone we choose determines which names can be matched, which taxonomic opinion governs synonym resolution, and which extra metadata columns are available downstream.

  1. Choose the backbones that cover the taxonomic scope of the list (see The backbones); list_backbones() reports each one’s current version, name count, and whether it is installed.
  2. Download them ahead of time, or pin a version, with taxify_download().
  3. Match against a single backbone with the backbone argument of taxify().
  4. Chain several backbones in fallback order by passing a vector to backbone.
  5. Audit which backbone matched each name through the backbone and backbone_version columns.
  6. Extend the result with backbone-specific columns from add_wfo_info(), add_col_info(), and add_gbif_info().
  7. Diagnose names outside a backbone’s scope with lookup_genus() and taxify_register_coverage().

Example

library(taxify)

Downloading backbones

taxify downloads a backbone on first use. When taxify(names, backbone = "wfo") finds no local WFO backbone, it fetches the pre-built .vtr file from the taxifydb release it is published under, writes it to taxify_data_dir(), and caches the path for the rest of the R session. Later calls, in the same session or in future ones, reuse the local copy without network access. taxify_download() fetches backbones ahead of time, which is useful on a shared server or in a Docker image:

# Download one backbone
taxify_download("wfo")

# Download several at once
taxify_download(c("wfo", "col", "worms"))

Pre-built .vtr files are published as GitHub Releases on taxifydb, one tag per backbone (gbif-2026.06, algaebase-2026.06), and range from 8 MB (AviList) to 2.0 GB (COL). They are compiled from the raw Darwin Core sources with precomputed matching keys, embedded synonym resolution, and genus-level indexes, so they can be queried as soon as the download completes.

A specific backbone version can be pinned:

taxify_download("wfo", version = "2024.01")

Pinned versions are stored in their own directory (taxify_data_dir()/wfo/2024.01/) and are never overwritten by future updates. The “latest” slot (taxify_data_dir()/wfo/latest/) is overwritten whenever a newer version becomes available. A project that locks a backbone version produces identical results regardless of when the analysis is re-run.

Matching against one backbone

The simplest case matches plant names against WFO:

plants <- c(
  "Quercus robur",
  "Quercus petraea",
  "Pinus sylvestris",
  "Acer pseudoplatanus",
  "Betula pendula",
  "Fagus sylvatica",
  "Picea abies"
)

result <- taxify(plants, backbone = "wfo")
result[, c("input_name", "accepted_name", "family", "match_type", "backbone")]

Every row in the output has 27 columns regardless of which backbone produced it: input_name, matched_name, accepted_name, taxon_id, accepted_id, rank, family, genus, epithet, authorship, accepted_authorship, is_synonym, taxonomic_status, is_hybrid, match_type, fuzzy_dist, n_ids, accepted_ids, backbone, backbone_version, kingdom_group, taxon_group, life_form, qualifier, qualifier_position, aggregate_fallback, and hybrid_type. The backbone column records "wfo" for matched rows and NA for unmatched ones. The backbone_version column records the backbone name, version, and download date (e.g., "wfo:2024-12 (2026-04-01)"), so we can cite the exact data snapshot used.

When a name matches a synonym, taxify resolves it to the accepted name. The matched_name column shows the string that matched in the backbone (which may be a synonym), accepted_name shows the current accepted name after resolution, and is_synonym is TRUE for resolved synonyms and FALSE for direct matches.

Fuzzy matching is on by default with a normalized Damerau-Levenshtein threshold of 0.2, roughly one edit per five characters, which catches typos such as transposed letters or missing diacritics. A lower threshold is stricter:

result <- taxify(plants, backbone = "wfo", fuzzy_threshold = 0.1)

fuzzy = FALSE turns fuzzy matching off and accepts only exact and case-insensitive matches:

result <- taxify(plants, backbone = "wfo", fuzzy = FALSE)

The fuzzy_method argument selects the distance metric: the default "dl" (Damerau-Levenshtein, which counts a transposition as one edit), "levenshtein" (standard Levenshtein, no transposition handling), or "jw" (Jaro-Winkler, which weights agreement in the first characters of the two strings).

Chaining backbones

A wetland monitoring dataset might contain vascular plants, invertebrates, amphibians, algae, and fungi, and no single backbone covers all of these equally well. Passing a vector of backbone names builds a fallback chain: names are matched against each backbone in order, and a name matched by an earlier backbone is removed from the pool and never re-matched by a later one.

mixed <- c(
  "Quercus robur",       # plant
  "Panthera leo",        # animal
  "Amanita muscaria",    # fungus
  "Salmo trutta",        # fish
  "Escherichia coli"     # bacterium
)

result <- taxify(mixed, backbone = c("wfo", "col", "gbif"))
result[, c("input_name", "accepted_name", "match_type", "backbone")]

The console output during matching shows the chain in action:

Matching 5 names against 3 backbones: wfo -> col -> gbif
  [wfo] Matching 5 names...
  [col] Matching 4 remaining names...
  [gbif] Matching 1 remaining names...

"Quercus robur" matches in WFO and leaves the pool. The remaining four names go to COL, and anything still unmatched after COL (perhaps an obscure bacterial name) goes to GBIF. The process continues until all names have been tried against all backbones or all names have matched. If earlier backbones have matched every name, later ones are skipped with a message:

  [gbif] Skipped (all names matched)

The order of the vector decides which taxonomic opinion wins for each name. If "Quercus robur" exists in both WFO and COL, putting WFO first means WFO’s accepted name, family assignment, and synonym resolution are used; putting COL first gives COL’s. For names that exist in several backbones, the first backbone in the chain always wins. With GBIF first, everything would match there (GBIF has ~6.4M names, the largest of any backbone) and the curated opinions of WFO, COL, or WoRMS would never be consulted. For a plant-heavy list with some non-plant taxa, c("wfo", "col") or c("wfo", "col", "gbif") gives WFO’s curated plant taxonomy for the plants while COL or GBIF picks up the rest.

Fuzzy matching runs independently within each backbone. A name that fails exact matching in WFO is fuzzy-matched against WFO, and only if that also fails does it move to the next backbone, where it is exact-matched and then fuzzy-matched again. A misspelled plant name therefore gets its best chance in WFO, the plant-specialist backbone, before falling through to COL or GBIF.

Plants: WFO with a COL fallback

WFO focuses on accepted vascular plant and bryophyte names. Names that appear only in older literature, belong to genera not yet integrated into WFO, or are nomenclaturally orphaned (no clear accepted name) may be absent. COL inherits WFO’s plant taxonomy as one of its sector databases and supplements it with names from other sources, including historical synonyms and cultivar names.

plants <- c(
  "Quercus robur",
  "Quercus petraea",
  "Pinus sylvestris",
  "Acer pseudoplatanus",
  "Coffea arabica",
  "Welwitschia mirabilis",
  "Lepidodendron aculeatum",   # extinct lycopsid
  "Nothofagus cunninghamii",
  "Dracaena draco"
)

# WFO alone
wfo_result <- taxify(plants, backbone = "wfo")
table(wfo_result$match_type)

If any names come back as "none", COL can be added as a fallback:

# WFO first, COL as fallback
both_result <- taxify(plants, backbone = c("wfo", "col"))
table(both_result$match_type)
both_result[, c("input_name", "accepted_name", "backbone")]

The backbone column now shows "wfo" for names matched by WFO and "col" for names only COL could resolve. A methods section can then state “plant names were resolved against WFO 2024-12, with unmatched names resolved against COL 2025.” In large vegetation plot datasets most names resolve in WFO, and the handful of edge cases (cultivars, historical names, genera recently moved between families) that fall through to COL would otherwise need manual resolution.

Mixed kingdoms: COL, GBIF and WoRMS

A monitoring dataset from a coastal estuary might contain vascular plants, invertebrates, fish, and marine algae. COL has broad expert-curated coverage across kingdoms, GBIF fills gaps with its larger name pool, including names from national checklists not yet incorporated into COL, and WoRMS is a final backstop for marine invertebrate synonyms, which can be slow to propagate to generalist databases. The chain runs from highest curation to broadest coverage, and most of these names resolve in COL.

estuary_species <- c(
  "Zostera marina",              # seagrass (plant)
  "Salicornia europaea",         # glasswort (plant)
  "Carcinus maenas",             # shore crab
  "Mytilus edulis",              # blue mussel
  "Platichthys flesus",          # European flounder
  "Nereis diversicolor",         # ragworm
  "Fucus vesiculosus",           # bladderwrack (brown alga)
  "Littorina littorea",          # common periwinkle
  "Arenicola marina",            # lugworm
  "Cerastoderma edule"           # common cockle
)

result <- taxify(estuary_species, backbone = c("col", "gbif", "worms"))
result[, c("input_name", "accepted_name", "family", "backbone")]

For a predominantly marine list with only a few terrestrial taxa, leading with WoRMS makes its marine taxonomy take precedence:

result <- taxify(estuary_species, backbone = c("worms", "col"))

Fungi: Species Fungorum with a COL fallback

Species Fungorum Plus is curated specifically for fungi, with ~315k names including anamorphs, teleomorphs, and the pleomorphic naming changes of the 2011 Melbourne Code. Its synonym coverage for fungal genera is better than in generalist databases, where fungal taxonomy is often a secondary concern.

fungi <- c(
  "Amanita muscaria",
  "Boletus edulis",
  "Cantharellus cibarius",
  "Tuber melanosporum",
  "Saccharomyces cerevisiae",
  "Aspergillus niger",
  "Penicillium chrysogenum",
  "Agaricus bisporus",
  "Trametes versicolor",
  "Cordyceps militaris"
)

result <- taxify(fungi, backbone = c("fungorum", "col"))
result[, c("input_name", "accepted_name", "is_synonym", "backbone")]

Species Fungorum resolves the standard names, and COL picks up any obscure, recently described, or historically orphaned species that fall through. For a list of fungi and plants, a three-backbone chain gives each group its specialist, with COL catching what falls through both:

mixed <- c(
  "Quercus robur",              # plant
  "Amanita muscaria",           # fungus
  "Lactarius deliciosus",       # fungus
  "Pinus sylvestris",           # plant
  "Russula emetica"             # fungus
)

result <- taxify(mixed, backbone = c("wfo", "fungorum", "col"))

Because the genus Amanita is not in WFO’s coverage table, "Amanita muscaria" is marked out of scope for WFO immediately and passed to the next backbone without fuzzy matching (see The genus register).

Algae: AlgaeBase

AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. Its curation is strong for freshwater and marine microalgae, where generalist databases often have thin coverage and outdated synonymy.

algae <- c(
  "Chlamydomonas reinhardtii",
  "Chlorella vulgaris",
  "Ulva lactuca",
  "Fucus vesiculosus",
  "Sargassum muticum"
)

result <- taxify(algae, backbone = c("algaebase", "col"))
result[, c("input_name", "accepted_name", "backbone")]

AlgaeBase is licensed CC BY-NC, and taxify prints a license notice during its download. For commercial applications, COL or WoRMS can serve instead, with less specialized algal coverage.

Molecular ecology: NCBI

Species lists from metabarcoding or eDNA studies are often linked to NCBI accessions. The NCBI backbone aligns taxify’s accepted names with the taxonomy used in GenBank and BOLD.

edna_hits <- c(
  "Salmo trutta",
  "Phoxinus phoxinus",
  "Anguilla anguilla",
  "Cottus gobio",
  "Lampetra planeri",
  "Chironomus riparius",     # midge (insect)
  "Potamopyrgus antipodarum" # New Zealand mud snail
)

result <- taxify(edna_hits, backbone = c("ncbi", "col"))
result[, c("input_name", "accepted_name", "taxon_id", "backbone")]

The taxon_id values of NCBI-matched rows are NCBI tax_ids, which link directly to GenBank records and NCBI taxonomy pages. COL covers names not found in NCBI, such as taxa without sequenced representatives.

Auditing the backbone column

backbone is a plain character column. In a single-backbone call every matched row shows the same name; in a chain it records which backbone produced each match; unmatched rows have backbone = NA. It can count how many names each backbone resolved:

result <- taxify(species_list, backbone = c("wfo", "col", "gbif"))
table(result$backbone, useNA = "ifany")

or filter to the rows matched by one backbone:

wfo_matches <- result[result$backbone == "wfo" & !is.na(result$backbone), ]
col_matches <- result[result$backbone == "col" & !is.na(result$backbone), ]

In a list expected to be purely plants, many names matched by COL instead of WFO point to names outside WFO’s scope, perhaps algae classified as plants in older literature, or animal-associated organisms such as plant parasites.

For a methods section, the unique backbone_version strings give the provenance of the whole result:

unique(result$backbone_version[!is.na(result$backbone_version)])
# e.g., c("wfo:2024-12 (2026-04-01)", "col:2025 (2026-04-01)")

Each string identifies both the taxonomic source and the snapshot used, so the result can be reproduced even if a backbone releases a new version between the analysis and a reviewer’s check.

Backbone-specific extras

Three backbones have enrichment functions that join extra backbone-specific columns to a taxify result. Each fills only the rows matched by its own backbone; rows from other backbones get NA in the new columns.

add_wfo_info() adds scientificNameID, parentNameUsageID, namePublishedIn, higherClassification, taxonRemarks, and infraspecificEpithet. namePublishedIn is useful for citing original descriptions, and higherClassification holds the full taxonomic hierarchy as a semicolon-separated string.

result <- taxify(plants, backbone = "wfo") |>
  add_wfo_info()

result[, c("input_name", "accepted_name", "namePublishedIn")]

add_col_info() adds COL classification columns (kingdom, phylum, col_class, order), nomenclatural metadata (notho, nomenclaturalCode, nomenclaturalStatus, namePublishedIn), infraspecificEpithet, and SpeciesProfile flags (is_extinct, is_marine, is_freshwater, is_terrestrial). The class column is renamed to col_class to avoid conflict with R’s class() function. The SpeciesProfile flags come from a separate file in the COL DwC-A archive and can, for instance, exclude extinct species from a contemporary biodiversity analysis or separate marine from terrestrial taxa in an estuarine dataset.

result <- taxify(species_list, backbone = "col") |>
  add_col_info()

# Check which species are marine
result[result$is_marine == TRUE & !is.na(result$is_marine),
       c("input_name", "accepted_name", "kingdom", "is_marine")]

add_gbif_info() adds notho_type (hybrid type), nom_status (nomenclatural status), bracket_authorship (basionym author), bracket_year, gbif_year, name_published_in, origin (how the name entered the GBIF backbone), and infra_specific_epithet. origin values such as "SOURCE", "DENORMED_CLASSIFICATION", or "VERBATIM_ACCEPTED" record how GBIF ingested the name.

result <- taxify(species_list, backbone = "gbif") |>
  add_gbif_info()

result[, c("input_name", "accepted_name", "origin", "nom_status")]

A multi-backbone result can pass through all three; each touches only rows from its own backbone:

result <- taxify(species_list, backbone = c("wfo", "col", "gbif")) |>
  add_wfo_info() |>
  add_col_info() |>
  add_gbif_info()

The result is a wide data.frame with the union of all extra columns, where each row has only its own backbone’s columns populated and the rest NA. For most workflows the base 27 columns are enough; the extras matter when an analysis needs nomenclatural details, habitat flags, or publication references that the standard output does not include.

Diagnosing out-of-scope names

lookup_genus() returns the genus register row for a single genus. The register is loaded into memory on the first call and cached for the session.

lookup_genus("Quercus")
#   genus   kingdom phylum class   order   family   life_form
# 1 Quercus Plantae ...    ...     Fagales Fagaceae vascular plant
lookup_genus("Panthera")
#   genus    kingdom  phylum   class    order     family  life_form
# 1 Panthera Animalia Chordata Mammalia Carnivora Felidae animal

taxify_register_coverage() shows which backbones contain a genus and at what version. If the genus does not appear for the requested backbone, a name in it is out of that backbone’s scope.

taxify_register_coverage("Quercus")
#     genus   backbone version date_added
# 1 Quercus  col     2025    2026-04-01
# 2 Quercus  gbif    current 2026-04-01
# 3 Quercus  wfo     2024-12 2026-04-01

A genus covered by all three backbones can be matched by any of them, while a genus covered only by GBIF (perhaps a recently described bacterial genus) will not match against WFO or COL.

When an unmatched name’s genus is in the register but covered by none of the requested backbones, taxify sets match_type = "out_of_scope" instead of "none". An out-of-scope result means the name likely exists in a different backbone, where a plain "none" leaves open whether it is a misspelling or an invalid name.

# Trying to match a marine invertebrate against WFO (plants only)
result <- taxify("Carcinus maenas", backbone = "wfo")
result$match_type
# [1] "out_of_scope"

result$life_form
# [1] "animal"

The life_form column places the genus among animals, so WFO is the wrong backbone for this name. In a pipeline, the "out_of_scope" rows can be filtered and re-run against a broader backbone; including the right backbones in the chain from the start avoids the second pass. The print() method of a taxify result tallies out-of-scope names by life_form.

The backbones

The table summarizes all 19. “Approx. names” is the total number of name strings in the compiled backbone (accepted names plus synonyms); the species count is lower because each accepted species may have several synonym entries pointing to it.

Backend Full name Scope Approx. names Source format
wfo World Flora Online Vascular plants, bryophytes ~1.6M Zenodo ZIP (classification.txt)
col Catalogue of Life All kingdoms ~5.3M ChecklistBank DwC-A (Taxon.tsv)
colxr Catalogue of Life Extended Release All kingdoms ~7.9M ChecklistBank DwC-A export of the XR dataset
gbif GBIF Backbone Taxonomy All kingdoms ~6.4M GBIF simple.txt.gz (30 positional cols)
itis Integrated Taxonomic Information System All kingdoms, US focus ~990k SQLite dump from itis.gov
ncbi NCBI Taxonomy All life incl. viruses ~2.7M Pipe-delimited .dmp files (taxdump)
ott Open Tree of Life All life (synthetic) ~3.7M Pipe-delimited taxonomy.tsv + synonyms.tsv
worms World Register of Marine Species Marine and brackish ~1.6M ChecklistBank DwC-A
euromed Euro+Med PlantBase European/Mediterranean plants ~147k Semicolon-delimited CSV
fungorum Species Fungorum Plus Fungi ~315k ChecklistBank DwC-A
algaebase AlgaeBase Algae and cyanobacteria ~170k ChecklistBank DwC-A (CC BY-NC)
fishbase FishBase Fishes ~100k rfishbase (load_taxa + synonyms)
sealifebase SeaLifeBase Non-fish marine and aquatic ~134k rfishbase (load_taxa + synonyms)
reptiledb Reptile Database Reptiles ~50k reptarium taxa.csv + synonym/checklist XLSX
lcvp Leipzig Catalogue of Vascular Plants Vascular plants ~1.3M idiv-biodiversity tab_lcvp.rda (R data package)
wcvp World Checklist of Vascular Plants (Kew) Vascular plants ~1.4M Kew wcvp.zip (wcvp_names.csv)
mdd Mammal Diversity Database Mammals ~62k MDD.zip of CSVs (species + synonyms)
avilist AviList (Global Avian Checklist) Birds ~41k AviList extended .xlsx
lpsn List of Prokaryotic names with Standing in Nomenclature Bacteria and archaea ~45k ChecklistBank ColDP (NameUsage.tsv)

WFO (Borsch et al. 2020) is the standard reference for plant taxonomy, maintained by the World Flora Online consortium and updated regularly. The backbone includes all taxonomic ranks from kingdom down to form, with full synonym resolution and authorship.

COL (Banki et al. 2024) and GBIF (GBIF Secretariat 2024) both cover all kingdoms with different curation strategies. COL is an expert-curated checklist assembled from over 160 sector databases, each maintained by a taxonomic authority for its group. GBIF’s backbone is assembled algorithmically from COL, ITIS, and dozens of other sources, which gives it broader raw coverage (~6.4M names against COL’s ~5.3M) and occasional inconsistencies where source databases disagree. In practice COL tends to give cleaner synonym resolution and GBIF tends to match more names.

ITIS (2025) was originally developed for North American fauna and remains particularly strong on freshwater invertebrates, insects, and US-listed species; its coverage of non-American taxa is uneven. It is distributed as a SQLite dump, so building the backbone from source requires the RSQLite package, a dependency the pre-built .vtr avoids.

NCBI Taxonomy is the reference taxonomy for sequence-linked work. Every GenBank, RefSeq, and BOLD sequence is linked to an NCBI tax_id, which makes this backbone essential for molecular ecology and metagenomics, and it is the only backbone that covers bacteria, archaea, and viruses in meaningful depth. NCBI Taxonomy stores no authorship data, so authorship is always NA for NCBI-matched rows.

OTT (Open Tree of Life) is a synthetic taxonomy that merges NCBI, GBIF, WoRMS, IRMNG, and several other sources into a single tree. It has the broadest coverage of any single source and cross-references all of its constituent databases through the sourceinfo field. Synthetic taxonomies can carry conflicts and inconsistencies at the edges, where source databases disagree about the placement of a taxon.

WoRMS is the authoritative source for marine species, curated by a network of over 300 taxonomic editors and covering marine, brackish, and some freshwater species. Beyond basic taxonomy, the WoRMS backbone stores habitat flags (marine, brackish, freshwater, terrestrial) and extinction status, some of which are accessible through the COL SpeciesProfile.

Euro+Med PlantBase (2026) is the taxonomic reference for the flora of Europe, the Mediterranean, and the Caucasus. It covers all native and introduced vascular plants in its geographic scope (~49k accepted names, ~83k synonyms). The backbone is built from the 2020 bulk download, updated by a PESI API delta refresh (April 2026) that resolved 1,014 reclassifications and synonym changes cross-referenced against WFO and POWO. It suits European vegetation surveys and datasets aligned with the European Vegetation Archive (EVA). Its data is licensed CC BY-SA 3.0.

Species Fungorum Plus is the specialist reference for fungal taxonomy, with ~315k names curated by the Royal Botanic Gardens, Kew. It covers Ascomycota, Basidiomycota, and other fungal phyla, including anamorphs and teleomorphs, and for purely mycological datasets it gives better synonym resolution than generalist databases.

AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. It is the only backbone licensed CC BY-NC (non-commercial use only); all other backbones are open-access.

FishBase covers fishes and SeaLifeBase the non-fish marine and aquatic groups (molluscs, crustaceans, marine mammals, and the rest). Both are compiled from the rfishbase package rather than a file download, taking accepted species from load_taxa() and synonym links from synonyms(). They carry species-rank rows only, so their genera enter the genus register derived from those species. The data is CC BY-NC 3.0.

The Reptile Database (Uetz et al. 2026) is the taxonomic reference for reptiles, with ~50k names covering snakes, lizards, turtles, crocodilians, and the tuatara. It is built from the reptarium bulk exports: the current accepted-species list, the periodic synonym snapshot, and the family-to-order checklist. The source carries no kingdom or class field, so the backbone stamps Animalia / Chordata / Reptilia on every row. The data is CC BY 4.0.

The Mammal Diversity Database is the American Society of Mammalogists’ reference mammal taxonomy, distributed as a zip of CSVs holding the accepted species with their full higher classification and every name ever applied to them. The data is MIT-licensed.

AviList is the global bird checklist that merged the long-standing IOC, Clements, and BirdLife split, shipped as one Excel workbook covering order through subspecies. It publishes no synonym table, so the backbone derives homotypic synonyms from the Protonym column: where a species has since moved genus, its original combination is a synonym of the current name (Parus caeruleus -> Cyanistes caeruleus). The data is CC BY 4.0.

LPSN is the nomenclatural authority for prokaryotes, hosted by DSMZ, and it records whether a bacterial or archaeal name is validly published, which NCBI, GBIF, and OTT leave open. LPSN’s own download route sits behind a free DSMZ account, so the backbone is built from the open ColDP mirror on GBIF ChecklistBank. The data is CC BY-SA 4.0.

Output differences between backbones

All 19 backbones produce the same 27-column output schema, so downstream code does not need to know which backbone produced a match. The content of those columns varies in several ways.

Authorship. WFO’s scientificName is already canonical (no authorship appended), so the authorship column comes from a separate scientificNameAuthorship field. COL and WoRMS store the full scientificName with authorship included; taxify strips it at build time to produce the canonical name used for matching and stores the stripped authorship separately. NCBI and OTT have no authorship data, so authorship is always NA for those backbones. GBIF and ITIS provide authorship, Euro+Med provides it from its AuthorString field, and Species Fungorum and AlgaeBase from their DwC-A archives.

Taxon IDs. Each backbone uses a different identifier system. WFO IDs look like "wfo-0000000123". COL IDs are opaque alphanumeric strings like "4LHBG". GBIF uses integer keys ("2878688"), ITIS TSN integers ("183671"), NCBI its Taxonomy IDs ("9606"), and OTT its own IDs ("770315"). WoRMS uses AphiaIDs extracted from LSIDs: at build time taxify strips the urn:lsid:marinespecies.org:taxname: prefix and stores just the numeric ID. Euro+Med uses TaxonUsageID integers from the PlantBase export, and Species Fungorum and AlgaeBase use ChecklistBank dataset-specific IDs. All IDs are stored as character strings in taxon_id and accepted_id, but their format is backbone-specific and meaningful only within that backbone’s ecosystem: a taxon_id from WFO cannot be looked up in the COL database, and vice versa.

Classification depth. The base output always includes family and genus. WFO provides these directly from its classification file. COL stores the full Linnaean hierarchy (kingdom through order) in the Taxon.tsv, though the extra columns require add_col_info(). GBIF provides family through a denormalized family_key self-join at build time. ITIS, NCBI, and OTT resolve family and genus by parent-hierarchy walks during backbone compilation, which traverse up to 25 levels of the taxonomic tree. WoRMS has denormalized classification columns in its DwC-A, and Euro+Med resolves family and genus by a hierarchy walk on IsChildTaxonOfID. The genus register fills in the higher classification fields (kingdom_group, taxon_group, life_form) for all backbones.

Synonym handling. WFO and COL use the Darwin Core field acceptedNameUsageID to point from a synonym row to its accepted name. GBIF encodes synonyms through parent_key pointing to the accepted taxon. NCBI represents synonyms as alternative name strings for the same tax_id; at build time taxify emits these as separate rows with synthetic IDs of the form "123456_syn_1", "123456_syn_2", etc. OTT uses a separate synonyms.tsv file with explicit synonym-to-accepted mappings. All of these representations are normalized at build time into the same is_synonym + accepted_name + accepted_id schema.

Synonym chains. Some backbones contain chained synonyms, where synonym A points to synonym B, which points to accepted name C. taxify resolves these chains at build time (up to 10 hops), so accepted_name always points to the terminal accepted name, at no cost to query-time performance.

The genus register

The genus register is a unified index of the genera across every supported backbone. It holds 503,262 genera, each with its family, higher classification (kingdom through order, where available), and a life_form label (e.g., "vascular plant", "animal", "fungus"). Where two backbones disagree about which family a genus belongs to, the classification is resolved by priority: WoRMS, COL Extended Release, COL, WCVP, Reptile Database, MDD, AviList, LPSN, GBIF, Euro+Med, LCVP, ITIS, NCBI, OTT, WFO, FishBase, SeaLifeBase, Species Fungorum, AlgaeBase. If COL and WFO disagree, COL’s assignment wins.

taxifydb builds the register over that fixed backbone set and publishes it as a versioned asset, the same way it publishes each backbone, so it arrives by download. taxify resolves it on first use: local disk, then the manifest download, then a local build through taxify_build_register() (equivalently taxify_download("register")) if taxifydb is installed. A register built from whichever backbones a machine happened to have installed made the same code report different kingdom_group and life_form labels on two machines, which the published asset prevents. A backend_coverage asset ships alongside it, one row per genus and backbone, and is what taxify_register_coverage() reads.

The register serves two purposes in matching. It supplies the life_form, kingdom_group, and taxon_group columns for every matched name, regardless of which backbone matched it, so results can be stratified by broad taxonomic group without looking up each family. It also enables out-of-scope detection: before fuzzy matching begins, taxify checks whether an unmatched name’s genus is known to the register but absent from the coverage table of every requested backbone, and if so marks the name "out_of_scope" immediately. This skips fuzzy matching against a backbone that could never produce a match and gives a more informative signal than a plain "none".

Choosing backbones

The right backbone depends on the taxonomic scope of the data. The general rule is specialist backbones first and generalist backbones second: lead with the backbone whose taxonomic opinion we trust most for the dominant taxon group, and add broader backbones as fallbacks for the remainder. The backbone column then records which taxonomic opinion was applied to each name.

Performance

Backbone size affects download time and, to a lesser extent, matching speed: WFO (~1.6M names) matches faster than GBIF (~6.4M names) for the same query. taxify uses index-accelerated genus-blocked joins at the C level (via vectra), so a list of 5,000 names resolves against GBIF in under a second on modern hardware. The difference only becomes noticeable at scale (100k+ names) or with heavy fuzzy matching against a large backbone.

In a multi-backbone chain, putting the most likely backbone first saves time, because names matched by the first backbone skip all later ones. If 90% of a list is plants, c("wfo", "col") is faster than c("col", "wfo"): WFO is smaller and resolves most names on the first pass, and COL, which is larger, only processes the remaining 10%.

Fuzzy matching is the most expensive step. It runs a genus-blocked fuzzy join with multi-threaded string distance computation. For names with misspelled genera, where genus blocking cannot help, backbones configured for it fall back to a 2-character prefix block that catches most genus-level typos while keeping the search space manageable.

Versions and reproducibility

taxify checks for backbone updates once per R session. The first taxify() call in a session fetches the manifest from GitHub, compares each requested backbone’s local version against the latest release, and downloads a new version only if one exists. If the network is unavailable, taxify falls back to the bundled manifest and uses whatever local copy is on disk. The version check and any update are logged to the console with the old and new version numbers, and the backbone version never changes mid-session. The manifest, which maps backbone names to their download URLs and latest versions, is cached per session and can be refreshed with taxify_refresh_manifest().

For a published analysis, the backbone_version strings belong in the methods section or supplementary material. A project that needs exact reproducibility can pin every backbone with taxify_download(version = ), as in Downloading backbones, and never use the “latest” slot, which keeps tracking new releases independently. A project that prefers to stay current can rely on the default “latest” behavior and cite the backbone_version strings from the output.

Backbones with large source files can also be built from source: taxify_build("gbif") downloads the raw 1.5 GB simple.txt.gz from GBIF, parses all 30 positional columns, denormalizes the family hierarchy through self-joins, and compiles the result into .vtr format. This is slower than downloading the pre-built file and produces the same output; it is mainly useful for CI pipelines or for customizing the compilation step.

References

Banki O, Roskov Y, Doring M, Ower G, Hernandez Robles DR, Plata Corredor CA, Stjernegaard Jeppesen T, Orrell TM, Pugh D, Kostichka J, et al. (2024). Catalogue of Life Checklist. https://doi.org/10.48580/d4t2

Borsch T, Berendsohn W, Dalcin E, Delmas M, Demissew S, Elliott A, Fritsch P, Fuchs A, Geltman D, Guner A, et al. (2020). World Flora Online: Placing taxonomists at the heart of a definitive and comprehensive global resource on the world’s plants. Taxon 69: 1311-1341. https://doi.org/10.1002/tax.12373

Euro+Med (2026). Euro+Med PlantBase: the information resource for Euro-Mediterranean plant diversity. https://europlusmed.org, accessed 2026-07-11.

GBIF Secretariat (2024). GBIF Backbone Taxonomy. https://doi.org/10.15468/39omei

ITIS (2025). Integrated Taxonomic Information System. https://doi.org/10.5066/F7KH0KBK

Uetz P, Freed P, Aguilar R, Reyes F, Kudera J, Hosek J (2026). The Reptile Database. http://www.reptile-database.org