This vignette shows how to pick a taxonomic backbone for a name list, how to chain several backbones so that each name is resolved by the source best placed to judge it, and how to check afterwards which source resolved which name. taxify matches names against locally stored Darwin Core backbone databases. 19 backbones are available, each compiled from a different authoritative source. The backbone we choose determines which names can be matched, which taxonomic opinion governs synonym resolution, and which extra metadata columns are available downstream.
list_backbones() reports each one’s current version, name
count, and whether it is installed.taxify_download().backbone argument of taxify().backbone.backbone and backbone_version columns.add_wfo_info(), add_col_info(), and
add_gbif_info().lookup_genus() and
taxify_register_coverage().taxify downloads a backbone on first use. When
taxify(names, backbone = "wfo") finds no local WFO
backbone, it fetches the pre-built .vtr file from the
taxifydb release it is published under, writes it to
taxify_data_dir(), and caches the path for the rest of the
R session. Later calls, in the same session or in future ones, reuse the
local copy without network access. taxify_download()
fetches backbones ahead of time, which is useful on a shared server or
in a Docker image:
# Download one backbone
taxify_download("wfo")
# Download several at once
taxify_download(c("wfo", "col", "worms"))Pre-built .vtr files are published as GitHub Releases on
taxifydb, one tag per
backbone (gbif-2026.06, algaebase-2026.06),
and range from 8 MB (AviList) to 2.0 GB (COL). They are compiled from
the raw Darwin Core sources with precomputed matching keys, embedded
synonym resolution, and genus-level indexes, so they can be queried as
soon as the download completes.
A specific backbone version can be pinned:
Pinned versions are stored in their own directory
(taxify_data_dir()/wfo/2024.01/) and are never overwritten
by future updates. The “latest” slot
(taxify_data_dir()/wfo/latest/) is overwritten whenever a
newer version becomes available. A project that locks a backbone version
produces identical results regardless of when the analysis is
re-run.
The simplest case matches plant names against WFO:
plants <- c(
"Quercus robur",
"Quercus petraea",
"Pinus sylvestris",
"Acer pseudoplatanus",
"Betula pendula",
"Fagus sylvatica",
"Picea abies"
)
result <- taxify(plants, backbone = "wfo")
result[, c("input_name", "accepted_name", "family", "match_type", "backbone")]Every row in the output has 27 columns regardless of which backbone
produced it: input_name, matched_name,
accepted_name, taxon_id,
accepted_id, rank, family,
genus, epithet, authorship,
accepted_authorship, is_synonym,
taxonomic_status, is_hybrid,
match_type, fuzzy_dist, n_ids,
accepted_ids, backbone,
backbone_version, kingdom_group,
taxon_group, life_form,
qualifier, qualifier_position,
aggregate_fallback, and hybrid_type. The
backbone column records "wfo" for matched rows
and NA for unmatched ones. The
backbone_version column records the backbone name, version,
and download date (e.g., "wfo:2024-12 (2026-04-01)"), so we
can cite the exact data snapshot used.
When a name matches a synonym, taxify resolves it to the accepted
name. The matched_name column shows the string that matched
in the backbone (which may be a synonym), accepted_name
shows the current accepted name after resolution, and
is_synonym is TRUE for resolved synonyms and
FALSE for direct matches.
Fuzzy matching is on by default with a normalized Damerau-Levenshtein threshold of 0.2, roughly one edit per five characters, which catches typos such as transposed letters or missing diacritics. A lower threshold is stricter:
fuzzy = FALSE turns fuzzy matching off and accepts only
exact and case-insensitive matches:
The fuzzy_method argument selects the distance metric:
the default "dl" (Damerau-Levenshtein, which counts a
transposition as one edit), "levenshtein" (standard
Levenshtein, no transposition handling), or "jw"
(Jaro-Winkler, which weights agreement in the first characters of the
two strings).
A wetland monitoring dataset might contain vascular plants, invertebrates, amphibians, algae, and fungi, and no single backbone covers all of these equally well. Passing a vector of backbone names builds a fallback chain: names are matched against each backbone in order, and a name matched by an earlier backbone is removed from the pool and never re-matched by a later one.
mixed <- c(
"Quercus robur", # plant
"Panthera leo", # animal
"Amanita muscaria", # fungus
"Salmo trutta", # fish
"Escherichia coli" # bacterium
)
result <- taxify(mixed, backbone = c("wfo", "col", "gbif"))
result[, c("input_name", "accepted_name", "match_type", "backbone")]The console output during matching shows the chain in action:
Matching 5 names against 3 backbones: wfo -> col -> gbif
[wfo] Matching 5 names...
[col] Matching 4 remaining names...
[gbif] Matching 1 remaining names...
"Quercus robur" matches in WFO and leaves the pool. The
remaining four names go to COL, and anything still unmatched after COL
(perhaps an obscure bacterial name) goes to GBIF. The process continues
until all names have been tried against all backbones or all names have
matched. If earlier backbones have matched every name, later ones are
skipped with a message:
[gbif] Skipped (all names matched)
The order of the vector decides which taxonomic opinion wins for each
name. If "Quercus robur" exists in both WFO and COL,
putting WFO first means WFO’s accepted name, family assignment, and
synonym resolution are used; putting COL first gives COL’s. For names
that exist in several backbones, the first backbone in the chain always
wins. With GBIF first, everything would match there (GBIF has ~6.4M
names, the largest of any backbone) and the curated opinions of WFO,
COL, or WoRMS would never be consulted. For a plant-heavy list with some
non-plant taxa, c("wfo", "col") or
c("wfo", "col", "gbif") gives WFO’s curated plant taxonomy
for the plants while COL or GBIF picks up the rest.
Fuzzy matching runs independently within each backbone. A name that fails exact matching in WFO is fuzzy-matched against WFO, and only if that also fails does it move to the next backbone, where it is exact-matched and then fuzzy-matched again. A misspelled plant name therefore gets its best chance in WFO, the plant-specialist backbone, before falling through to COL or GBIF.
WFO focuses on accepted vascular plant and bryophyte names. Names that appear only in older literature, belong to genera not yet integrated into WFO, or are nomenclaturally orphaned (no clear accepted name) may be absent. COL inherits WFO’s plant taxonomy as one of its sector databases and supplements it with names from other sources, including historical synonyms and cultivar names.
plants <- c(
"Quercus robur",
"Quercus petraea",
"Pinus sylvestris",
"Acer pseudoplatanus",
"Coffea arabica",
"Welwitschia mirabilis",
"Lepidodendron aculeatum", # extinct lycopsid
"Nothofagus cunninghamii",
"Dracaena draco"
)
# WFO alone
wfo_result <- taxify(plants, backbone = "wfo")
table(wfo_result$match_type)If any names come back as "none", COL can be added as a
fallback:
# WFO first, COL as fallback
both_result <- taxify(plants, backbone = c("wfo", "col"))
table(both_result$match_type)
both_result[, c("input_name", "accepted_name", "backbone")]The backbone column now shows "wfo" for
names matched by WFO and "col" for names only COL could
resolve. A methods section can then state “plant names were resolved
against WFO 2024-12, with unmatched names resolved against COL 2025.” In
large vegetation plot datasets most names resolve in WFO, and the
handful of edge cases (cultivars, historical names, genera recently
moved between families) that fall through to COL would otherwise need
manual resolution.
A monitoring dataset from a coastal estuary might contain vascular plants, invertebrates, fish, and marine algae. COL has broad expert-curated coverage across kingdoms, GBIF fills gaps with its larger name pool, including names from national checklists not yet incorporated into COL, and WoRMS is a final backstop for marine invertebrate synonyms, which can be slow to propagate to generalist databases. The chain runs from highest curation to broadest coverage, and most of these names resolve in COL.
estuary_species <- c(
"Zostera marina", # seagrass (plant)
"Salicornia europaea", # glasswort (plant)
"Carcinus maenas", # shore crab
"Mytilus edulis", # blue mussel
"Platichthys flesus", # European flounder
"Nereis diversicolor", # ragworm
"Fucus vesiculosus", # bladderwrack (brown alga)
"Littorina littorea", # common periwinkle
"Arenicola marina", # lugworm
"Cerastoderma edule" # common cockle
)
result <- taxify(estuary_species, backbone = c("col", "gbif", "worms"))
result[, c("input_name", "accepted_name", "family", "backbone")]For a predominantly marine list with only a few terrestrial taxa, leading with WoRMS makes its marine taxonomy take precedence:
Species Fungorum Plus is curated specifically for fungi, with ~315k names including anamorphs, teleomorphs, and the pleomorphic naming changes of the 2011 Melbourne Code. Its synonym coverage for fungal genera is better than in generalist databases, where fungal taxonomy is often a secondary concern.
fungi <- c(
"Amanita muscaria",
"Boletus edulis",
"Cantharellus cibarius",
"Tuber melanosporum",
"Saccharomyces cerevisiae",
"Aspergillus niger",
"Penicillium chrysogenum",
"Agaricus bisporus",
"Trametes versicolor",
"Cordyceps militaris"
)
result <- taxify(fungi, backbone = c("fungorum", "col"))
result[, c("input_name", "accepted_name", "is_synonym", "backbone")]Species Fungorum resolves the standard names, and COL picks up any obscure, recently described, or historically orphaned species that fall through. For a list of fungi and plants, a three-backbone chain gives each group its specialist, with COL catching what falls through both:
mixed <- c(
"Quercus robur", # plant
"Amanita muscaria", # fungus
"Lactarius deliciosus", # fungus
"Pinus sylvestris", # plant
"Russula emetica" # fungus
)
result <- taxify(mixed, backbone = c("wfo", "fungorum", "col"))Because the genus Amanita is not in WFO’s coverage
table, "Amanita muscaria" is marked out of scope for WFO
immediately and passed to the next backbone without fuzzy matching (see
The genus register).
AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. Its curation is strong for freshwater and marine microalgae, where generalist databases often have thin coverage and outdated synonymy.
algae <- c(
"Chlamydomonas reinhardtii",
"Chlorella vulgaris",
"Ulva lactuca",
"Fucus vesiculosus",
"Sargassum muticum"
)
result <- taxify(algae, backbone = c("algaebase", "col"))
result[, c("input_name", "accepted_name", "backbone")]AlgaeBase is licensed CC BY-NC, and taxify prints a license notice during its download. For commercial applications, COL or WoRMS can serve instead, with less specialized algal coverage.
Species lists from metabarcoding or eDNA studies are often linked to NCBI accessions. The NCBI backbone aligns taxify’s accepted names with the taxonomy used in GenBank and BOLD.
edna_hits <- c(
"Salmo trutta",
"Phoxinus phoxinus",
"Anguilla anguilla",
"Cottus gobio",
"Lampetra planeri",
"Chironomus riparius", # midge (insect)
"Potamopyrgus antipodarum" # New Zealand mud snail
)
result <- taxify(edna_hits, backbone = c("ncbi", "col"))
result[, c("input_name", "accepted_name", "taxon_id", "backbone")]The taxon_id values of NCBI-matched rows are NCBI
tax_ids, which link directly to GenBank records and NCBI taxonomy pages.
COL covers names not found in NCBI, such as taxa without sequenced
representatives.
backbone is a plain character column. In a
single-backbone call every matched row shows the same name; in a chain
it records which backbone produced each match; unmatched rows have
backbone = NA. It can count how many names each backbone
resolved:
result <- taxify(species_list, backbone = c("wfo", "col", "gbif"))
table(result$backbone, useNA = "ifany")or filter to the rows matched by one backbone:
wfo_matches <- result[result$backbone == "wfo" & !is.na(result$backbone), ]
col_matches <- result[result$backbone == "col" & !is.na(result$backbone), ]In a list expected to be purely plants, many names matched by COL instead of WFO point to names outside WFO’s scope, perhaps algae classified as plants in older literature, or animal-associated organisms such as plant parasites.
For a methods section, the unique backbone_version
strings give the provenance of the whole result:
unique(result$backbone_version[!is.na(result$backbone_version)])
# e.g., c("wfo:2024-12 (2026-04-01)", "col:2025 (2026-04-01)")Each string identifies both the taxonomic source and the snapshot used, so the result can be reproduced even if a backbone releases a new version between the analysis and a reviewer’s check.
Three backbones have enrichment functions that join extra
backbone-specific columns to a taxify result. Each fills only the rows
matched by its own backbone; rows from other backbones get
NA in the new columns.
add_wfo_info() adds scientificNameID,
parentNameUsageID, namePublishedIn,
higherClassification, taxonRemarks, and
infraspecificEpithet. namePublishedIn is
useful for citing original descriptions, and
higherClassification holds the full taxonomic hierarchy as
a semicolon-separated string.
result <- taxify(plants, backbone = "wfo") |>
add_wfo_info()
result[, c("input_name", "accepted_name", "namePublishedIn")]add_col_info() adds COL classification columns
(kingdom, phylum, col_class,
order), nomenclatural metadata (notho,
nomenclaturalCode, nomenclaturalStatus,
namePublishedIn), infraspecificEpithet, and
SpeciesProfile flags (is_extinct, is_marine,
is_freshwater, is_terrestrial). The
class column is renamed to col_class to avoid
conflict with R’s class() function. The SpeciesProfile
flags come from a separate file in the COL DwC-A archive and can, for
instance, exclude extinct species from a contemporary biodiversity
analysis or separate marine from terrestrial taxa in an estuarine
dataset.
result <- taxify(species_list, backbone = "col") |>
add_col_info()
# Check which species are marine
result[result$is_marine == TRUE & !is.na(result$is_marine),
c("input_name", "accepted_name", "kingdom", "is_marine")]add_gbif_info() adds notho_type (hybrid
type), nom_status (nomenclatural status),
bracket_authorship (basionym author),
bracket_year, gbif_year,
name_published_in, origin (how the name
entered the GBIF backbone), and infra_specific_epithet.
origin values such as "SOURCE",
"DENORMED_CLASSIFICATION", or
"VERBATIM_ACCEPTED" record how GBIF ingested the name.
result <- taxify(species_list, backbone = "gbif") |>
add_gbif_info()
result[, c("input_name", "accepted_name", "origin", "nom_status")]A multi-backbone result can pass through all three; each touches only rows from its own backbone:
result <- taxify(species_list, backbone = c("wfo", "col", "gbif")) |>
add_wfo_info() |>
add_col_info() |>
add_gbif_info()The result is a wide data.frame with the union of all extra columns,
where each row has only its own backbone’s columns populated and the
rest NA. For most workflows the base 27 columns are enough;
the extras matter when an analysis needs nomenclatural details, habitat
flags, or publication references that the standard output does not
include.
lookup_genus() returns the genus register row for a
single genus. The register is loaded into memory on the first call and
cached for the session.
lookup_genus("Quercus")
# genus kingdom phylum class order family life_form
# 1 Quercus Plantae ... ... Fagales Fagaceae vascular plantlookup_genus("Panthera")
# genus kingdom phylum class order family life_form
# 1 Panthera Animalia Chordata Mammalia Carnivora Felidae animaltaxify_register_coverage() shows which backbones contain
a genus and at what version. If the genus does not appear for the
requested backbone, a name in it is out of that backbone’s scope.
taxify_register_coverage("Quercus")
# genus backbone version date_added
# 1 Quercus col 2025 2026-04-01
# 2 Quercus gbif current 2026-04-01
# 3 Quercus wfo 2024-12 2026-04-01A genus covered by all three backbones can be matched by any of them, while a genus covered only by GBIF (perhaps a recently described bacterial genus) will not match against WFO or COL.
When an unmatched name’s genus is in the register but covered by none
of the requested backbones, taxify sets
match_type = "out_of_scope" instead of "none".
An out-of-scope result means the name likely exists in a different
backbone, where a plain "none" leaves open whether it is a
misspelling or an invalid name.
# Trying to match a marine invertebrate against WFO (plants only)
result <- taxify("Carcinus maenas", backbone = "wfo")
result$match_type
# [1] "out_of_scope"
result$life_form
# [1] "animal"The life_form column places the genus among animals, so
WFO is the wrong backbone for this name. In a pipeline, the
"out_of_scope" rows can be filtered and re-run against a
broader backbone; including the right backbones in the chain from the
start avoids the second pass. The print() method of a
taxify result tallies out-of-scope names by life_form.
The table summarizes all 19. “Approx. names” is the total number of name strings in the compiled backbone (accepted names plus synonyms); the species count is lower because each accepted species may have several synonym entries pointing to it.
| Backend | Full name | Scope | Approx. names | Source format |
|---|---|---|---|---|
wfo |
World Flora Online | Vascular plants, bryophytes | ~1.6M | Zenodo ZIP (classification.txt) |
col |
Catalogue of Life | All kingdoms | ~5.3M | ChecklistBank DwC-A (Taxon.tsv) |
colxr |
Catalogue of Life Extended Release | All kingdoms | ~7.9M | ChecklistBank DwC-A export of the XR dataset |
gbif |
GBIF Backbone Taxonomy | All kingdoms | ~6.4M | GBIF simple.txt.gz (30 positional cols) |
itis |
Integrated Taxonomic Information System | All kingdoms, US focus | ~990k | SQLite dump from itis.gov |
ncbi |
NCBI Taxonomy | All life incl. viruses | ~2.7M | Pipe-delimited .dmp files (taxdump) |
ott |
Open Tree of Life | All life (synthetic) | ~3.7M | Pipe-delimited taxonomy.tsv + synonyms.tsv |
worms |
World Register of Marine Species | Marine and brackish | ~1.6M | ChecklistBank DwC-A |
euromed |
Euro+Med PlantBase | European/Mediterranean plants | ~147k | Semicolon-delimited CSV |
fungorum |
Species Fungorum Plus | Fungi | ~315k | ChecklistBank DwC-A |
algaebase |
AlgaeBase | Algae and cyanobacteria | ~170k | ChecklistBank DwC-A (CC BY-NC) |
fishbase |
FishBase | Fishes | ~100k | rfishbase (load_taxa + synonyms) |
sealifebase |
SeaLifeBase | Non-fish marine and aquatic | ~134k | rfishbase (load_taxa + synonyms) |
reptiledb |
Reptile Database | Reptiles | ~50k | reptarium taxa.csv + synonym/checklist XLSX |
lcvp |
Leipzig Catalogue of Vascular Plants | Vascular plants | ~1.3M | idiv-biodiversity tab_lcvp.rda (R data package) |
wcvp |
World Checklist of Vascular Plants (Kew) | Vascular plants | ~1.4M | Kew wcvp.zip (wcvp_names.csv) |
mdd |
Mammal Diversity Database | Mammals | ~62k | MDD.zip of CSVs (species + synonyms) |
avilist |
AviList (Global Avian Checklist) | Birds | ~41k | AviList extended .xlsx |
lpsn |
List of Prokaryotic names with Standing in Nomenclature | Bacteria and archaea | ~45k | ChecklistBank ColDP (NameUsage.tsv) |
WFO (Borsch et al. 2020) is the standard reference for plant taxonomy, maintained by the World Flora Online consortium and updated regularly. The backbone includes all taxonomic ranks from kingdom down to form, with full synonym resolution and authorship.
COL (Banki et al. 2024) and GBIF (GBIF Secretariat 2024) both cover all kingdoms with different curation strategies. COL is an expert-curated checklist assembled from over 160 sector databases, each maintained by a taxonomic authority for its group. GBIF’s backbone is assembled algorithmically from COL, ITIS, and dozens of other sources, which gives it broader raw coverage (~6.4M names against COL’s ~5.3M) and occasional inconsistencies where source databases disagree. In practice COL tends to give cleaner synonym resolution and GBIF tends to match more names.
ITIS (2025) was
originally developed for North American fauna and remains particularly
strong on freshwater invertebrates, insects, and US-listed species; its
coverage of non-American taxa is uneven. It is distributed as a SQLite
dump, so building the backbone from source requires the RSQLite package,
a dependency the pre-built .vtr avoids.
NCBI Taxonomy is the reference taxonomy for sequence-linked work.
Every GenBank, RefSeq, and BOLD sequence is linked to an NCBI tax_id,
which makes this backbone essential for molecular ecology and
metagenomics, and it is the only backbone that covers bacteria, archaea,
and viruses in meaningful depth. NCBI Taxonomy stores no authorship
data, so authorship is always NA for
NCBI-matched rows.
OTT (Open Tree of Life) is a synthetic taxonomy that merges NCBI,
GBIF, WoRMS, IRMNG, and several other sources into a single tree. It has
the broadest coverage of any single source and cross-references all of
its constituent databases through the sourceinfo field.
Synthetic taxonomies can carry conflicts and inconsistencies at the
edges, where source databases disagree about the placement of a
taxon.
WoRMS is the authoritative source for marine species, curated by a network of over 300 taxonomic editors and covering marine, brackish, and some freshwater species. Beyond basic taxonomy, the WoRMS backbone stores habitat flags (marine, brackish, freshwater, terrestrial) and extinction status, some of which are accessible through the COL SpeciesProfile.
Euro+Med PlantBase (2026) is the taxonomic reference for the flora of Europe, the Mediterranean, and the Caucasus. It covers all native and introduced vascular plants in its geographic scope (~49k accepted names, ~83k synonyms). The backbone is built from the 2020 bulk download, updated by a PESI API delta refresh (April 2026) that resolved 1,014 reclassifications and synonym changes cross-referenced against WFO and POWO. It suits European vegetation surveys and datasets aligned with the European Vegetation Archive (EVA). Its data is licensed CC BY-SA 3.0.
Species Fungorum Plus is the specialist reference for fungal taxonomy, with ~315k names curated by the Royal Botanic Gardens, Kew. It covers Ascomycota, Basidiomycota, and other fungal phyla, including anamorphs and teleomorphs, and for purely mycological datasets it gives better synonym resolution than generalist databases.
AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. It is the only backbone licensed CC BY-NC (non-commercial use only); all other backbones are open-access.
FishBase covers fishes and SeaLifeBase the non-fish marine and
aquatic groups (molluscs, crustaceans, marine mammals, and the rest).
Both are compiled from the rfishbase package rather than a file
download, taking accepted species from load_taxa() and
synonym links from synonyms(). They carry species-rank rows
only, so their genera enter the genus register derived from those
species. The data is CC BY-NC 3.0.
The Reptile Database (Uetz et al. 2026) is the taxonomic reference for reptiles, with ~50k names covering snakes, lizards, turtles, crocodilians, and the tuatara. It is built from the reptarium bulk exports: the current accepted-species list, the periodic synonym snapshot, and the family-to-order checklist. The source carries no kingdom or class field, so the backbone stamps Animalia / Chordata / Reptilia on every row. The data is CC BY 4.0.
The Mammal Diversity Database is the American Society of Mammalogists’ reference mammal taxonomy, distributed as a zip of CSVs holding the accepted species with their full higher classification and every name ever applied to them. The data is MIT-licensed.
AviList is the global bird checklist that merged the long-standing
IOC, Clements, and BirdLife split, shipped as one Excel workbook
covering order through subspecies. It publishes no synonym table, so the
backbone derives homotypic synonyms from the Protonym
column: where a species has since moved genus, its original combination
is a synonym of the current name (Parus caeruleus ->
Cyanistes caeruleus). The data is CC BY 4.0.
LPSN is the nomenclatural authority for prokaryotes, hosted by DSMZ, and it records whether a bacterial or archaeal name is validly published, which NCBI, GBIF, and OTT leave open. LPSN’s own download route sits behind a free DSMZ account, so the backbone is built from the open ColDP mirror on GBIF ChecklistBank. The data is CC BY-SA 4.0.
All 19 backbones produce the same 27-column output schema, so downstream code does not need to know which backbone produced a match. The content of those columns varies in several ways.
Authorship. WFO’s scientificName is
already canonical (no authorship appended), so the
authorship column comes from a separate
scientificNameAuthorship field. COL and WoRMS store the
full scientificName with authorship included; taxify strips
it at build time to produce the canonical name used for matching and
stores the stripped authorship separately. NCBI and OTT have no
authorship data, so authorship is always NA
for those backbones. GBIF and ITIS provide authorship, Euro+Med provides
it from its AuthorString field, and Species Fungorum and
AlgaeBase from their DwC-A archives.
Taxon IDs. Each backbone uses a different identifier
system. WFO IDs look like "wfo-0000000123". COL IDs are
opaque alphanumeric strings like "4LHBG". GBIF uses integer
keys ("2878688"), ITIS TSN integers
("183671"), NCBI its Taxonomy IDs ("9606"),
and OTT its own IDs ("770315"). WoRMS uses AphiaIDs
extracted from LSIDs: at build time taxify strips the
urn:lsid:marinespecies.org:taxname: prefix and stores just
the numeric ID. Euro+Med uses TaxonUsageID integers from
the PlantBase export, and Species Fungorum and AlgaeBase use
ChecklistBank dataset-specific IDs. All IDs are stored as character
strings in taxon_id and accepted_id, but their
format is backbone-specific and meaningful only within that backbone’s
ecosystem: a taxon_id from WFO cannot be looked up in the
COL database, and vice versa.
Classification depth. The base output always
includes family and genus. WFO provides these
directly from its classification file. COL stores the full Linnaean
hierarchy (kingdom through order) in the Taxon.tsv, though the extra
columns require add_col_info(). GBIF provides family
through a denormalized family_key self-join at build time.
ITIS, NCBI, and OTT resolve family and genus by parent-hierarchy walks
during backbone compilation, which traverse up to 25 levels of the
taxonomic tree. WoRMS has denormalized classification columns in its
DwC-A, and Euro+Med resolves family and genus by a hierarchy walk on
IsChildTaxonOfID. The genus register fills in the higher
classification fields (kingdom_group,
taxon_group, life_form) for all backbones.
Synonym handling. WFO and COL use the Darwin Core
field acceptedNameUsageID to point from a synonym row to
its accepted name. GBIF encodes synonyms through parent_key
pointing to the accepted taxon. NCBI represents synonyms as alternative
name strings for the same tax_id; at build time taxify
emits these as separate rows with synthetic IDs of the form
"123456_syn_1", "123456_syn_2", etc. OTT uses
a separate synonyms.tsv file with explicit
synonym-to-accepted mappings. All of these representations are
normalized at build time into the same is_synonym +
accepted_name + accepted_id schema.
Synonym chains. Some backbones contain chained
synonyms, where synonym A points to synonym B, which points to accepted
name C. taxify resolves these chains at build time (up to 10 hops), so
accepted_name always points to the terminal accepted name,
at no cost to query-time performance.
The genus register is a unified index of the genera across every
supported backbone. It holds 503,262 genera, each with its family,
higher classification (kingdom through order, where available), and a
life_form label (e.g., "vascular plant",
"animal", "fungus"). Where two backbones
disagree about which family a genus belongs to, the classification is
resolved by priority: WoRMS, COL Extended Release, COL, WCVP, Reptile
Database, MDD, AviList, LPSN, GBIF, Euro+Med, LCVP, ITIS, NCBI, OTT,
WFO, FishBase, SeaLifeBase, Species Fungorum, AlgaeBase. If COL and WFO
disagree, COL’s assignment wins.
taxifydb builds the register over that fixed backbone set and
publishes it as a versioned asset, the same way it publishes each
backbone, so it arrives by download. taxify resolves it on first use:
local disk, then the manifest download, then a local build through
taxify_build_register() (equivalently
taxify_download("register")) if taxifydb is installed. A
register built from whichever backbones a machine happened to have
installed made the same code report different kingdom_group
and life_form labels on two machines, which the published
asset prevents. A backend_coverage asset ships alongside
it, one row per genus and backbone, and is what
taxify_register_coverage() reads.
The register serves two purposes in matching. It supplies the
life_form, kingdom_group, and
taxon_group columns for every matched name, regardless of
which backbone matched it, so results can be stratified by broad
taxonomic group without looking up each family. It also enables
out-of-scope detection: before fuzzy matching begins, taxify checks
whether an unmatched name’s genus is known to the register but absent
from the coverage table of every requested backbone, and if so marks the
name "out_of_scope" immediately. This skips fuzzy matching
against a backbone that could never produce a match and gives a more
informative signal than a plain "none".
The right backbone depends on the taxonomic scope of the data. The
general rule is specialist backbones first and generalist backbones
second: lead with the backbone whose taxonomic opinion we trust most for
the dominant taxon group, and add broader backbones as fallbacks for the
remainder. The backbone column then records which taxonomic
opinion was applied to each name.
backbone = "wfo". If
some names fall through (horticultural cultivars, nomenclaturally
complex genera, or names from older floras that use outdated synonymy),
add COL: backbone = c("wfo", "col").backbone = c("euromed", "wfo"). Euro+Med PlantBase is the
taxonomic reference used by EVA and covers all native and introduced
vascular plants of Europe, the Mediterranean, and the Caucasus. Leading
with it makes European synonym resolution follow Euro+Med’s opinion,
with WFO as fallback for non-European taxa or names outside its scope.
Euro+Med data is CC BY-SA 3.0.backbone = "worms", which is curated by domain experts and
includes habitat and extinction flags. For estuarine or transitional
lists with some terrestrial taxa, add COL:
backbone = c("worms", "col").backbone = "fungorum", with
COL as fallback for obscure or recently described species:
backbone = c("fungorum", "col").backbone = "algaebase", with
COL or WoRMS as fallback: backbone = c("algaebase", "col").
AlgaeBase is CC BY-NC.backbone = c("col", "gbif"), or lead with a specialist
backbone for the dominant taxon group: c("wfo", "col") for
a plant-dominated dataset with some animals and fungi,
c("worms", "col", "gbif") for a marine biodiversity survey,
and c("wfo", "fungorum", "col") for a forest inventory that
includes trees, fungi, insects, and epiphytes.backbone = "ncbi", the reference for GenBank, BOLD, and
other sequence databases, which covers the bacteria, archaea, and
viruses other backbones lack. For mixed molecular and ecological work:
backbone = c("ncbi", "col").backbone = "ott", the
backbone of the Open Tree of Life. Its cross-references to NCBI, GBIF,
WoRMS, and IRMNG bridge different identifier systems.backbone = c("col", "gbif").
COL provides expert-curated taxonomy for ~5.3M names and GBIF’s backbone
adds ~6.4M names from additional sources; together they cover virtually
all described species with a nomenclatural record. This combination is a
reasonable default when the taxonomic composition of the dataset is
unknown.Backbone size affects download time and, to a lesser extent, matching speed: WFO (~1.6M names) matches faster than GBIF (~6.4M names) for the same query. taxify uses index-accelerated genus-blocked joins at the C level (via vectra), so a list of 5,000 names resolves against GBIF in under a second on modern hardware. The difference only becomes noticeable at scale (100k+ names) or with heavy fuzzy matching against a large backbone.
In a multi-backbone chain, putting the most likely backbone first
saves time, because names matched by the first backbone skip all later
ones. If 90% of a list is plants, c("wfo", "col") is faster
than c("col", "wfo"): WFO is smaller and resolves most
names on the first pass, and COL, which is larger, only processes the
remaining 10%.
Fuzzy matching is the most expensive step. It runs a genus-blocked fuzzy join with multi-threaded string distance computation. For names with misspelled genera, where genus blocking cannot help, backbones configured for it fall back to a 2-character prefix block that catches most genus-level typos while keeping the search space manageable.
taxify checks for backbone updates once per R session. The first
taxify() call in a session fetches the manifest from
GitHub, compares each requested backbone’s local version against the
latest release, and downloads a new version only if one exists. If the
network is unavailable, taxify falls back to the bundled manifest and
uses whatever local copy is on disk. The version check and any update
are logged to the console with the old and new version numbers, and the
backbone version never changes mid-session. The manifest, which maps
backbone names to their download URLs and latest versions, is cached per
session and can be refreshed with
taxify_refresh_manifest().
For a published analysis, the backbone_version strings
belong in the methods section or supplementary material. A project that
needs exact reproducibility can pin every backbone with
taxify_download(version = ), as in Downloading backbones, and never use
the “latest” slot, which keeps tracking new releases independently. A
project that prefers to stay current can rely on the default “latest”
behavior and cite the backbone_version strings from the
output.
Backbones with large source files can also be built from source:
taxify_build("gbif") downloads the raw 1.5 GB
simple.txt.gz from GBIF, parses all 30 positional columns,
denormalizes the family hierarchy through self-joins, and compiles the
result into .vtr format. This is slower than downloading
the pre-built file and produces the same output; it is mainly useful for
CI pipelines or for customizing the compilation step.
Banki O, Roskov Y, Doring M, Ower G, Hernandez Robles DR, Plata Corredor CA, Stjernegaard Jeppesen T, Orrell TM, Pugh D, Kostichka J, et al. (2024). Catalogue of Life Checklist. https://doi.org/10.48580/d4t2
Borsch T, Berendsohn W, Dalcin E, Delmas M, Demissew S, Elliott A, Fritsch P, Fuchs A, Geltman D, Guner A, et al. (2020). World Flora Online: Placing taxonomists at the heart of a definitive and comprehensive global resource on the world’s plants. Taxon 69: 1311-1341. https://doi.org/10.1002/tax.12373
Euro+Med (2026). Euro+Med PlantBase: the information resource for Euro-Mediterranean plant diversity. https://europlusmed.org, accessed 2026-07-11.
GBIF Secretariat (2024). GBIF Backbone Taxonomy. https://doi.org/10.15468/39omei
ITIS (2025). Integrated Taxonomic Information System. https://doi.org/10.5066/F7KH0KBK
Uetz P, Freed P, Aguilar R, Reyes F, Kudera J, Hosek J (2026). The Reptile Database. http://www.reptile-database.org