GBIF (the Global Biodiversity Information Facility) assigns each taxon in its backbone an integer identifier called a taxon key. taxify uses these keys to request occurrence records for a species list.
taxify() matches names against a local copy of the GBIF
backbone. gbif_request() submits the keys to GBIF, and
gbif_fetch() waits for the download and imports the records
with the input names attached. GBIF caps a download query at about 1,000
keys, so gbif_request() splits a longer list into several
downloads, each with its own DOI (details).
Name matching runs offline. Requests and downloads require rgbif. The default download method also requires a free GBIF account; see Setting up rgbif.
Match the names, submit the request, fetch the records and retrieve the citations:
matched <- taxify(spp, backbone = "gbif") # resolve the names to GBIF keys
dl <- gbif_request(matched) # ask GBIF for every key
recs <- gbif_fetch(dl) # wait, download, attach the input names
cite(recs) # the download's DOIThe matching result is a data frame. Inspect it before submitting the request:
matched <- taxify(spp, backbone = "gbif")
matched[, c("input_name", "accepted_name", "accepted_id", "taxonomic_status", "n_ids")]
#> input_name accepted_name accepted_id taxonomic_status n_ids
#> 1 Quercus robur Quercus robur 2878688 ACCEPTED 3
#> 2 Pinus sylvestris Pinus sylvestris 5285637 ACCEPTED 6
#> 3 Bellis perennis Bellis perennis 3117424 ACCEPTED 1
#> 4 Morus alba Morus alba 5361889 ACCEPTED 1
#> Sources: GBIF 2023.08 | cite() for full citationsaccepted_id contains the selected GBIF key.
n_ids gives the number of keys associated with the name;
taxify_ids() lists them individually.
By default, gbif_request() requests all associated keys.
GBIF prepares the download in the background, and
gbif_fetch() waits for it to finish:
The returned input_name column links records to the
queried names for counting or grouping. cite(recs) reports
the package, backbone and download citations; see Citing the download.
gbif_request() also accepts a vector of names. It calls
taxify() internally, passing through matching arguments
such as fuzzy, kingdom and
region:
dl <- gbif_request(spp) # match, then ask GBIF for every key
recs <- gbif_fetch(dl) # wait, download, import, attach the input namesThe matched table is available afterwards as
attr(dl, "taxa"). Call taxify() separately if
you want to inspect the matches before submitting a request.
Use dry_run = TRUE to inspect the keys and estimated
record count without contacting GBIF:
A name can have several GBIF keys because it is a homonym shared by
different taxa, or because the backbone contains doubtful or duplicate
entries alongside an accepted entry. taxify_ids() returns
one row per name and key, including its occurrence count:
taxify(spp, backbone = "gbif") |>
taxify_ids()
#> input_name accepted_id accepted_name taxonomic_status n_occurrences is_pick
#> 1 Quercus robur 2878688 Quercus robur ACCEPTED 1725528 TRUE
#> 2 Quercus robur 7911626 Quercus robur DOUBTFUL 2 FALSE
#> 3 Quercus robur 7586523 Quercus robur DOUBTFUL 9 FALSE
#> 4 Pinus sylvestris 5285637 Pinus sylvestris ACCEPTED 1204947 TRUE
#> 5 Pinus sylvestris 7718215 Pinus sylvestris DOUBTFUL 0 FALSE
#> 6 Pinus sylvestris 8116613 Pinus sylvestris DOUBTFUL 0 FALSE
#> 7 Pinus sylvestris 9079676 Pinus sylvestris DOUBTFUL 0 FALSE
#> 8 Pinus sylvestris 8251433 Pinus sylvestris DOUBTFUL 0 FALSE
#> 9 Pinus sylvestris 7393648 Pinus sylvestris DOUBTFUL 0 FALSE
#> 10 Bellis perennis 3117424 Bellis perennis ACCEPTED 1107989 TRUE
#> 11 Morus alba 5361889 Morus alba ACCEPTED 51610 TRUEis_pick marks the key reported in
accepted_id. The other rows are further records of the same
name. Requesting only the pick would have missed the eleven records
under the two doubtful Quercus robur keys.
n_occurrences is the count taken when the backbone was
built. For an accepted key it includes the records of that taxon’s
synonyms and descendants, which is what a download by that key returns,
so it gives the approximate size of a request.
Every key is requested by default. strict = TRUE narrows
the request to the single key reported in accepted_id:
gbif_request(spp, strict = TRUE, dry_run = TRUE)
#> 4 GBIF key(s) from 4 matched name(s). About 4,090,074 occurrence record(s) across them.
#> [1] 2878688 5285637 3117424 5361889For this list the two choices differ by 11 records out of 4.09
million, so requesting every key costs almost nothing and keeps records
that would otherwise be dropped. For a list with many homonyms the
difference is larger, and taxify_ids() shows it before
choosing.
Sending the request needs rgbif:
method = "search" works with nothing further.
method = "download" is an authenticated call and needs a
GBIF account, which is free: register at gbif.org.
rgbif reads that account from three environment variables. Its own
documentation recommends keeping them in .Renviron, which
also keeps them out of scripts:
Add the three lines, with no quotes and no spaces around the
=:
GBIF_USER=yourname
GBIF_PWD=yourpassword
GBIF_EMAIL=name@example.org
Save, then restart R so the file is read. rgbif also accepts the
lower-case names gbif_user, gbif_pwd and
gbif_email in .Rprofile as an alternative; see
?rgbif::occ_download.
To confirm GBIF accepts the account, ask it for the account’s past downloads. That call is authenticated, so it only succeeds when the credentials are right:
If a variable is missing, gbif_request() stops before
contacting GBIF and names which one. If the password is wrong, GBIF
answers 401.
method = "search" sends unauthenticated searches and
returns the records directly. It needs no account, and the GBIF search
API limits how deep it will page, so it suits a look at a few taxa:
The records come back as one data frame with a
taxon_key_requested column, so a row can be traced back to
the key that asked for it.
method = "download" is the default. It submits an
asynchronous GBIF download, which has no record cap and is issued a
citable DOI. It needs a GBIF account:
gbif_fetch() waits for the download, retrieves it,
imports it, and attaches the queried names.
All of the keys go into one download: they are sent as a single
taxonKey IN (...) predicate, so the whole list comes back
as one dataset under one DOI. GBIF prepares a download of four million
records on its own servers, which takes a while, so the call returns as
soon as the request is queued.
GBIF caps a download query at 12,000 characters. A taxonKey predicate costs about 10 characters per key, so roughly 1,187 keys fit, and a checklist of a few hundred names can pass that once its homonyms are counted.
Above the limit gbif_request() splits the keys into
chunks of 1,000 and submits them through rgbif’s
occ_download_queue(), which respects GBIF’s rule of three
concurrent downloads per user:
dl <- gbif_request(long_species_list, method = "download")
#> 2431 keys exceed GBIF's query limit; splitting into 3 downloads of up to
#> 1000 keys. Each gets its own DOI.dl is then a vector of download keys, and
gbif_fetch() waits for all of them and stacks the records,
so fetching a split request looks the same as fetching a single one:
The cap applies to the query naming the keys; a single download can hold tens of millions of records.
The records carry GBIF taxon keys. Linking them to the names asked
for is a second step, and gbif_backmatch() does it, adding
requested_key, input_name and
accepted_name:
recs <- gbif_request(spp, method = "search", limit = 100)
recs <- gbif_backmatch(recs, recs)
table(recs$input_name, useNA = "ifany")gbif_fetch() calls it itself, so this is only needed
when the records were fetched some other way or with
backmatch = FALSE:
It takes the object gbif_request() returned, which
carries the key table as an attribute. A taxify_ids() table
or a taxify() result works too, so records imported by hand
can still be linked back:
GBIF returns the records of a key’s descendants along with its own,
and a record identified below species level carries its own key.
Requesting Pinus nigra (5284809) returns records whose
taxonKey is 5686674, P. nigra subsp.
salzmannii; those rows carry the requested key in
speciesKey and nowhere else.
A join on taxonKey misses those rows. In a sample of 300
records from that request, 50 were identified to a subspecies, and a
join on taxonKey linked the other 250.
gbif_backmatch() linked all 300. For each record it reads
the keys in order, taxonKey, then
acceptedTaxonKey, speciesKey,
genusKey and familyKey, and takes the first
one that was requested.
Some records carry none of the requested keys. Those rows are kept,
with NA in requested_key,
input_name and accepted_name, and a message
reports how many there were.
A GBIF download gets a DOI once it is ready, and citing it is what
lets someone else retrieve the same records. The download key travels
with the records, so cite() reports the DOI next to the
backbone the names were matched against:
recs <- gbif_fetch(dl)
cite(recs)
#> ── taxify citations ────────────────────────────────────────────────
#> [1] Colling G (2026). taxify: Offline Taxonomic Name Matching (version 0.6.0).
#> [2] GBIF Secretariat (2024). GBIF Backbone Taxonomy. doi:10.15468/39omei
#> [3] GBIF.org (2026-09-23) GBIF Occurrence Download https://doi.org/10.15468/dl.ckzs32
#> ────────────────────────────────────────────────────────────cite() reads the DOI from GBIF when it is called,
because GBIF issues it only once the download has finished preparing. A
download still running is reported as having no DOI yet.
cite(recs, file = "refs.bib") writes the same entries as
BibTeX, the download included.
Every match above named backbone = "gbif". taxify’s
default chain starts at COL XR, so a plain taxify(spp)
resolves most names against another backbone and the result carries no
GBIF keys. Passing such a result works: the rows without a GBIF key are
re-matched against GBIF, with a message, because GBIF only accepts its
own keys.
taxify(spp) |>
gbif_request(dry_run = TRUE)
#> No rows matched by GBIF; re-matching the names against the GBIF backbone, whose keys a GBIF request needs.
#> 11 GBIF key(s) from 4 matched name(s). About 4,090,085 occurrence record(s) across them.
#> [1] 2878688 7911626 7586523 5285637 7718215 8116613 9079676 8251433 7393648
#> [10] 3117424 5361889A mixed result, where the one name only GBIF carries landed on
gbif and the rest elsewhere, is handled the same way: the
rows without a key are re-matched and the rest are kept, so every name
in the list reaches the request.
gbif_request() takes matching arguments only when it
does the matching itself. With a taxify() result as input,
they go in the taxify() call:
taxify_ids() lists the keys and their occurrence
counts.
vignette("large-scale") covers lists too large to
hold in memory.
vignette("quickstart") introduces name
matching.