---
title: "Requesting GBIF data for a species list"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Requesting GBIF data for a species list}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  eval = FALSE
)
```

GBIF (the Global Biodiversity Information Facility) assigns each taxon in its
backbone an integer identifier called a *taxon key*. taxify uses these keys to
request occurrence records for a species list.

`taxify()` matches names against a local copy of the GBIF backbone.
`gbif_request()` submits the keys to GBIF, and `gbif_fetch()` waits for the
download and imports the records with the input names attached. GBIF caps a
download query at about 1,000 keys, so `gbif_request()` splits a longer list
into several downloads, each with its own DOI
([details](#long-lists-are-split-across-downloads)).

Name matching runs offline. Requests and downloads require
[rgbif](https://CRAN.R-project.org/package=rgbif). The default download method
also requires a free GBIF account; see [Setting up rgbif](#setting-up-rgbif).

```{r}
library(taxify)

spp <- c("Quercus robur", "Pinus sylvestris", "Bellis perennis", "Morus alba")
```

## Match first, then request

Match the names, submit the request, fetch the records and retrieve the citations:

```{r}
matched <- taxify(spp, backbone = "gbif")   # resolve the names to GBIF keys
dl      <- gbif_request(matched)            # ask GBIF for every key
recs    <- gbif_fetch(dl)                   # wait, download, attach the input names
cite(recs)                                  # the download's DOI
```

The matching result is a data frame. Inspect it before submitting the request:

```{r}
matched <- taxify(spp, backbone = "gbif")

matched[, c("input_name", "accepted_name", "accepted_id", "taxonomic_status", "n_ids")]
#>         input_name    accepted_name accepted_id taxonomic_status n_ids
#> 1    Quercus robur    Quercus robur     2878688         ACCEPTED     3
#> 2 Pinus sylvestris Pinus sylvestris     5285637         ACCEPTED     6
#> 3  Bellis perennis  Bellis perennis     3117424         ACCEPTED     1
#> 4       Morus alba       Morus alba     5361889         ACCEPTED     1
#> Sources: GBIF 2023.08 | cite() for full citations
```

`accepted_id` contains the selected GBIF key. `n_ids` gives the number of keys
associated with the name; `taxify_ids()` lists them individually.

By default, `gbif_request()` requests all associated keys. GBIF prepares the
download in the background, and `gbif_fetch()` waits for it to finish:

```{r}
dl   <- gbif_request(matched)
recs <- gbif_fetch(dl)
```

The returned `input_name` column links records to the queried names for counting
or grouping. `cite(recs)` reports the package, backbone and download citations;
see [Citing the download](#citing-the-download).

## Match and request in one call

`gbif_request()` also accepts a vector of names. It calls `taxify()` internally,
passing through matching arguments such as `fuzzy`, `kingdom` and `region`:

```{r}
dl   <- gbif_request(spp)   # match, then ask GBIF for every key
recs <- gbif_fetch(dl)      # wait, download, import, attach the input names
```

The matched table is available afterwards as `attr(dl, "taxa")`. Call `taxify()`
separately if you want to inspect the matches before submitting a request.

Use `dry_run = TRUE` to inspect the keys and estimated record count without
contacting GBIF:

```{r}
gbif_request(spp, dry_run = TRUE)
#> 11 GBIF key(s) from 4 matched name(s). About 4,090,085 occurrence record(s) across them.
#>  [1] 2878688 7911626 7586523 5285637 7718215 8116613 9079676 8251433 7393648
#> [10] 3117424 5361889
```

## Why one name sometimes gives several keys

A name can have several GBIF keys because it is a homonym shared by different
taxa, or because the backbone contains doubtful or duplicate entries alongside
an accepted entry. `taxify_ids()` returns one row per name and key, including
its occurrence count:

```{r}
taxify(spp, backbone = "gbif") |>
  taxify_ids()
#>          input_name accepted_id    accepted_name taxonomic_status n_occurrences is_pick
#> 1     Quercus robur     2878688    Quercus robur         ACCEPTED       1725528    TRUE
#> 2     Quercus robur     7911626    Quercus robur         DOUBTFUL             2   FALSE
#> 3     Quercus robur     7586523    Quercus robur         DOUBTFUL             9   FALSE
#> 4  Pinus sylvestris     5285637 Pinus sylvestris         ACCEPTED       1204947    TRUE
#> 5  Pinus sylvestris     7718215 Pinus sylvestris         DOUBTFUL             0   FALSE
#> 6  Pinus sylvestris     8116613 Pinus sylvestris         DOUBTFUL             0   FALSE
#> 7  Pinus sylvestris     9079676 Pinus sylvestris         DOUBTFUL             0   FALSE
#> 8  Pinus sylvestris     8251433 Pinus sylvestris         DOUBTFUL             0   FALSE
#> 9  Pinus sylvestris     7393648 Pinus sylvestris         DOUBTFUL             0   FALSE
#> 10  Bellis perennis     3117424  Bellis perennis         ACCEPTED       1107989    TRUE
#> 11       Morus alba     5361889       Morus alba         ACCEPTED         51610    TRUE
```

`is_pick` marks the key reported in `accepted_id`. The other rows are further
records of the same name. Requesting only the pick would have missed the eleven
records under the two doubtful *Quercus robur* keys.

`n_occurrences` is the count taken when the backbone was built. For an accepted
key it includes the records of that taxon's synonyms and descendants, which is
what a download by that key returns, so it gives the approximate size of a
request.

## Choosing what to request

Every key is requested by default. `strict = TRUE` narrows the request to
the single key reported in `accepted_id`:

```{r}
gbif_request(spp, strict = TRUE, dry_run = TRUE)
#> 4 GBIF key(s) from 4 matched name(s). About 4,090,074 occurrence record(s) across them.
#> [1] 2878688 5285637 3117424 5361889
```

For this list the two choices differ by 11 records out of 4.09 million, so
requesting every key costs almost nothing and keeps records that would
otherwise be dropped. For a list with many homonyms the difference is larger,
and `taxify_ids()` shows it before choosing.

## Setting up rgbif

Sending the request needs rgbif:

```{r}
install.packages("rgbif")
```

`method = "search"` works with nothing further. `method = "download"` is an
authenticated call and needs a GBIF account, which is free: register at
[gbif.org](https://www.gbif.org/).

rgbif reads that account from three environment variables. Its own
documentation recommends keeping them in `.Renviron`, which also keeps them out
of scripts:

```{r}
usethis::edit_r_environ()
```

Add the three lines, with no quotes and no spaces around the `=`:

```
GBIF_USER=yourname
GBIF_PWD=yourpassword
GBIF_EMAIL=name@example.org
```

Save, then restart R so the file is read. rgbif also accepts the lower-case
names `gbif_user`, `gbif_pwd` and `gbif_email` in `.Rprofile` as an
alternative; see `?rgbif::occ_download`.

To confirm GBIF accepts the account, ask it for the account's past downloads.
That call is authenticated, so it only succeeds when the credentials are right:

```{r}
rgbif::occ_download_list(limit = 5)
```

If a variable is missing, `gbif_request()` stops before contacting GBIF and
names which one. If the password is wrong, GBIF answers 401.

## Sending the request

`method = "search"` sends unauthenticated searches and returns the records
directly. It needs no account, and the GBIF search API limits how deep it will
page, so it suits a look at a few taxa:

```{r}
recs <- gbif_request(spp, method = "search", strict = TRUE, limit = 50)
nrow(recs)
#> [1] 200
```

The records come back as one data frame with a `taxon_key_requested` column, so
a row can be traced back to the key that asked for it.

`method = "download"` is the default. It submits an asynchronous GBIF download,
which has no record cap and is issued a citable DOI. It needs a GBIF
account:

```{r}
dl <- gbif_request(spp, method = "download")
recs <- gbif_fetch(dl)
```

`gbif_fetch()` waits for the download, retrieves it, imports it, and attaches
the queried names.

All of the keys go into one download: they are sent as a single
`taxonKey IN (...)` predicate, so the whole list comes back as one dataset
under one DOI. GBIF prepares a download of four million records on its own
servers, which takes a while, so the call returns as soon as the request is
queued.

### Long lists are split across downloads

GBIF caps a download query at 12,000 characters. A taxonKey predicate costs
about 10 characters per key, so roughly 1,187 keys fit, and a checklist of a
few hundred names can pass that once its homonyms are counted.

Above the limit `gbif_request()` splits the keys into chunks of 1,000 and
submits them through rgbif's `occ_download_queue()`, which respects GBIF's rule
of three concurrent downloads per user:

```{r}
dl <- gbif_request(long_species_list, method = "download")
#> 2431 keys exceed GBIF's query limit; splitting into 3 downloads of up to
#> 1000 keys. Each gets its own DOI.
```

`dl` is then a vector of download keys, and `gbif_fetch()` waits for all of
them and stacks the records, so fetching a split request looks the same as
fetching a single one:

```{r}
recs <- gbif_fetch(dl)
```

The cap applies to the query naming the keys; a single download can hold tens
of millions of records.

## Linking the records back to the queried names

The records carry GBIF taxon keys. Linking them to the names asked for is a
second step, and `gbif_backmatch()` does it, adding `requested_key`,
`input_name` and `accepted_name`:

```{r}
recs <- gbif_request(spp, method = "search", limit = 100)
recs <- gbif_backmatch(recs, recs)

table(recs$input_name, useNA = "ifany")
```

`gbif_fetch()` calls it itself, so this is only needed when the records were
fetched some other way or with `backmatch = FALSE`:

```{r}
recs <- gbif_fetch(dl, backmatch = FALSE)
recs <- gbif_backmatch(recs, dl)
```

It takes the object `gbif_request()` returned, which carries the key table as
an attribute. A `taxify_ids()` table or a `taxify()` result works too, so
records imported by hand can still be linked back:

```{r}
matched <- taxify(spp, backbone = "gbif")
recs <- gbif_backmatch(recs, matched)
```

### Why not just join on taxonKey

GBIF returns the records of a key's descendants along with its own, and a
record identified below species level carries its own key. Requesting
*Pinus nigra* (5284809) returns records whose `taxonKey` is 5686674,
*P. nigra* subsp. *salzmannii*; those rows carry the requested key in
`speciesKey` and nowhere else.

A join on `taxonKey` misses those rows. In a sample of 300 records from that
request, 50 were identified to a subspecies, and a join on `taxonKey` linked
the other 250. `gbif_backmatch()` linked all 300. For each record it reads the
keys in order, `taxonKey`, then `acceptedTaxonKey`, `speciesKey`, `genusKey`
and `familyKey`, and takes the first one that was requested.

Some records carry none of the requested keys. Those rows are kept, with `NA`
in `requested_key`, `input_name` and `accepted_name`, and a message reports how
many there were.

## Citing the download

A GBIF download gets a DOI once it is ready, and citing it is what lets someone
else retrieve the same records. The download key travels with the records, so
`cite()` reports the DOI next to the backbone the names were matched against:

```{r}
recs <- gbif_fetch(dl)

cite(recs)
#> ── taxify citations ────────────────────────────────────────────────
#>   [1] Colling G (2026). taxify: Offline Taxonomic Name Matching (version 0.6.0).
#>   [2] GBIF Secretariat (2024). GBIF Backbone Taxonomy. doi:10.15468/39omei
#>   [3] GBIF.org (2026-09-23) GBIF Occurrence Download https://doi.org/10.15468/dl.ckzs32
#>   ────────────────────────────────────────────────────────────
```

`cite()` reads the DOI from GBIF when it is called, because GBIF issues it only
once the download has finished preparing. A download still running is reported
as having no DOI yet.

`cite(recs, file = "refs.bib")` writes the same entries as BibTeX, the
download included.

## Results matched against another backbone

Every match above named `backbone = "gbif"`. taxify's default chain starts at
COL XR, so a plain `taxify(spp)` resolves most names against another backbone
and the result carries no GBIF keys. Passing such a result works: the rows
without a GBIF key are re-matched against GBIF, with a message, because GBIF
only accepts its own keys.

```{r}
taxify(spp) |>
  gbif_request(dry_run = TRUE)
#> No rows matched by GBIF; re-matching the names against the GBIF backbone, whose keys a GBIF request needs.
#> 11 GBIF key(s) from 4 matched name(s). About 4,090,085 occurrence record(s) across them.
#>  [1] 2878688 7911626 7586523 5285637 7718215 8116613 9079676 8251433 7393648
#> [10] 3117424 5361889
```

A mixed result, where the one name only GBIF carries landed on `gbif` and the
rest elsewhere, is handled the same way: the rows without a key are re-matched
and the rest are kept, so every name in the list reaches the request.

`gbif_request()` takes matching arguments only when it does the matching
itself. With a `taxify()` result as input, they go in the `taxify()` call:

```{r}
taxify(spp, backbone = "gbif", kingdom = "Plantae") |>
  gbif_request()
```

## See also

- `taxify_ids()` lists the keys and their occurrence counts.

- `vignette("large-scale")` covers lists too large to hold in memory.

- `vignette("quickstart")` introduces name matching.
