Package {sfhotspot}


Title: Hot-Spot Analysis with Simple Features
Version: 1.1.1
Description: Identify and understand clusters of points (typically representing the locations of places or events) stored in simple-features (SF) objects. This is useful for analysing, for example, hot-spots of crime events. The package emphasises producing results from point SF data in a single step using reasonable default values for all other arguments, to aid rapid data analysis by users who are starting out. Functions available include kernel density estimation (for details, see Yip (2020) <doi:10.22224/gistbok/2020.1.12>), analysis of spatial association (Getis and Ord (1992) <doi:10.1111/j.1538-4632.1992.tb00261.x>) and hot-spot classification (Chainey (2020) ISBN:158948584X).
License: MIT + file LICENSE
Language: en-GB
URL: https://pkgs.lesscrime.info/sfhotspot/
BugReports: https://github.com/mpjashby/sfhotspot/issues
Encoding: UTF-8
Imports: classInt, cli, dbscan, ggplot2, ggspatial, isoband, rlang, sf, SpatialKDE, spdep, tibble
Depends: R (≥ 4.1.0)
Suggests: testthat (≥ 3.0.0), knitr, lubridate, rmarkdown, quarto
LazyData: true
Config/testthat/edition: 3
VignetteBuilder: knitr
Config/roxygen2/version: 8.1.0
NeedsCompilation: no
Packaged: 2026-09-30 22:10:18 UTC; mattashby
Author: Matt Ashby ORCID iD [aut, cre]
Maintainer: Matt Ashby <matthew.ashby@ucl.ac.uk>
Repository: CRAN
Date/Publication: 2026-10-01 00:40:02 UTC

Plot map of hotspot classifications

Description

Plot the output produced by hotspot_classify() with reasonable defaults.

Usage

## S3 method for class 'hspt_c'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_c'
autolayer(object, ...)

## S3 method for class 'hspt_c'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_c, e.g. as produced by hotspot_classify().

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot map of changes in grid counts

Description

Plot the output produced by hotspot_change() with reasonable defaults.

Usage

## S3 method for class 'hspt_d'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_d'
autolayer(object, ...)

## S3 method for class 'hspt_d'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_d, e.g. as produced by hotspot_change().

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot map of dual kernel-density values

Description

Plot the output produced by hotspot_dual_kde() using a scale appropriate to the comparison method. Ratios, logged ratios and sums use sequential scales. Only differences use a diverging scale centred on zero. Legends label the lower and upper ends of each scale as "low" and "high", respectively.

Usage

## S3 method for class 'hspt_dk'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_dk'
autolayer(object, ...)

## S3 method for class 'hspt_dk'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_dk, e.g. as produced by hotspot_dual_kde(). The object must have a valid method attribute.

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot map of Getis-Ord Gi* results

Description

If object contains a kde column, density is shown only in cells in which the Gi*/Gi result passes the specified significance and sign conditions. The pvalue column is used as supplied and is not adjusted by the plotting methods. Cells that do not satisfy the conditions are transparent. When sign = "both", cold-spot densities are negated for plotting and a diverging scale distinguishes cold spots from hot spots. When only one sign is shown, a medium-to-dark sequential scale avoids making the least-dense significant cells appear nearly white. If object does not contain a kde column, the Gi*/Gi value is plotted using a diverging scale centred on zero and critical_p and sign do not affect the mapped values.

Usage

## S3 method for class 'hspt_g'
autoplot(
  object,
  critical_p = 0.05,
  sign = c("both", "hot", "cold"),
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_g'
autolayer(object, critical_p = 0.05, sign = c("both", "hot", "cold"), ...)

## S3 method for class 'hspt_g'
hotspot_map(
  object,
  mapping = NULL,
  critical_p = 0.05,
  sign = c("both", "hot", "cold"),
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_g, e.g. as produced by hotspot_gistar().

critical_p

A single numeric value specifying the largest p-value to treat as statistically significant when plotting density.

sign

Which significant results should show density: "both" (the default), "hot" for positive Gi*/Gi values, or "cold" for negative values.

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot isobands

Description

Plot the output produced by hotspot_isoband() using a sequential or diverging discrete scale appropriate to the original hotspot result and selected value. Legend entries use the ordered label column containing concise, automatically formatted ranges.

Usage

## S3 method for class 'hspt_ib'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_ib'
autolayer(object, ...)

## S3 method for class 'hspt_ib'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_ib, as produced by hotspot_isoband().

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot map of kernel-density values

Description

Plot the output produced by hotspot_kde() with reasonable default values.

Usage

## S3 method for class 'hspt_k'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_k'
autolayer(object, ...)

## S3 method for class 'hspt_k'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_k, e.g. as produced by hotspot_kde().

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot map of grid counts

Description

Plot the output produced by hotspot_count() with reasonable default values. Weighted counts are plotted when the object contains a sum column; otherwise unweighted counts in the n column are plotted.

Usage

## S3 method for class 'hspt_n'
autoplot(
  object,
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_n'
autolayer(object, ...)

## S3 method for class 'hspt_n'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_n, e.g. as produced by hotspot_count().

...

Further arguments passed to ggplot2::geom_sf(), e.g. alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns a layer that can be added to a ggplot2::ggplot object.

Functions


Plot DBSCAN hotspot clusters

Description

Plot the polygon clusters produced by hotspot_dbscan() with reasonable defaults. Polygons can be filled according to their count, proportion or rank, and optionally labelled with the same values.

Usage

## S3 method for class 'hspt_s'
autoplot(
  object,
  col_fill = c("n", "prop", "rank", "none"),
  col_label = "none",
  ...,
  basemap_type = "none",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  quiet = TRUE
)

## S3 method for class 'hspt_s'
autolayer(
  object,
  col_fill = c("n", "prop", "rank", "none"),
  col_label = "none",
  ...
)

## S3 method for class 'hspt_s'
hotspot_map(
  object,
  mapping = NULL,
  col_fill = c("n", "prop", "rank", "none"),
  col_label = "none",
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An object with class hspt_s, as produced by hotspot_dbscan().

col_fill

A single string specifying the column used for the fill aesthetic: "n" (the default), "prop" or "rank". Use "none" to draw unfilled polygons with the default ggplot2 border colour. Set colour in ... to use a different border colour.

col_label

One or more strings specifying the columns used for labels shown in each cluster: "none" (the default), or any combination of "n", "prop" and "rank". Multiple labels are separated by newlines. Proportions are formatted as percentages and ranks as ordinal numbers. Labels use a semi-transparent background chosen to contrast with the base map.

...

Static aesthetics and other arguments passed to ggplot2::geom_sf(), e.g. colour, fill or alpha.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map. autoplot() defaults to "none" and hotspot_map() defaults to "osm". A base map requires an internet connection unless the required tiles are cached. When a base map is used, the hotspot layer is drawn with alpha = 0.75, overriding any alpha value supplied in ....

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

mapping

For hotspot_map(), this must be NULL. Hotspot-result objects determine their own aesthetic mappings.

caption

For hotspot_map(), additional text to display in the map caption, or NULL for no additional text. Base-map attribution and any explanatory caption generated by the plotting method are retained, with explanatory text above this caption and attribution below it.

Value

autoplot() and hotspot_map() return ggplot2::ggplot objects. autolayer() returns one or more layers that can be added to a ggplot2::ggplot object.

Functions


Identify change in hotspots over time

Description

Identify change in the number of points (typically representing events) between two periods (before and after a specified date) or in two groups (e.g. on weekdays or at weekends).

Usage

hotspot_change(
  data,
  time = NULL,
  boundary = NULL,
  groups = NULL,
  cell_size = NULL,
  grid_type = "rect",
  grid = NULL,
  quiet = FALSE
)

Arguments

data

sf data frame containing points.

time

Name of the column in data containing Date or POSIXt values representing the date associated with each point. Ignored if groups is not NULL. If this argument is NULL and data contains a single column of Date or POSIXt values, that column will be used automatically.

boundary

A single Date or POSIXt value representing the point after which points should be treated as having occurred in the second time period. See 'Details'.

groups

Name of a column in data containing exactly two unique non-missing values, which will be used to identify whether each row should be counted in the first (before) or second (after) groups. Which groups to use will be determined by calling sort(unique(groups)). If groups is not a factor, a message will be printed confirming which value has been used for which group. See 'Details'.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

grid

sf data frame containing points containing polygons, which will be used as the grid for which counts are made.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

Details

This function creates a regular two-dimensional grid of cells (unless a custom grid is specified with grid) and calculates the difference between the number of points in each grid cell:

If both boundary and groups are not NULL, the value of boundary will be ignored.

Coverage of the output data

The grid produced by this function covers the convex hull of the input data layer. This means the result may include zero counts for cells that are outside the area for which data were provided, which could be misleading. To handle this, consider cropping the output layer to the area for which data are available. For example, if you only have crime data for a particular district, crop the output dataset to the district boundary using st_intersection.

Automatic cell-size selection

If no cell size is given then the cell size will be set so that there are 50 cells on the shorter side of the grid. If the data SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

Value

An sf tibble of regular grid cells with corresponding hot-spot classifications for each cell. This can be plotted using autoplot.

See Also

hotspot_dual_kde() for comparing the density of two layers, which will often be more useful than comparing counts if the point locations represent and underlying continuous distribution.

Examples


# Compare counts from the first half of the period covered by the data to
# counts from the second half

hotspot_change(memphis_robberies)


# Create a grouping variable, then compare counts across values of that
# variable

memphis_robberies$weekend <-
  weekdays(memphis_robberies$date) %in% c("Saturday", "Sunday")
hotspot_change(memphis_robberies, groups = weekend)



Classify hot-spots

Description

Classify cells in a grid based on changes in the clustering of points (typically representing events) in a two-dimensional regular grid over time.

Usage

hotspot_classify(
  data,
  time = NULL,
  period = NULL,
  start = NULL,
  cell_size = NULL,
  grid_type = "rect",
  grid = NULL,
  collapse = FALSE,
  params = hotspot_classify_params(),
  quiet = FALSE
)

Arguments

data

sf data frame containing points.

time

Name of the column in data containing Date or POSIXt values representing the date associated with each point. If this argument is NULL and data contains a single column of Date or POSIXt values, that column will be used automatically.

period

A character value containing a number followed by a unit of time, e.g. for example, "12 months" or "3.5 days", where the unit of time is one of second, minute, hour, day, week, month, quarter or year (or their plural forms).

start

A Date or POSIXt value specifying when the first temporal period should start. If NULL (the default), the first period will start at the beginning of the earliest date found in the data (if period is specified in days, weeks, months, quarters or years) or at the earliest time found in the data otherwise.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

grid

sf data frame containing points containing polygons, which will be used as the grid for which counts are made.

collapse

If the range of dates in the data is not a multiple of period, the final period will be shorter than the others. In that case, should this shorter period be collapsed into the penultimate period?

params

A list of optional parameters that can affect the output. The list can be produced most easily using the hotspot_classify_params helper function.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

Value

An sf tibble of regular grid cells with corresponding hot-spot classifications for each cell. This can be plotted using autoplot.

Hot-spots are spatial areas that contain more points than would be expected by chance; cold-spots are areas that contain fewer points than would be expected. Whether an area is a hot-spot can vary over time. This function creates a space-time cube, determines whether an area is a hot-spot for each of several consecutive time periods and uses that to classify areas according to whether they are persistent, intermittent, emerging or former hot- or cold-spots.

Hot and cold spots

Hot- and cold-spots are identified by calculating the Getis-Ord Gi* (gi-star) or Gi* Z-score statistic for each cell in a regular grid for each time period. The corresponding p-values are adjusted once within each period using p.adjustSP and the method specified by p_adjust_method in params. They are not adjusted again across time periods, to avoid overly-conservative results. Cells are classified as follows, using the parameters provided in the params argument:

Coverage of the output data

The grid produced by this function covers the convex hull of the input data layer. This means the result may include Gi* or Gi* values for cells that are outside the area for which data were provided, which could be misleading. To handle this, consider cropping the output layer to the area for which data are available. For example, if you only have crime data for a particular district, crop the output dataset to the district boundary using st_intersection.

Automatic cell-size selection

If no cell size is given then the cell size will be set so that there are 50 cells on the shorter side of the grid. If the data SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

References

Chainey, S. (2020). Understanding Crime: Analyzing the Geography of Crime. Redlands, CA: ESRI.


Control the parameters used to classify hotspots

Description

This function allows specification of parameters that affect the output from hotspot_classify.

Usage

hotspot_classify_params(
  hotspot_prop = 0.1,
  persistent_prop = 0.8,
  recent_prop = 0.2,
  critical_p = 0.05,
  nb_dist = NULL,
  include_self = TRUE,
  p_adjust_method = NULL
)

Arguments

hotspot_prop

A single numeric value specifying the minimum proportion of periods for which a cell must contain significant clusters of points before the cell can be classified as a hot or cold spot of any type.

persistent_prop

A single numeric value specifying the minimum proportion of periods for which a cell must contain significant clusters of points before the cell can be classified as a persistent hot or cold spot.

recent_prop

A single numeric value specifying the proportion of periods that should be treated as being recent in the classification of emerging and former hotspots.

critical_p

A threshold p-value below which values should be treated as being statistically significant.

nb_dist

The distance around a cell that contains the neighbours of that cell, which are used in calculating the statistic. If this argument is NULL (the default), nb_dist is set as cell_size * sqrt(2) so that only the cells immediately adjacent to each cell are treated as being its neighbours.

include_self

Should points in a given cell be counted as well as counts in neighbouring cells when calculating the values of Gi* (if include_self = TRUE, the default) or Gi* (if include_self = FALSE) values? You are unlikely to want to change the default value.

p_adjust_method

The method to be used to adjust p-values for multiple comparisons using p.adjustSP. NULL (the default) uses the default method used by p.adjust (currently "holm"), but any of the character values in stats::p.adjust.methods may be specified. Set this argument to "none" to return unadjusted p-values. The adjustment is applied separately to the cells in each time period.

Value

A list that can be used as the input to the params argument to hotspot_classify.


Extract spatial features inside a polygon

Description

Extract spatial features inside a polygon

Usage

hotspot_clip(data, boundary, quiet = FALSE, ...)

Arguments

data

sf data frame containing spatial features.

boundary

sf data frame containing polygons.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

...

Further arguments passed to st_intersection.

Details

This function is a wrapper around st_intersection that performs some additional checks and reports useful information. If data has a specialised result class produced by this package, that class is preserved in the clipped result. A warning is produced if clipping reduces the dimension of any geometry, such as from a polygon to a line.

Value

an SF data frame containing those spatial features that are covered by the polygons.


Count points in cells in a two-dimensional grid

Description

Count points in cells in a two-dimensional grid

Usage

hotspot_count(
  data,
  cell_size = NULL,
  grid_type = "rect",
  grid = NULL,
  weights = NULL,
  quiet = FALSE
)

Arguments

data

sf data frame containing points.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

grid

sf data frame containing polygons, which will be used as the grid for which counts are made.

weights

NULL or the name of a column in data to be used as weights for weighted counts.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

Details

This function counts the number of points in each cell in a regular grid. If a column name in data is supplied with the weights argument, weighted counts will also be produced.

Automatic cell-size selection

If grid is NULL and no cell size is given, the cell size will be set so that there are 50 cells on the shorter side of the grid. If the data SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

Value

An sf tibble of regular grid cells with corresponding point counts for each cell. This can be plotted using autoplot.

Examples


# Set cell size automatically

hotspot_count(memphis_robberies_jan)


# Transform data to UTM zone 15N so that cell_size and bandwidth can be set
# in metres
library(sf)
memphis_robberies_utm <- st_transform(memphis_robberies_jan, 32615)

# Manually set grid-cell size in metres, since the `memphis_robberies_utm`
# dataset uses a co-ordinate reference system (UTM zone 15 north) that is
# specified in metres

hotspot_count(memphis_robberies_utm, cell_size = 200)



Identify hotspots using DBSCAN

Description

Identify clusters of points using density-based spatial clustering and represent each cluster as a buffered convex or concave hull. Clusters are ranked by the number of points they contain.

Usage

hotspot_dbscan(
  data,
  eps = NULL,
  min_pts = NULL,
  density_adjust = 2,
  hull = c("concave", "convex"),
  hull_ratio = 0.75,
  transform = TRUE,
  quiet = FALSE,
  ...
)

Arguments

data

An sf::sf object containing point geometries.

eps

A single positive number specifying the DBSCAN neighbourhood radius in the units of the analysis co-ordinate reference system (CRS). If NULL, the radius is calculated automatically from nearest-neighbour distances. See Automatic parameter selection.

min_pts

A single integer specifying the minimum number of points in an eps neighbourhood, including the point itself, for a point to be a core point. If NULL, the value is calculated automatically from the number of point coordinates in data. See Automatic parameter selection.

density_adjust

A single positive number controlling automatic selection of eps. Ignored unless eps = NULL. A value of 1 (the reference neighbourhood density) uses the median nearest-neighbour distance; the default of 2 requires approximately twice that density.

hull

The type of hull used to represent each cluster: "convex" or "concave" (the default).

hull_ratio

For concave hulls, a number from zero to one specifying the fraction convex. Zero produces a maximally concave hull and one produces a convex hull. The default is 0.75. Ignored when hull = "convex".

transform

DBSCAN uses Euclidean distances and therefore requires a projected CRS. If TRUE (the default), geographic input is transformed automatically with st_transform_auto() before analysis and the result is transformed back afterwards. If FALSE, geographic input produces an error.

quiet

If TRUE, suppress informative messages about automatically selected parameters and geometry preparation.

...

Further arguments passed to dbscan::dbscan(), such as borderPoints and nearest-neighbour search controls. The arguments x, eps, minPts, and weights cannot be supplied through ....

Details

MULTIPOINT geometries are cast to individual points before analysis. The output n column counts all input point coordinates intersecting each final cluster polygon, not only points assigned to that DBSCAN cluster. DBSCAN includes border points in clusters by default; supplying borderPoints = FALSE through ... instead performs DBSCAN* and excludes border points from the points used to construct the hull.

Each cluster geometry is the selected hull of all assigned points, buffered by eps and clipped to the convex hull of all input points. Different cluster polygons may overlap after hull construction and buffering.

Value

An sf tibble with class hspt_s and one row per non-noise cluster. It contains cluster, the original DBSCAN cluster identifier; rank, the priority rank; n, the number of input point coordinates intersecting the polygon; prop, n as a proportion of all input point coordinates; prop_area, the polygon area as a proportion of the area covered by a hull of the complete dataset; and geometry. The complete-dataset hull uses the same hull, hull_ratio, buffering and clipping rules as the cluster hulls. Since cluster polygons may overlap, the sums of prop and prop_area can exceed one. Ranking is by decreasing n, then increasing polygon area, then increasing cluster identifier.

Automatic parameter selection

When min_pts = NULL, it is calculated as

min(n, 50, max(5, ceiling(sqrt(n)))),

where n is the number of point coordinates after empty geometries have been removed and MULTIPOINT geometries have been expanded. For datasets with two to four coordinates, all coordinates are required. At least two coordinates are needed for automatic selection. This rule increases the evidence required to identify a hotspot in larger datasets, while the upper limit prevents the required number of neighbours becoming excessively large.

When eps = NULL, the function calculates the distance from every point to its (min_pts - 1)th nearest other point. The median of those distances is a typical local neighbourhood radius. The value used for clustering is

eps = median_neighbour_distance / sqrt(density_adjust).

min_pts - 1 other points are used because DBSCAN counts the focal point itself. Since circular area is proportional to the square of its radius, the default density_adjust = 2 searches for the required number of points in approximately half the typical neighbourhood area, corresponding to approximately twice the typical local point density. Larger values identify denser concentrations; values below one allow less-dense concentrations.

Automatic values provide a starting point for exploratory analysis. DBSCAN results can be sensitive to both parameters, so users should consider whether the resulting neighbourhood size and minimum density are meaningful for their application.

Examples


# The default concave hulls require GEOS 3.11 or later.
if (utils::compareVersion(sf::sf_extSoftVersion()[["GEOS"]], "3.11.0") >= 0) {
  hotspot_dbscan(memphis_robberies_jan)
}

hotspot_dbscan(
  memphis_robberies_jan,
  density_adjust = 3,
  hull = "convex"
)



Estimate the relationship between the kernel density of two layers of points

Description

Estimate the relationship between the kernel density of two layers of points

Usage

hotspot_dual_kde(
  x,
  y,
  cell_size = NULL,
  grid_type = "rect",
  bandwidth = NULL,
  bandwidth_adjust = 1,
  method = "ratio",
  grid = NULL,
  weights = NULL,
  transform = TRUE,
  quiet = FALSE,
  ...
)

Arguments

x, y

sf data frames containing points.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the x argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

bandwidth

either a single numeric value specifying the bandwidth to be used in calculating the kernel density estimates, or a list of exactly 2 such values. If this argument is NULL (the default), a single bandwidth will be determined automatically using the result of bandwidth.nrd called on the co-ordinates of the points in x, then used for both x and y. If this argument is list(NULL, NULL), separate bandwidths will be determined automatically for x and y based on each layer. For lon/lat data that are transformed automatically, bandwidths are specified or calculated in the units of the projected co-ordinate reference system.

bandwidth_adjust

single positive numeric value by which the value of bandwidth for both x and y will be multiplied, or a list of two such values. Useful for setting the bandwidth relative to the default.

method

character specifying the method by which the densities, d(), of x and y will be related:

ratio

(the default) calculates the density of x divided by the density of y, i.e. d(x) / d(y).

log

calculates the natural logarithm of the density of x divided by the density of y, i.e. log(d(x) / d(y)).

diff

calculates the difference between the density of x and the density of y, i.e. d(x) - d(y).

sum

calculates the sum of the density of x and the density of y, i.e. d(x) + d(y).

The result of this calculation will be returned in the kde column of the return value.

grid

sf data frame containing polygons, which will be used as the grid for which densities are estimated. Both x and y must overlap the grid and use the same co-ordinate reference system.

weights

NULL (the default) or a vector of length two giving either NULL or the name of a column in each of x and y to be used as weights for weighted counts and KDE values.

transform

the underlying SpatialKDE package cannot calculate kernel density for lon/lat data, so this must be transformed to use a projected co-ordinate reference system. If this argument is TRUE (the default) and sf::st_is_longlat(x) is TRUE, x, y and grid (if provided) will be transformed to a common projected co-ordinate reference system chosen using st_transform_auto before bandwidths and kernel densities are estimated. The result is returned using the original co-ordinate reference system. Set this argument to FALSE to suppress automatic transformation of the data.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

...

Further arguments passed to kde.

Value

An sf tibble of grid cells with corresponding point counts and dual kernel density estimates for each cell. The result has classes hspt_dk and hspt_k, and a method attribute containing the comparison method used.

The output can be plotted with reasonable defaults using autoplot(). The colour scale is selected according to the comparison method.

This function creates a regular two-dimensional grid of cells (unless a custom grid is specified with grid), calculates the density of points in each cell for each of x and y using functions from the SpatialKDE package, then produces a value representing a relation between the two densities. The count of points in each cell is also returned.

Dual kernel density values can be useful for understanding the relationship between the distributions of two sets of point locations. For example:

Coverage of the output data

The grid produced by this function covers the convex hull of the points in x. This means the result may include KDE values for cells that are outside the area for which data were provided, which could be misleading. To handle this, consider cropping the output layer to the area for which data are available. For example, if you only have crime data for a particular district, crop the output dataset to the district boundary using st_intersection.

Automatic cell-size selection

If no cell size is given then the cell size will be set so that there are 50 cells on the shorter side of the grid. If the x SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

References

Yin, P. (2020). Kernels and Density Estimation. The Geographic Information Science & Technology Body of Knowledge (1st Quarter 2020 Edition), John P. Wilson (ed.). doi:doi:10.22224/gistbok/2020.1.12

Examples

# See also the examples for `hotspot_kde()` for examples of how to specify
# `cell_size`, `bandwidth`, etc.

library(sf)

# Transform data to UTM zone 15N so that cell_size and bandwidth can be set
# in metres
memphis_robberies_utm <- st_transform(memphis_robberies, 32615)
memphis_population_utm <- st_transform(memphis_population, 32615)

# Calculate burglary risk based on residential population. `weights` is set
# to `c(NULL, population)` so that the robberies layer is not weighted and
# the population layer is weighted according to the number of residents in
# each census block.

hotspot_dual_kde(
  memphis_robberies_utm,
  memphis_population_utm,
  bandwidth = list(NULL, NULL),
  weights = c(NULL, population)
)



Identify significant spatial clusters of points

Description

Identify hotspot and coldspot locations, that is cells in a regular grid in which there are more/fewer points than would be expected if the points were distributed randomly.

Usage

hotspot_gistar(
  data,
  cell_size = NULL,
  grid_type = "rect",
  kde = TRUE,
  bandwidth = NULL,
  bandwidth_adjust = 1,
  grid = NULL,
  weights = NULL,
  nb_dist = NULL,
  include_self = TRUE,
  p_adjust_method = NULL,
  transform = TRUE,
  quiet = FALSE,
  ...
)

Arguments

data

sf data frame containing points.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

kde

TRUE (the default) or FALSE indicating whether kernel density estimates (KDE) should be produced for each grid cell.

bandwidth

numeric value specifying the bandwidth to be used in calculating the kernel density estimates. If this argument is NULL (the default), the bandwidth will be specified automatically using the mean result of bandwidth.nrd called on the x and y co-ordinates separately.

bandwidth_adjust

single positive numeric value by which the value of bandwidth is multiplied. Useful for setting the bandwidth relative to the default.

grid

sf data frame containing polygons, which will be used as the grid for which counts are made.

weights

NULL or the name of a column in data to be used as weights for weighted counts, Gi*/Gi statistics and KDE values.

nb_dist

The distance around a cell that contains the neighbours of that cell, which are used in calculating the statistic. If this argument is NULL (the default), nb_dist is set as cell_size * sqrt(2) so that only the cells immediately adjacent to each cell are treated as being its neighbours.

include_self

Should points in a given cell be counted as well as counts in neighbouring cells when calculating the values of Gi* (if include_self = TRUE, the default) or Gi* (if include_self = FALSE) values? You are unlikely to want to change the default value.

p_adjust_method

The method to be used to adjust p-values for multiple comparisons using p.adjustSP. NULL (the default) uses the default method used by p.adjust (currently "holm"), but any of the character values in stats::p.adjust.methods may be specified. Set this argument to "none" to return unadjusted p-values.

transform

the underlying SpatialKDE package cannot calculate kernel density for lon/lat data, so this must be transformed to use a projected co-ordinate reference system. If this argument is TRUE (the default) and sf::st_is_longlat(data) is TRUE, data (and grid if provided) will be transformed automatically using link{st_transform_auto} before the kernel density is estimated and transformed back afterwards. Set this argument to FALSE to suppress automatic transformation of the data.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

...

Further arguments passed to kde or ignored if kde = FALSE.

Details

This function calculates the Getis-Ord Gi* (gi-star) or Gi* Z-score statistic for identifying clusters of point locations. The underlying implementation uses the localG function to calculate the Z scores and then p.adjustSP function to adjust the corresponding p-values for multiple comparison. The function also returns counts of points in each cell and (by default but optionally) kernel density estimates using the kde function. If weights is supplied, the Gi*/Gi statistics are calculated from the weighted counts; otherwise, they are calculated from the unweighted counts.

Coverage of the output data

The grid produced by this function covers the convex hull of the input data layer. This means the result may include Gi* or Gi* values for cells that are outside the area for which data were provided, which could be misleading. To handle this, consider cropping the output layer to the area for which data are available. For example, if you only have crime data for a particular district, crop the output dataset to the district boundary using st_intersection.

Automatic cell-size selection

If no cell size is given then the cell size will be set so that there are 50 cells on the shorter side of the grid. If the data SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

Value

An sf tibble with class hspt_g containing regular grid cells with corresponding point counts, Gi* or Gi* values and (optionally) kernel density estimates for each cell. Values greater than zero indicate more points than would be expected for randomly distributed points and values less than zero indicate fewer points. Critical values of Gi* and Gi* are given in the manual page for localG.

The output from this function can be plotted with reasonable defaults using autoplot(). When KDE values are returned, the default map shows density only in areas identified as having significantly more or fewer points than expected by chance. When kde = FALSE, the Gi*/Gi statistic is plotted on a diverging scale centred on zero.

References

Getis, A. & Ord, J. K. (1992). The Analysis of Spatial Association by Use of Distance Statistics. Geographical Analysis, 24(3), 189-206. doi:doi:10.1111/j.1538-4632.1992.tb00261.x

Examples

library(sf)

# Transform data to UTM zone 15N so that cell_size and bandwidth can be set
# in metres
memphis_robberies_utm <- st_transform(memphis_robberies_jan, 32615)

# Automatically set grid-cell size, bandwidth and neighbour distance

hotspot_gistar(memphis_robberies_utm)


# Manually set grid-cell size in metres, since the `memphis_robberies`
# dataset uses a co-ordinate reference system (UTM zone 15 north) that is
# specified in metres

hotspot_gistar(memphis_robberies_utm, cell_size = 200)


# Automatically set grid-cell size and bandwidth for lon/lat data. The data
# will be transformed automatically before KDE values are calculated and the
# result will then be transformed back to the original CRS.

hotspot_gistar(memphis_robberies)



Create either a rectangular or hexagonal two-dimensional grid

Description

Create either a rectangular or hexagonal two-dimensional grid

Usage

hotspot_grid(data, cell_size = NULL, grid_type = "rect", quiet = FALSE, ...)

Arguments

data

sf data frame.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. If this argument is NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex").

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

...

Further arguments passed to link[sf]{st_make_grid}.

Value

A simple features tibble containing polygons representing grid cells.

The grid will be based on the convex hull of data, expanded by a buffer of cell_size / 2 to ensure all the points in data fall within the resulting grid.


Convert a hotspot grid to isobands

Description

Generalise values in a regular square grid into polygon bands. The result is an sf::sf object with one row for each non-empty band and can be plotted with autoplot() or hotspot_layer().

Usage

hotspot_isoband(
  data,
  value = NULL,
  breaks = 5,
  style = "equal",
  critical_p = 0.05,
  quiet = FALSE
)

Arguments

data

An sf::sf object containing a regular square grid and at least one numeric column, typically produced by a ⁠hotspot_*()⁠ function.

value

The unquoted name of the numeric column to convert. If NULL, a suitable column is inferred from the sfhotspot class of data, or from the sole numeric column in an otherwise unrecognised SF object.

breaks

Either a single positive integer giving the requested number of bands (five by default), or a strictly increasing numeric vector giving the band boundaries. If multiple boundaries are supplied, style is silently ignored.

style

Method used to calculate breaks when breaks has length one. Values other than "pconventions" are passed to the style argument of classInt::classIntervals(). "pconventions" uses two-sided normal-theory thresholds corresponding to p-values of 0.1, 0.05, 0.01 and 0.001 and ignores breaks.

critical_p

For output from hotspot_gistar() containing a kde column, the largest p-value for a cell to be included. The default is 0.05. This argument is ignored for other inputs.

quiet

If TRUE, suppress informative messages about automatically selected values and potentially unhelpful intervals.

Details

The default style = "equal" divides the observed range into equal-width intervals, providing a discrete analogue of a linear continuous colour scale. Extreme values can make equal-width bands uninformative; a message is produced when at least 90% of finite values fall in one calculated band.

For change values, Gi*/Gi statistics, and dual-KDE differences or logged ratios, automatically calculated breaks place zero at a boundary or at the centre of a band. Explicit break vectors are never altered.

When data is output from hotspot_gistar() and contains KDE values, KDE is converted and cells with pvalue >= critical_p are excluded before the bands are calculated. Without KDE values, the Gi*/Gi statistic is converted.

Values are interpolated between grid-cell centres. Lattice positions absent from a grid clipped to its original analysis area are treated as missing. Consequently, the isobands need not cover the complete area of every outer grid cell.

Value

An sf tibble with class hspt_ib. It has one row per non-empty band and columns lower, upper, band, label, and geometry. The band and label columns are ordered factors; band contains technical interval notation and label contains ranges formatted for display.

Examples


memphis_robberies_jan |>
  hotspot_kde() |>
  hotspot_isoband() |>
  autoplot()



Estimate two-dimensional kernel density of points

Description

Estimate two-dimensional kernel density of points

Usage

hotspot_kde(
  data,
  cell_size = NULL,
  grid_type = "rect",
  bandwidth = NULL,
  bandwidth_adjust = 1,
  grid = NULL,
  weights = NULL,
  transform = TRUE,
  quiet = FALSE,
  ...
)

Arguments

data

sf data frame containing points.

cell_size

numeric value specifying the size of each equally spaced grid cell, using the same units (metres, degrees, etc.) as used in the sf data frame given in the data argument. Ignored if grid is not NULL. If this argument and grid are NULL (the default), the cell size will be calculated automatically (see Details).

grid_type

character specifying whether the grid should be made up of squares ("rect", the default) or hexagons ("hex"). Ignored if grid is not NULL.

bandwidth

numeric value specifying the bandwidth to be used in calculating the kernel density estimates. If this argument is NULL (the default), the bandwidth will be determined automatically using the result of bandwidth.nrd called on the co-ordinates of data.

bandwidth_adjust

single positive numeric value by which the value of bandwidth is multiplied. Useful for setting the bandwidth relative to the default.

grid

sf data frame containing polygons, which will be used as the grid for which densities are estimated.

weights

NULL or the name of a column in data to be used as weights for weighted counts and KDE values.

transform

the underlying SpatialKDE package cannot calculate kernel density for lon/lat data, so this must be transformed to use a projected co-ordinate reference system. If this argument is TRUE (the default) and sf::st_is_longlat(data) is TRUE, data (and grid if provided) will be transformed automatically using link{st_transform_auto} before the kernel density is estimated and transformed back afterwards. Set this argument to FALSE to suppress automatic transformation of the data.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

...

Further arguments passed to kde.

Details

This function creates a regular two-dimensional grid of cells (unless a custom grid is specified with grid) and calculates the density of points in each cell on that grid using functions from the SpatialKDE package. The count of points in each cell is also returned.

Coverage of the output data

The grid produced by this function covers the convex hull of the input data layer. This means the result may include KDE values for cells that are outside the area for which data were provided, which could be misleading. To handle this, consider cropping the output layer to the area for which data are available. For example, if you only have crime data for a particular district, crop the output dataset to the district boundary using st_intersection.

Automatic cell-size selection

If no cell size is given then the cell size will be set so that there are 50 cells on the shorter side of the grid. If the data SF object is projected in metres or feet, the number of cells will be adjusted upwards so that the cell size is a multiple of 100.

Value

An sf tibble of grid cells with corresponding point counts and kernel density estimates for each cell. This can be plotted using autoplot.

References

Yin, P. (2020). Kernels and Density Estimation. The Geographic Information Science & Technology Body of Knowledge (1st Quarter 2020 Edition), John P. Wilson (ed.). doi:doi:10.22224/gistbok/2020.1.12

Examples

library(sf)

# Transform data to UTM zone 15N so that cell_size and bandwidth can be set
# in metres
memphis_robberies_utm <- st_transform(memphis_robberies_jan, 32615)

# Automatically set grid-cell size, bandwidth and neighbour distance

hotspot_kde(memphis_robberies_utm)


# Manually set grid-cell size and bandwidth in metres, since the
# `memphis_robberies_utm` dataset uses a co-ordinate reference system (UTM
# zone 15 north) that is specified in metres

hotspot_kde(memphis_robberies_utm, cell_size = 200, bandwidth = 1000)



Create a ggplot layer from hotspot results

Description

hotspot_layer() is a student-friendly wrapper around ggplot2::autolayer(). It uses the plotting method for the class of object, so results from different ⁠hotspot_*()⁠ functions are displayed using the appropriate variable and aesthetic mapping.

Usage

hotspot_layer(object, ...)

Arguments

object

An object returned by an sfhotspot analysis.

...

Arguments passed to the corresponding ggplot2::autolayer() method.

Details

Use hotspot_layer() when combining hotspot results with other ggplot2 layers. To create a complete map with suitable scales, labels and an optional base map, use hotspot_map() instead.

Value

A ggplot2 layer, or a list of layers when the selected plotting method requires more than one layer.

See Also

hotspot_map() for creating a complete map and ggplot2::autolayer() for the ggplot2 generic that this function wraps.

Examples

robbery_counts <- hotspot_count(
  memphis_robberies,
  cell_size = 0.01,
  quiet = TRUE
)

ggplot2::ggplot() +
  hotspot_layer(robbery_counts) +
  ggplot2::scale_fill_distiller(palette = "Blues", direction = 1) +
  ggplot2::theme_void()


Create an automatic map

Description

hotspot_map() creates a map from an sf object or the result of an sfhotspot analysis using reasonable defaults. Maps include an OpenStreetMap base map by default and are returned as ggplot2::ggplot objects, so they can be customised with ggplot2 layers, scales and themes.

Usage

hotspot_map(object, mapping = NULL, ...)

## S3 method for class 'sf'
hotspot_map(
  object,
  mapping = NULL,
  ...,
  basemap_type = "osm",
  basemap_zoom = NULL,
  basemap_attribution = NULL,
  caption = NULL,
  quiet = TRUE
)

Arguments

object

An sf object or an object returned by an sfhotspot analysis.

mapping

An aesthetic mapping created by ggplot2::aes(), or NULL for no aesthetic mappings. This argument can be used only for ordinary sf objects, not objects with an ⁠hspt_*⁠ class created by one of the ⁠hotspot_*()⁠ family of functions.

...

For ordinary sf objects, fixed aesthetics and other arguments passed to ggplot2::geom_sf(). For hotspot-result objects, arguments passed to the corresponding ggplot2::autoplot() method.

basemap_type

A map type passed to the type argument of ggspatial::annotation_map_tile(), or "none" for no base map.

basemap_zoom

The zoom level passed to ggspatial::annotation_map_tile(), or NULL to choose it automatically.

basemap_attribution

Attribution for the tile provider, or NULL to use the statement known for basemap_type. A non-empty string is required for a custom or otherwise unknown map type.

caption

Additional text to display in the map caption, or NULL for no additional text. Any explanatory caption generated by the plotting method is retained above this text and the base-map attribution is retained below it.

quiet

If TRUE (the default), suppress progress bars and other messages. If FALSE, show tile-download progress and other messages.

Details

For an ordinary sf object, mapping supplies aesthetic mappings for the feature layer and ... supplies fixed values and other arguments to ggplot2::geom_sf(). Hotspot-result objects determine their own aesthetic mappings, so mapping cannot be supplied for those classes.

Value

A ggplot2::ggplot object.

Examples

hotspot_map(memphis_robberies, basemap_type = "none")
hotspot_map(
  memphis_robberies,
  mapping = ggplot2::aes(colour = offense_type),
  basemap_type = "none",
  caption = "Memphis robberies"
)

Populations of census blocks in Memphis in 2020

Description

A dataset containing records of populations associated with the centroids of census blocks in Memphis, Tennessee, in 2020.

Usage

memphis_population

Format

A simple-features tibble with 10,393 rows and three variables:

geoid

the census GEOID for each block

population

the number of people residing in each block

geometry

the co-ordinates of the centroid of each block, stored in simple-features point format

Source

US Census Bureau. Census 2020, Redistricting Data summary file. https://www.census.gov/programs-surveys/decennial-census/about/rdo/summary-files.html


Memphis Police Department Precincts

Description

A dataset containing the boundaries of Memphis Police Department precincts.

Usage

memphis_precincts

Format

A simple-features tibble with 9 rows and two variables:

precinct

the precinct name

geometry

the boundary of each precinct, stored in simple-features polygon format

Licence: Public domain https://data.memphistn.gov/d/tdws-78iq

Source

City of Memphis https://data.memphistn.gov/d/rqqz-pj4u


Personal robberies in Memphis in 2019

Description

A dataset containing records of personal robberies recorded by police in Memphis, Tennessee, in 2019.

Usage

memphis_robberies

Format

A simple-features tibble with 2,245 rows and four variables:

uid

a unique identifier for each robbery

offense_type

the type of crime (always 'personal robbery')

date

the date and time at which the crime occurred

geometry

the co-ordinates at which the crime occurred, stored in simple-features point format

Source

Crime Open Database, https://osf.io/zyaqn/


Personal robberies in Memphis in January 2019

Description

A dataset containing records of personal robberies recorded by police in Memphis, Tennessee, in January 2019. This dataset is too small for some types of analysis but is included for testing purposes.

Usage

memphis_robberies_jan

Format

A simple-features tibble with 206 rows and four variables:

uid

a unique identifier for each robbery

offense_type

the type of crime (always 'personal robbery')

date

the date and time at which the crime occurred

geometry

the co-ordinates at which the crime occurred, stored in simple-features point format

Source

Crime Open Database, https://osf.io/zyaqn/


Objects exported from other packages

Description

These objects are imported from other packages. Follow the links below to see their documentation.

ggplot2

autolayer(), autoplot()


Toggle between lon/lat and UTM co-ordinates

Description

It is often useful to convert geographic data to a projected co-ordinate system, e.g. for specifying a buffer distance in metres. st_transform_auto converts between lon/lat co-ordinates and the UTM or UPS zone for the centroid of the data.

Usage

st_transform_auto(data, check = TRUE, quiet = FALSE)

Arguments

data

An sf object.

check

Should the function check if the chosen UTM zone covers all the data? If so then a warning is issued if the centroid of any row in the data is more than 1ยบ outside the area covered by the UTM zone.

quiet

if set to TRUE, messages reporting the values of any parameters set automatically will be suppressed. The default is FALSE.

Value

SF object transformed to new CRS.