Starting with misha 5.3.0, databases can be stored in two formats:
The indexed format provides better performance and scalability, especially for genomes with many contigs (>50 chromosomes).
The indexed format uses unified files:
Sequence data: - seq/genome.seq - All
chromosome sequences concatenated - seq/genome.idx - Index
mapping chromosome names to positions
Track data: -
tracks/mytrack.track/track.dat - All chromosome data
concatenated - tracks/mytrack.track/track.idx - Index with
offset/length per chromosome
Advantages: - Fewer file descriptors (important for genomes with 100+ contigs) - Better performance for large workloads (14% faster) - Smaller disk footprint - Faster track creation and conversion
The per-chromosome format uses separate files:
Sequence data: - seq/chr1.seq,
seq/chr2.seq, … - One file per chromosome
Track data: -
tracks/mytrack.track/chr1.track, chr2.track, …
- One file per chromosome
When to use: - Compatibility with older misha versions (<5.3.0) - Small genomes (<25 chromosomes) where performance difference is negligible
By default, new databases use the indexed format:
Use gdb.info() to check your database format:
gdb.init_examples() # the bundled examples database
info <- gdb.info()
print(info$format) # "indexed" or "per-chromosome"
#> [1] "per-chromosome"The full record:
Convert all tracks and sequences to indexed format:
This will: 1. Convert sequence files (chr*.seq →
genome.seq + genome.idx) 2. Convert all tracks to indexed
format 3. Validate conversions 4. Remove old files after successful
conversion
Convert specific tracks while keeping others in legacy format:
Convert interval sets to indexed format:
# Only "big" interval sets are stored per chromosome, so lower the threshold
# to get big sets out of the small example database (default: 1,000,000).
options(gbig.intervals.size = 10)
gintervals.save("myintervals", gscreen("dense_track > 0.3"))
gintervals.save("my2dintervals", gextract("rects_track", gintervals.2d.all())[, 1:6])
options(gbig.intervals.size = 1e6)
# 1D intervals
gintervals.convert_to_indexed("myintervals")
# 2D intervals
gintervals.2d.convert_to_indexed("my2dintervals")High priority (significant benefits): - Genomes with many contigs (>50 chromosomes) - Large-scale analyses (10M+ bp regions frequently) - 2D track workflows - File descriptor limit issues
Medium priority (moderate benefits): - Repeated extraction workflows - Regular analyses on medium-sized regions (1-10M bp)
Low priority (minimal benefits): - Small genomes (<25 chromosomes) - One-off analyses - Simple queries on small regions
Step 1: Backup (optional but recommended)
# eval = FALSE: copies a whole database outside the vignette's temp directory.
# Create backup of important database
system("cp -r /path/to/mydb /path/to/mydb.backup")Step 2: Check current format
gdb.init_examples()
info <- gdb.info()
print(paste("Current format:", info$format))
#> [1] "Current format: per-chromosome"Step 3: Convert
Step 4: Verify
# Check format changed
info <- gdb.info()
print(paste("New format:", info$format))
#> [1] "New format: indexed"
# Test a few operations
result <- gextract("dense_track", gintervals(1, 0, 1000))
print(head(result))
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1Step 5: Remove backup (after validation)
You can freely copy tracks between databases with different formats.
gtrack.copy()gtrack.copy(src, db = target_db) is the idiom. It reads
the source track directly, reconciles per-chromosome vs. indexed storage
and any difference in the two databases’ chromosome sets, and preserves
the track’s type. src may be a character vector, so a batch
copy is a single call.
gdb.init_examples()
source_db <- .misha$GROOT
# A second database to copy into
target_db <- file.path(tempdir(), "target_db")
unlink(target_db, recursive = TRUE)
gdb.create_linked(target_db, parent = source_db)
#> Created linked database at /dev/shm/aviezerl-rbuild/RtmpQxWCjh/target_db (linked to /dev/shm/aviezerl-rbuild/RtmpQxWCjh/trackdb/test)
# One track, or many
gtrack.copy("dense_track", db = target_db)
gtrack.copy(c("sparse_track", "array_track"), db = target_db)
gsetroot(target_db)
gtrack.ls()
#> [1] "array_track" "dense_track" "sparse_track"
gtrack.info("dense_track")[c("type", "size.in.bytes")]
#> $type
#> [1] "dense"
#>
#> $size.in.bytes
#> [1] 80012The destination does not have to be in the same format: converting it afterwards leaves the tracks readable.
gdb.convert_to_indexed()
gdb.info()$format
#> [1] "indexed"
head(gextract("dense_track", gintervals(1, 0, 300)))
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1Exporting with gextract(..., file = ...) and
re-importing with gtrack.import() also works, but it is
lossy: binsize = 0 builds a sparse track,
so a dense source track comes back sparse and about 5x larger (20 bytes
per value against 4 bytes per bin). Reach for it only when the
destination is not a misha database, or when you actually want a
different track type.
gsetroot(source_db)
exported <- file.path(tempdir(), "dense_track.txt")
gextract("dense_track", gintervals.all(), iterator = "dense_track", file = exported)
gsetroot(target_db)
gtrack.import("dense_track_roundtrip", "Round-tripped", exported, binsize = 0)
gtrack.info("dense_track")[c("type", "size.in.bytes")]
#> $type
#> [1] "dense"
#>
#> $size.in.bytes
#> [1] 80012
gtrack.info("dense_track_roundtrip")[c("type", "size.in.bytes")]
#> $type
#> [1] "sparse"
#>
#> $size.in.bytes
#> [1] 400120
gtrack.rm("dense_track_roundtrip", force = TRUE)Based on comprehensive benchmarks comparing indexed vs legacy formats:
# Work with both formats in same session
gsetroot(source_db) # per-chromosome
data1 <- gextract("dense_track", gintervals(1, 0, 1000))
head(data1)
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1
gsetroot(target_db) # indexed
data2 <- gextract("dense_track", gintervals(1, 0, 1000))
head(data2)
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1This occurs with many-contig genomes in legacy format:
Solution: Convert to indexed format
After manually copying track directories:
Solution: Reload database
gdb.create_genome() for standard genomesgdb.create() with multi-FASTA for custom
genomesgdb.info()gdb.convert_to_indexed()