Disk caching
Results that are expensive to compute can also be kept on disk, so that they survive the process. The disk is a second level below the cache in memory: a call looks in RAM first (as chosen by CacheStyle), then on disk, and only computes when both miss, after which the result is written to both.
Disk caching is a package extension on SQLite.jl: load it with using SQLite, or, in a package, add SQLite to its dependencies and import SQLite.
using MemoizationKit, SQLite
computations = Ref(0)
@cached function expensive(n::Int)::Matrix{Float64}
computations[] += 1
[Float64(i == j) for i in 1:n, j in 1:n]
end
MemoizationKit.DiskCacheStyle(::typeof(expensive), args...) = DiskCache()
expensive(2)2×2 Matrix{Float64}:
1.0 0.0
0.0 1.0Clear RAM and call again to read the stored result from disk:
empty_caches!(expensive)
expensive(2)
computations[] # the body ran only once1Choosing what goes to disk
DiskCacheStyle(f, args...) selects the disk strategy per function and argument types, independently of CacheStyle: NoCache() (the default) or a DiskCache. Both resolve at compile time, so functions without a disk cache are not affected, and a RAM hit costs the same with or without one.
CacheStyle | DiskCacheStyle | a call |
|---|---|---|
GlobalCache() (default) | NoCache() (default) | RAM, else compute |
GlobalCache() | DiskCache() | RAM, else disk, else compute |
NoCache() | DiskCache() | disk, else compute, on every call |
Where the data goes
Each function has one SQLite database per machine, <Module>.<f>-v<version>-<host>.sqlite. The directory is the disk_path preference of the function, its package, or MemoizationKit (see Configuration), and by default a scratch space of the package that owns f. Functions outside packages (in scripts or the REPL) use a scratch space of MemoizationKit.
set_cache_preferences!(MyPackage; disk_path = "/path/to/cache")
println(read("LocalPreferences.toml", String))
[MyPackage.MemoizationKit]
disk_path = "/path/to/cache"- One file per function and machine, whatever the number of entries, which matters on file systems that limit the number of files.
- Processes on one machine share the database, reading and writing at the same time. A write is a transaction, so an interrupted process never leaves a partial entry.
- Machines write separate files, named after the host, since SQLite's locking does not work across machines on a shared file system. Results are not shared between machines.
Disk lookups take file locks, which can be slow on network storage. Use a local disk_path when possible, and reserve disk caching for expensive computations. Disk storage has no automatic size limit or eviction; use empty_disk_caches! to reclaim space in the current database.
Versions
The file name holds a version, MemoizationKit.diskversion(f), "1" by default. Keeping it current is up to you: bump it when the results of f change, and the old file is no longer read.
MemoizationKit.diskversion(::typeof(expensive)) = "2"
MemoizationKit.diskversion(expensive)"2"The default format is that of Serialization, which is not guaranteed to be readable by other Julia versions, nor after the definition of a stored type changes. Bump the version when that happens, or write a stable format with a serializer of your own (below).
Entries that cannot be read, or are not of the value type of the call, count as misses and are overwritten. Errors of the disk itself, such as a full disk, never fail a call: they are reported once, and the call goes on without the disk.
Formats
Keys and values are written with Serialization, indexed by the SHA-256 hash of the serialized key; the key is stored too, to guard against hash collisions.
MemoizationKit.cachekey selects the key for both RAM and disk caching. To share disk entries between equivalent inputs, return a canonical representation that serializes to the same bytes. Hashed changes RAM hashing and equality, but its wrapped value is still serialized, so custom equality alone does not merge disk entries. See Custom cache keys for examples and the equality contract.
The serializer is a parameter of the style, DiskCache(; serializer = Serializer). To choose the format of some types, define a serializer type with its own serialize and deserialize methods for them; everything else is written as by Serialization. A serializer is a mutable AbstractSerializer with these fields and a constructor from an IO:
using Serialization
using Serialization: AbstractSerializer
mutable struct MySerializer{I <: IO} <: AbstractSerializer
io::I
counter::Int
table::IdDict{Any, Any}
pending_refs::Vector{Int}
version::Int
MySerializer(io::I) where {I <: IO} = new{I}(io, 0, IdDict(), Int[], 0)
end
struct Point
x::Int
y::Int
end
function Serialization.serialize(s::MySerializer, p::Point)
Serialization.writetag(s.io, Serialization.OBJECT_TAG)
serialize(s, Point)
serialize(s, "point $(p.x) $(p.y)")
end
function Serialization.deserialize(s::MySerializer, ::Type{Point})
_, x, y = split(deserialize(s)::String)
Point(parse(Int, x), parse(Int, y))
end
@cached point(x::Int) = Point(x, 2x)
MemoizationKit.DiskCacheStyle(::typeof(point), ::Int) = DiskCache(; serializer = MySerializer)
point(4)
empty_caches!(point)
p = point(4) # read with the custom serializer
(p.x, p.y)(4, 8)writetag and OBJECT_TAG are internals of Serialization, used the same way by Distributed's ClusterSerializer.
Turning it off
disable_disk_caches!() turns all disk caches off in this process, and enable_disk_caches!() back on; while off, nothing is read from or written to disk. The disk = false preference turns them off per function, package, or globally:
set_cache_preferences!(MyPackage; disk = false)
println(read("LocalPreferences.toml", String))
[MyPackage.MemoizationKit]
disk = false
disk_path = "/path/to/cache"disk and disk_path are read when a function first uses its disk cache in a session. Disk caches are never used during precompilation.
Managing disk caches
info = only(disk_cache_info(expensive)).second
(; entries = info.entries, bytes = info.bytes)(entries = 1, bytes = 24728)Empty the database and check the entry count:
empty_disk_caches!(expensive)
only(disk_cache_info(expensive)).second.entries0disk_cache_info(MyPackage) lists disk caches of functions owned by a module and its submodules. MemoizationKit.disk_cache_stats reports hit/miss counters for this process.
disk_cache_info and empty_disk_caches! act on this machine's current database; files of older versions or other machines are left alone. The Disk tab of the dashboard lists the open disk caches with these numbers.
Precomputed results as an artifact
A package can ship precomputed results, read-only, as a Pkg artifact: fill the disk cache, export it with export_disk_cache, and point MemoizationKit.disk_artifact at the artifact.
using Pkg.Artifacts
artifact_hash = create_artifact() do dir
empty_caches!(expensive) # make every call reach disk
foreach(expensive, 1:3) # fill the disk cache
export_disk_cache(expensive, dir)
end
readdir(artifact_path(artifact_hash)) # the exported database1-element Vector{String}:
"Main.expensive-v1.sqlite"Point the function at the exported artifact, then empty RAM and the node database:
MemoizationKit.disk_artifact(::typeof(expensive)) = artifact_path(artifact_hash)
empty_caches!(expensive)
empty_disk_caches!(expensive)
before = computations[]
result = expensive(2)
println((; result, new_computations = computations[] - before))(result = [1.0 0.0; 0.0 1.0], new_computations = 0)To ship the artifact, archive, upload, and bind it in your package's Artifacts.toml. Its package definition can then return artifact"expensive" from MemoizationKit.disk_artifact.
A call then looks in RAM, the artifact, and this machine's database, in that order, and only then computes. New results go to the database; the artifact is never written, and holds the results of one version of f.