Hold the build conditions constant
Use a repository you can build repeatedly, one commit, one task command, and the same runner image, CPU, memory, region, tool versions, and dependency lockfile. Install dependencies before starting the build timer. Record checkout and installation separately if you also want the end-to-end CI result.
- Use isolated runners or disposable checkouts. Do not clear a shared production cache to create a baseline.
- Keep the task set fixed. An affected-only run that selects fewer tasks is a different experiment.
- Keep package downloads equally warm or equally cold in every run; they are not task-output cache hits.
- Do not restore local build caches or generated outputs into the remote replay runner.
- Check that restored outputs pass the same tests as freshly built outputs.
Run the three-part experiment
1. Measure an uncached baseline
On a fresh runner, disable task-cache reads and writes using your build tool's documented options. Run the chosen task command and record its elapsed time, task count, and exit status. Keep this command and configuration with the results.
2. Seed a separate remote cache
Use a dedicated benchmark workspace and a trusted write token. On another fresh runner at the same commit, enable remote caching and run the same tasks. Record this cold-cache run separately: it includes execution and uploads, so it may take longer than the baseline.
3. Replay from a fresh runner
Use the same remote workspace with a read-only token on a third fresh runner. Keep remote reads enabled and local build outputs absent. Run the same tasks and verify remote hits in the tool's output. A second build on the seed runner may use its local cache and is not proof of remote reuse.
Repeat the baseline and replay at least five times on equivalent fresh runners. Alternate their order to reduce time-of-day bias. Keep every sample, report the median and range, and explain failures or exclusions. A seed run is preparation, not a warm-replay sample.
Use the native setup and verification instructions for Nx, Lerna, Turborepo, Gradle, or Bazel. On a Linux CI runner, wrap the actual build command with /usr/bin/time -p and use its real value for elapsed seconds.
Calculate hit rate and time saved
Remote task hit rate = remote-hit tasks / tasks that attempted a remote lookup
Elapsed time saved = median uncached seconds - median remote replay seconds
Elapsed reduction (%) = 100 * elapsed time saved / median uncached secondsState the denominator. HTTP hit rate, task hit rate, and the share of all selected tasks that reused remote outputs measure different things. Bazel may read action metadata and several CAS blobs for one action; counting those requests as separate saved builds inflates the result. Exclude local hits from the remote-hit numerator and report non-cacheable tasks separately.
Sum of avoided task durations measures compute work. It is not pipeline wall time when tasks run in parallel. For monetary savings, use actual runner billing and the measured critical-path reduction; the CI savings calculator is an estimate, not a benchmark result.
Record enough to reproduce the result
Measured on (UTC):
Repository revision and task command:
Build tool, runtime, and runner image versions:
Runner CPU, memory, region, and concurrency:
Dependency cache state:
Local task cache state:
Remote backend, region, and read/write policy:
Baseline seconds (all samples):
Seed seconds:
Remote replay seconds (all samples):
Remote-hit / remote-lookup task counts:
Output verification command and result:
Median, range, exclusions, and limitations:Keep tokens, private repository paths, and customer artifacts out of anything you publish. If publishing an aggregate, also state the exact date range, number of workspaces, inclusion rule, and whether you averaged workspace rates or pooled all requests. This page provides a method; it does not claim a measured Cachely speedup.
Decide whether the cache helps
An identical-commit replay measures the best opportunity for reuse. Also repeat the experiment with a representative code change, test-only change, and dependency update. Include a cold-cache build in the decision: upload overhead, artifact size, network latency, and tasks that always run can outweigh reuse on small projects.
If the replay misses, work through cache miss diagnosis. If it hits but is slower, compare download and extraction time with execution time before caching more outputs. Keep remote caching where repeated work costs more than transfer.