DocuGate

Caching markdown from GitHub per commit, not per minute

How DocuGate caches documentation read from GitHub under the commit SHA it came from, why private repositories are never cached, and what that still costs.

By ·

The short answer

Cache content read from a repository under the commit it was read from, not under a time. A commit SHA names content that never changes, so a new commit is a new key, served on the next request with nothing to invalidate. The rule that matters more: never put content read with a private credential into a shared cache.

What a page view costs without a cache

DocuGate does not store your documentation. Every space is read from GitHub, and the sidebar needs every page's title, which lives inside each file. So reading a space is: look up the repository, look up the branch's current commit, read docugate.json, list the tree at that commit, then fetch every markdown file in the docs folder, eight at a time. A docs folder of 60 pages is about 64 GitHub requests.

Signed-out readers of public spaces are served with a server token, which GitHub allows 5,000 requests an hour; without one, it is 60. Either way, doing all of that on every page view would spend the budget on reading the same bytes again.

A time-based cache answers the wrong question

The first idea is a cache with a lifetime: keep the space for five minutes. That answers "how old may the docs be?", and every answer is wrong for somebody. Too long and the person who just fixed a typo reloads and sees the typo. Too short and the cache saves nothing.

The question we actually care about is "has the repository changed?", and Git already answers it. A commit SHA names one exact state of the tree. Content read at a SHA is the same forever, so the SHA is the key, and freshness is the cost of one request: ask GitHub what the branch points at now.

Keyed on the commit

This is getSpace in server/github.ts, shortened:

const info = await getRepo(owner, repo, token)
if (!info) return null

const ref = branch ?? info.default_branch
const sha = await getHeadSha(owner, repo, ref, token)
if (!sha) return null

const requested = normaliseDir(docsDir)
const key = info.private ? null : `tree:${owner}/${repo}@${sha}:${requested ?? '@config'}`

const content = await cached(key, async () => {
  // read docugate.json, the tree at `sha`, and every markdown file
})

Three decisions are in that one key line.

The folder is part of the key. Two spaces can point at the same public repository and read different folders. A key naming only the commit would hand one of them the other's pages. When the space does not name a folder, @config stands for "whatever docugate.json says", which is fixed for a given commit, so it is a stable key rather than a wildcard.

Everything is read at the SHA, not at the branch. The config, the tree and each file are fetched with ref=sha. If somebody pushes halfway through a read, the result is still one consistent commit, and it is stored under that commit's name.

A private repository gets no key at all. cached(null, load) just calls load.

Private content is never cached

The cache is a Map in the server process, shared by every request that process handles. Content read with a reader's own token, or the owner's, put in that map would be served to the next visitor without either. The comment above the cache says it in two lines: public content is shared between readers, private content never is.

We could have built a per-user cache instead. We did not, because the per-user key is exactly the kind of logic that is easy to get wrong once and expensive forever. The decision instead lives in one place, a null key, and is easy to check.

The same thinking keeps two other things out of the cache:

  • Whether the reader may edit. GitHub reports permissions.push on the repository read, but it describes the token, not the repository. It is added to the result after the cache, so one reader's write access is never handed to the next.
  • Merged spaces. A space can merge several repositories. The merge is not cached as a unit; each repository is cached under its own SHA. A space that merges a public repository with a private one therefore caches the public half and reads the private half live, with no new cache logic.

Failures are not cached

return load().then((value) => {
  // Never cache a failure.
  if (value == null) return value
  if (cache.size > 500) cache.clear()
  cache.set(key, { at: Date.now(), value })
  return value
})

This line was a fix. Before it, a moment's trouble at GitHub was cached as an empty result, and in a merged space that removed a whole section from the sidebar for five minutes after GitHub had recovered. It shipped in the changelog on 2026-09-24. The OpenAPI reader follows the same rule: a failed download is not the repository's answer, so it is not stored.

The trade-off

What we have is simple, and simple has costs we can name.

  • Two requests on every view, cache hit or not. The repository lookup and the head-commit lookup are not cached; they are what makes freshness free. A cached public space still costs two GitHub calls per view.
  • Private spaces pay full price. Every view reads every file. That is why a docs folder is capped at 300 files per source and 600 per space, and why server/db.ts notes that private repositories are both the paid feature and the expensive one.
  • The cache is per process. It is an in-memory map, and the API runs on serverless instances that share nothing but the database. Each instance warms its own copy, and a cold instance starts empty.
  • Eviction is blunt. Past 500 entries the whole map is cleared, not the oldest entry.

What we would do differently

  • Treat the five-minute lifetime as what it is. Entries carry a five-minute TTL, but under a SHA key the content cannot go stale, so the TTL only forces a re-read of identical bytes. It survives as a crude memory bound. A least-recently-used map with a size limit would bound memory properly and drop both the TTL and the clear-everything rule.
  • Make the freshness check cheaper. GitHub's REST API supports conditional requests with ETags. We do not send them yet. Remembering the last response for the head-commit lookup and asking "has this changed?" is the obvious next step for the two calls every view still makes.
  • Share the cache once traffic justifies it. A store outside the process would let instances share warm entries. We have not added one: a SHA-keyed cache needs no coordination to be correct, only to be efficient, and that can wait for traffic that shows it is needed.

Questions people also ask

How quickly does a pushed change appear on a DocuGate space?

On the next request. The server asks GitHub for the branch's current commit on every read, and a new commit is a cache key it has not seen, so the pages are read again.

Does DocuGate cache private repositories?

No. Content read from a private repository is never put in the cache, so every page view of a private space is a live read from GitHub.

Why key on the docs folder as well as the commit?

Two spaces can read different folders of the same public repository. A key that named only the commit would serve one of them the other's pages.

What happens if GitHub fails during a read?

Nothing is cached. A failure is returned to that request only, so a brief problem at GitHub does not hide a section of the docs for the length of the cache.

Read next

Get started with DocuGate, free.