Engram · Fast path

Engram

The fast path in front of your expensive calls.

Engram is a lightweight retrieval layer that sits between a query and the expensive work behind it: an LLM call, a database hit, an external API. When a query matches something Engram already holds, it returns that answer immediately and the costly call never happens. It is keyword-indexed, learns which answers actually get used, and is small and transparent by design.

Open source · Apache-2.0 · available now

The problem

The same question, paid for twice.

Conversational systems ask the same things over and over, and every repeat pays full price: another inference, another query, another round trip.

The usual fix is a cache, but a naive cache matches exact strings and breaks on the first reworded question. A vector store answers every near-miss with an opaque similarity score you cannot read and cannot tune by hand.

You want the speed of a cache with retrieval you can actually reason about.

What it does

It answers before the expensive call runs.

Engram sits in front of expensive computation. A sufficient match returns stored knowledge and skips the downstream call entirely.

  • Matches on keyword overlap, so retrieval is fast and predictable, with lemmatization, stemming, and synonym fallbacks that catch reworded questions an exact-match cache would miss.
  • Learns from use. Hit-rate tracking records which stored answers actually get retrieved, a signal that sharpens retrieval over time and surfaces dead weight to prune.
  • Separates the permanent from the disposable. Protected statements stay put; evictable ones age out under a capacity bound, so durable facts survive while transient ones clear.
  • Carries a conversation. Sessions expand context across turns, so a follow-up like what is its population resolves against what came before.
  • Persists as plain JSON. Save, load, inspect, and move state without ceremony.

Why it’s different

Retrieval you can read.

A vector store hides its judgment inside an embedding distance you have to trust. Engram is deliberately legible: the reason it returned an answer is something you can read, predict, and tune by hand.

There is an optional semantic resolver for teams that want one. It is off by default, it runs a pinned model on your own hardware, and the keyword path runs ahead of it either way. The retrieval you get out of the box stays readable end to end.

Legible

Keyword-first retrieval you can read and tune, rather than a black-box embedding distance.

Learns from use

Sharpens on what actually gets retrieved, rather than guessing at relevance.

Stays small

A library and a command-line tool you run yourself, not a service to stand up.

You own it

It runs where you run.

Engram is a library and a CLI, not a service you have to host. State persists as plain JSON you can inspect, version, and move. Nothing leaves your environment, and there is nothing to stand up. It is Apache-2.0 licensed, so you can read the source, fork it, and ship it.

Available now · Apache-2.0

Put it in front of your next call.

Engram is open source and usable as a library or from the command line. Read the source, run it yourself, and skip the calls you have already paid for.

View on GitHub
View on GitHub