Blog · September 4, 2026 · Updated September 4, 2026
A similarity score is not a citation
Nearest-neighbor confidence is not documentary authority. Production AI needs named sources you can open, version, and revoke.
BY HIVE FORENSICS AI

Nearest-neighbor confidence is not documentary authority. Production AI needs named sources you can open, version, and revoke — not a cosine that looks like proof.
Nearest-neighbor confidence is not documentary authority.
Teams ship retrieval stacks and treat a high similarity score as a citation. The UI shows a percentage. The demo nods. Procurement hears "grounded." What they actually bought is a ranking heuristic over chunks — not a named document, not a version, not a path counsel can reopen.
Hive Forensics AI builds for buyers who need the opposite: answers that point at sources you can hash, open again, and revoke. A score is a hint. A citation is a claim you can defend.
What a citation actually is
A citation names a source in scope. You can say which document, which version, and whether that source was allowed for the run. Someone else can open the same evidence later. If the source must leave the system, you revoke it — you do not wait for an embedding to "forget."
A similarity score is a distance in a vector space. It can be useful for candidate recall. It does not prove the model read the right page, that the page was current, or that the workflow was allowed to use it. Dressing the score up as a footnote only makes the liability quieter until an audit starts.
Why buyers mix them up
Vendors sell retrieval as grounding. Demos look sharp when the top chunk matches the question. Operators see a green bar and stop asking for the receipt. When SIU, fraud review, or a security owner asks what the system knew, "0.87 cosine" is not an answer.
You need a portable unit of knowledge — a Knowledge Image you can hash and replay — plus a runtime that fails closed when the cited source is missing, stale, or revoked. Retrieval can help find candidates. It does not replace ownership of the corpus or the evidence path.
That is the work Hive Forensics AI ships. Production systems on a verifiable knowledge foundation, with receipts and default-deny controls where the workflow demands them. It is not a hosted chatbot pitch. The commercial path stays the same: prove one real workflow first.
How work starts
Unscoped AI programs grow indexes. They rarely grow citations you can defend.
We do not sell chatbot SKUs, hour rental, or brochure demos with no owned outcome. Work starts on a five-day Bootcamp: one corpus, one workflow, representative source material, and a Friday recommendation with an evidence path. If the citation boundary holds, a bounded deploy can follow. If it does not, you learn that early, on purpose.
If your workflow cannot afford an answer justified only by a similarity score, start there.