Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to ground an AI assistant in your own documents

Reading time
7 minutes
Assumes
You have documents an assistant should answer from
Updated
Sep 6, 2026

What grounding fixes, and what it doesn't

An ungrounded model answers from what it absorbed in training. For anything specific to your business — your refund window, your escalation path, your product's actual limits — it has no information, so it produces something plausible. Plausible is the dangerous failure, because it reads exactly like knowledge.

Grounding fixes this by retrieving relevant passages from your own material and requiring the answer to come from them.

What it does not fix is a model that reasons badly, a question your documents don't answer, or documentation that is itself wrong. That third one catches teams by surprise: grounding faithfully surfaces your out-of-date policy page, and now the wrong answer arrives with a citation attached, which is worse than the confident guess because it is more convincing.

Before you build anything

Read ten of the documents you're about to index, chosen at random. If more than one is out of date, fix the documents first. Retrieval will amplify whatever is in there, including the parts your team has been quietly working around.

Choose a narrow corpus on purpose

The instinct is to index everything, on the theory that more knowledge is better. More documents means more chances for the retriever to surface something adjacent-but-wrong, and adjacent-but-wrong is the hardest failure to catch, because the answer is topically correct and factually inapplicable.

Start with the smallest corpus that answers the questions you actually get. For a support assistant that is usually the policy documents and the top fifty help articles, not the entire wiki, and definitely not five years of internal strategy decks.

Add material when you can point at a real question it would have answered. That discipline keeps the corpus small enough to audit, and auditability is what lets you diagnose a bad answer in ten minutes rather than a day.

Chunking is where quality is won or lost

Retrieval doesn't fetch documents, it fetches passages. How you cut them up determines what the model sees.

Too small and passages lose their context. A chunk reading "this does not apply to enterprise customers" is worse than useless when the sentence it qualifies is in the previous chunk.

Too large and the relevant sentence arrives buried in three pages of adjacent material, where it competes with everything else for the model's attention.

Cut across a boundary and you get the most insidious version: a chunk containing the end of one policy and the start of another, which reads as a single coherent rule that does not exist.

The heuristic that survives contact with real documents: cut on structural boundaries first — sections, headings, list items — and only fall back to length when a section is genuinely too long. Overlap adjacent chunks slightly so a sentence near a boundary appears in both.

Tell the model what to do when the answer isn't there

The single highest-value instruction in a grounded system is the one covering absence. Without it, a model handed passages that don't contain the answer will answer anyway, from training, in the same confident register it uses when the passage supports it.

Say explicitly: answer only from the provided context; if the context does not contain the answer, say so and hand off. Then test that path deliberately, by asking something you know is not in the corpus. This is the test teams skip and the behavior users hit in week one.

The related instruction worth adding: cite the passage. Not for the user's benefit primarily — for yours. When an answer is wrong, a citation tells you in seconds whether the retriever fetched the wrong passage or the model misread the right one. Those are completely different bugs with completely different fixes.

Diagnose retrieval and generation separately

When a grounded assistant gives a bad answer, there are exactly two places to look, and conflating them wastes days.

Retrieval failed. The right passage was never fetched. Look at what came back for that query. If the passage isn't there, no amount of prompt work will help — the fix is chunking, the corpus, or how the query is formed.

Generation failed. The right passage was fetched and the answer still contradicts it. Now it is a prompt problem, or the passage was ambiguous, or it arrived alongside a more prominent distractor.

Always check retrieval first. It is faster to check and it is wrong more often. Teams that skip this step spend a week rewriting prompts to fix a chunking problem.

Test grounding with questions that should fail

Your test set needs three kinds of question, and most teams only write the first.

Answerable questions. The passage exists and the answer is unambiguous. These confirm the basic loop works.

Absent questions. Reasonable questions your corpus does not cover. The correct answer is a decline and a handoff, and this is the category that catches regressions when somebody loosens a prompt to make the assistant "more helpful."

Adjacent questions. Questions where a topically similar but inapplicable passage exists — the enterprise policy when the user is on the standard plan, last year's rate when they need this year's. These are the hardest and the most representative of real traffic.

If your evaluation set contains only the first kind, it will report high scores right up until a customer asks something you didn't index.

Common mistake

Judging a grounded assistant on whether the answer sounds right. Judge it on whether the answer follows from the retrieved passage. A response that is correct despite the passage is a model recalling training data, and it will be wrong the next time on the same setup.

Keep the corpus current

A knowledge base is a copy, and copies drift. The policy page gets updated and the indexed version doesn't, which returns you to confidently-wrong-with-a- citation.

Two things prevent this. Decide who owns re-ingestion when a source document changes, and write it down — this is a process question rather than a technical one, and it is where most grounded systems quietly rot. And keep a handful of evaluation items whose answers you know change, so a stale index shows up as a failing test rather than as a customer complaint.

Before you point users at it

0 of 6 checked