Ragable

← All articles

From Pilot to Production: Rolling Out a Document Assistant

5 min read
From Pilot to Production: Rolling Out a Document Assistant

Every document assistant project we have run passed its pilot comfortably, then hit a completely different set of problems the week after. The gap isn't model quality. A pilot and a production rollout answer separate questions, and the second one is mostly about documents, permissions and operations rather than retrieval mathematics. What follows is the sequence we now use, in the order the decisions actually arrive.

Why the Pilot Was Never the Hard Part

A pilot answers one narrow question: can the assistant find the right passage in a small curated set and cite it. Production asks something harder. Does it keep doing that when the corpus triples, permissions differ per user, and nobody curated anything? Pilots run on files someone hand-picked last week. Production runs on a shared drive untouched since 2019. A demo audience forgives a wrong answer. The finance team filing a claim does not. So we treat the pilot as a feasibility test, not a dress rehearsal. The engineering starts once it passes.

Decide What the Assistant Is Allowed Not to Know

Retrieval augmented generation answers from your documents, so its honest failure mode is "I could not find this" rather than a fluent invention. Define the refusal threshold before launch: below what retrieval confidence should it decline? Scope the corpus deliberately too. Dump every company PDF into the index and you dilute ranking, burying the material people actually ask about. Each answer carries the passage it came from, which turns verification from a trust exercise into a two-second click.

Tip: write down five questions the assistant should refuse and test them on every release. Silent scope creep shows up there first.

Cloud or Your Own Infrastructure: Making the Call Once

The deciding factor is usually contractual, not technical: what your data classification and existing agreements permit. Self-hosted deployment keeps documents, embeddings and query logs inside the customer network, including the model itself where that is required. Cloud deployment takes GPU capacity planning, model upgrades and index maintenance off the customer's plate. What differs between them:

  • Latency profile: local hardware limits versus provider-side scaling
  • Upgrade cadence: scheduled by you, or by the platform
  • Vector index custody: customer network or managed service
  • Runtime patching: your team or ours

Migrating later works. It also costs a full re-index and a fresh evaluation run, so decide before the corpus grows.

Ingestion Is Where Production Quality Is Won or Lost

Real corpora are scanned contracts, spreadsheets used as databases, decks with text trapped in images, and six versions of one policy. Our pipeline:

  1. Inventory the sources
  2. Decide the authoritative version of each document
  3. Extract and chunk
  4. Attach metadata
  5. Index
  6. Verify a sample by hand

Metadata does more work than chunk size. Owner, effective date, department and superseded-by let the assistant filter before it ranks. Withdrawn documents are the quiet failure here: citing a retired policy is worse than saying nothing at all. And plan re-ingestion on day one, because a one-off import goes stale within a quarter.

Permissions Are Part of Retrieval, Not a Layer on Top

Apply access control after retrieval and the model has already read the restricted passage, ready to leak it through paraphrase. Filter at the index instead. Every chunk carries the access groups of its source, and each query runs only against the caller's permitted subset. Two people asking the same question should get different answers, and that has to be designed behaviour rather than a surprise. Query logs hold document content and user intent, so they inherit the corpus retention rules. Test with a low-privilege account - admin logins hide exactly the bugs that matter.

Build the Evaluation Set Before You Need It

Collect genuine questions from pilot users, pair each with the passage that should surface, then freeze the result as a regression set. Measure retrieval separately from generation. If the right chunk never appeared, no prompt tweak will rescue the answer. Run the set on every corpus update and model change, because both move results quietly.

Tip: keep a handful of adversarial questions - near-duplicates, outdated policies, questions with no answer in the corpus. They expose regressions that ordinary questions sail straight past. Human review covers a defined sample, and that rate falls as the set earns trust.

Rollout, Support and the First Month

Launch to one department with a named owner who can actually fix documents. Not to everyone at once. Give people a one-click way to flag a bad answer and route those flags to whoever maintains the corpus. Most early complaints turn out to be document problems rather than model problems: a missing file, a stale version, a scan nobody could read. Watch which questions get asked and which get refused, because the refusal log is the highest-value backlog you will have. And say plainly what this is - a fast reader of your documents, not an oracle, with citations there to be checked.

Frequently Asked Questions

How long does it take to move from pilot to production?

That depends almost entirely on document readiness, not on the assistant. Retrieval and generation are configuration work. The corpus cleanup is what sets your schedule.

Can the assistant run without sending anything to an external provider?

Yes. A fully self-hosted deployment keeps documents, index, model and logs on your own infrastructure.

What happens when the assistant does not know the answer?

It says so and points to the closest material it found. That is intended behaviour, not something to tune away.

What We Would Do Again

The sequence holds up: scope the corpus, settle deployment early, invest in ingestion and metadata, wire permissions into retrieval, freeze an evaluation set, then roll out narrowly. The recurring lesson is that a document assistant is a document project with a model attached, not the reverse. Citation-backed answers change adoption too, because someone who can verify in seconds stops treating the tool as a black box. Give corpus maintenance a named owner and treat it as ongoing operational work.

Build your own AI assistant on your documents

Ragable indexes your files and answers from them, with citations. Start on SaaS or run it on your own infrastructure.

Start from $99/month