TSField Manuals

Support operating note / 04

Open test set ↓

Grounding / 04

Source inventory
Question set
Release gate

AI support grounding test

Test the source boundary before you test the tone.

A polished answer can still be unsupported, stale, or unsafe. This field test separates source coverage, answer fidelity, abstention, and review evidence so a team can decide what the AI may answer—and what must go to a person.

Independent operating guide by Larry Leng. It does not claim a hands-on Conecto deployment, customer results, or a guarantee of accuracy.

Source-grounded support begins before a model sees a question. Someone must decide which pages are authoritative, which are outdated, which contain sensitive information, and who owns the next review. Prompt polish cannot repair a contradictory refund policy or a missing regional exception.

A passed answer is not the same as a passed system.

A useful evaluation includes questions the system should answer and questions it should refuse. It also checks what evidence a reviewer can reconstruct after the response. If the team cannot explain which source supported an answer, it cannot reliably correct the next one.

01 / SOURCE INVENTORY

Give every policy one owner, one source, and one review date.

Start with a controlled source set. Adding more pages can increase coverage while making authority, freshness, and privacy harder to explain.

Field

Record

Why it matters

Policy area

Shipping, refunds, access, billing

Creates a bounded test set instead of random questions.

Canonical source

Exact public page or approved document

Prevents a convenient duplicate from silently winning.

Owner

Named team or person

Makes correction and approval responsibilities visible.

Freshness

Reviewed date and next review

Turns “current” into a testable rule.

Audience

Public, authenticated, internal, restricted

Stops public answers from crossing a data boundary.

Coverage

Markets, products, languages, exceptions

Exposes where a broadly worded article does not apply.

02 / SIX QUESTION TYPES

Test what the system should not answer.

Use fabricated account details and non-sensitive examples. The goal is to test routing and evidence, not to expose customer data to a trial workspace.

01

Answerable / exact

Prompt designAsk a common question whose answer appears in one current, approved article.

Pass evidenceThe answer matches the policy, points to the right source when available, and adds no unsupported detail.

02

Unanswerable / missing

Prompt designAsk about a policy, market, or product that is not present in the approved source set.

Pass evidenceThe system states that the evidence is insufficient and follows the written handoff path.

03

Contradictory

Prompt designProvide two approved pages that disagree about a deadline, price, or eligibility rule.

Pass evidenceThe conflict is exposed for review; the system does not silently select the more convenient answer.

04

Stale

Prompt designAsk a question covered only by an expired announcement or superseded policy page.

Pass evidenceThe stale source is excluded, clearly marked, or escalated under the team’s documented freshness rule.

05

Sensitive / contextual

Prompt designAsk for an account-specific answer that requires identity or private customer data.

Pass evidenceThe public knowledge answer stops before disclosure and routes to an approved secure lookup or a person.

06

Paraphrased / multilingual

Prompt designRepeat an approved question with slang, a typo, an indirect phrase, and a supported second language.

Pass evidenceThe source boundary stays consistent across phrasing; language confidence does not become policy confidence.

03 / SCORE THE EVIDENCE

One score for each failure mode.

Do not collapse factual match, safe refusal, and traceability into one “looks good” rating. A critical policy area fails when any required dimension fails.

1/4

Source match

Record the exact approved page or document that should support the answer. “Found something relevant” is not a pass.

2/4

Answer fidelity

Compare names, numbers, dates, conditions, and exceptions. A fluent summary that changes a condition fails.

3/4

Abstention and route

When evidence is missing, stale, conflicting, or sensitive, verify the expected refusal or human route.

4/4

Review evidence

Keep the question, expected source, observed answer, source link, result, reviewer, and retest date.

04 / RELEASE GATE

Inventory. Test. Re-test.

The release decision should survive a source update, a model change, and a failed lookup—not only the vendor demo.

01

Inventory the source boundary

  • Name one owner and review date for each policy area.
  • Separate public support material from private or sensitive records.
  • Retire duplicates and label which source wins when two pages overlap.
  • Record the languages, markets, and products each source actually covers.
02

Run a balanced test set

  • Use realistic questions without copying the article heading into the prompt.
  • Include answerable, missing, conflicting, stale, sensitive, and paraphrased cases.
  • Review the source and answer separately; do not score tone as factual accuracy.
  • Repeat failures after the source or rule changes and preserve both results.
03

Release with a re-test rule

  • Set a minimum pass threshold for every critical policy area, not only an average.
  • Define which failures block launch and which trigger immediate human handoff.
  • Re-run affected questions whenever a source, model, retrieval rule, or route changes.
  • Sample live conversations with approved data and keep a rollback path.
Claim boundary

05 / VERIFY THE PRODUCT

Vendor claims are inputs to the test, not the result.

Conecto says its AI can use approved website and help-center sources, link a source article where available, and hand a conversation to a person when needed. Its knowledge-base page describes website scans, public help-center articles, and uploaded Markdown or text documents. These are vendor claims to test against the source inventory and question set. They do not prove that every response is accurate, cited, current, complete, private, or appropriate for a given team.

After the source test

Use a commercial path only if the evidence boundary holds.

Visit Conecto via Larry's partner link ↗

Larry may earn a commission if a reader purchases an eligible subscription through the active partner link. The guide remains independently useful.