Project architecture · Applied AI

Documentation &
evaluation assistants

Connecting product documentation, scientific evidence, and evaluation results to an assistant that can explain where its answers come from.

Three places to ask

A biosensor analysis application needs different evidence for product instructions, scientific explanations, and questions about a particular result. The assistants share retrieval and presentation while keeping those contexts distinct.

Product documentation

Answers questions using generated documentation and curated scientific references, with linked sources and a visible evidence-coverage label.

Example questionHow do I configure a single-cycle experiment?

Evaluation setup

The Kinetics and Affinity wizard adds the current goal, step, and selection to the initial prompt. It reuses the documentation endpoint and drawer.

Example questionWhich data do I need for this evaluation?

Completed results

A separate service adds deterministic findings and tools for stored fits, responses, and rejection evidence within the authorized evaluation.

Example questionWhy was this fit excluded from the result?

Follow a question

The assistant runs inside the existing FastAPI Hub. The browser sends a question, the Hub prepares the evidence, and the configured LLM provider generates the answer.

Select an entry point to inspect its request path.

Browser · React / TypeScript

Documentation drawer

A question and session identifier enter the shared assistant stream hook.

POST /documentation/chat/stream
FastAPI Hub · application backendAuthenticated user · conversation scope
  1. 01

    Route the question

    Deterministic rules select answer mode, knowledge intent, and technology context.

  2. 02

    Retrieve eligible evidence

    Metadata filters run before weighted lexical ranking. Up to six chunks enter actionable and background pools.

  3. 03

    Build prompt & stream

    Evidence, answer policy, and bounded history go to the shared OpenAI-compatible LLM client.

Local knowledge corpus → retrieval

Generated documentation JSON plus curated external references, indexed when loaded.

Optional catalog branch → prompt

A separate AI routing decision can authorize a bounded, cached first-party catalog read.

NDJSON response

token sends text increments. done carries the full answer, session, route, coverage, and resolved sources. error reports a failure.

Shared answer renderer

Markdown and equations, source links, grounding and intent labels, rendered inside the reusable assistant drawer.

The ordinary documentation path uses local retrieval. Catalog routes can add live first-party data or return a shop link or catalog-unavailable notice.

The browser consumes one stream contract

Both services emit newline-delimited JSON. This abbreviated illustration shows the event shapes, without a generated answer or a measured result.

{"type":"token","content":"Text increment"}
{"type":"done","session_id":"…","ai_message":"…","answer_mode":"product","knowledge_intent":"product_operation","source_coverage":"low","sources":[…]}
{"type":"error","detail":"User-facing error message"}

Evidence has a job

A relevant scientific reference may explain a mechanism without supporting a product procedure. Retrieval policy makes that distinction before the model sees an excerpt.

Actionable evidence

Support a product recommendation

Product documentation and code facts eligible under the policy support product-specific actions. Troubleshooting prioritizes evidence for the product's focal-molography technology.

Background evidence

Explain the science

Curated sources carry technology, claim-type, and transferability metadata. Analogy-only material can explain a concept, but the prompt prohibits turning it into a product instruction.

Representative local-context allocations. The complete policy also considers the named technology and source applicability.
Knowledge intentActionable poolBackground pool
Product operation / troubleshootingUp to 4 product chunksUp to 2 validated or shared focal-molography chunks
Product surface chemistry / experiment designUp to 3 product chunksUp to 3 compatible chemistry, mechanism, or design chunks
Technology comparisonNo actionable poolUp to 6 chunks, with representation reserved for named technologies

Source coverage is an evidence-count bucket. For product-grounded routes, scientific background does not increase the product-grounding count. This label describes retrieved support; it is not a calibrated probability that an answer is correct.

Routing separates answer mode, intent, and technology

Answer modes are product, science, hybrid, and low_confidence. Knowledge intents include product operation, troubleshooting, experiment design, target-specific experiment design, surface chemistry, science, comparison, and unknown. Technology context constrains which evidence is eligible.

Citation resolution stays within the supplied evidence

The backend resolves source markers against the excerpts and catalog records available for that answer. If no valid markers resolve, it returns the retrieved corpus sources as a fallback. Catalog sources appear only when cited. Resolving a marker verifies source membership, not whether the source supports every claim.

Completed evaluations use bounded, read-only tools

The study endpoint authorizes the evaluation and loads bounded, typed snapshots before streaming. Database connections close before the model's tool loop. Hub-executed tools list entities, retrieve fit results, retrieve responses, explain rejections, and run an explicitly parameterized kinetic-design what-if. The tools cannot mutate experiments or select another study.

Rounds, calls, argument size, result size, and time are bounded. After budget exhaustion, a final model turn runs with tools disabled. Rejected fits provide exclusion evidence and cannot support binding conclusions. Stored plot metadata does not mean the assistant has read plot pixels.

Conversation history commits only after an answer completes

Sessions are keyed by user, assistant, resource, and session identifier. One turn runs at a time. The question and answer commit together after successful generation; interrupted turns leave the committed history unchanged. History and active-session counts are bounded. Sessions expire after one hour of inactivity and live only in the Hub process.

Test the boundaries

The regression evaluator shares the production question-preparation and prompt-building functions. It can inspect retrieval without a provider call, then use the same prompt in optional live runs.

Corpus freshness

The build extracts documentation into a committed JSON corpus. CI checks source and generator hashes so documentation changes cannot silently leave retrieval on an old cache.

Offline regression gate

Deterministic cases check route, intent, technology, retrieval hits, and forbidden evidence. Gating cases run through the API test suite; diagnostic cases preserve known gaps.

Optional live evaluation

Provider runs record citations, time to first token, total duration, usage, and reported cost. Missing provider usage stays missing. The evaluator does not score semantic answer quality.

Inspectable retrieval

Local lexical ranking and explicit filters make source selection reproducible. Coverage still depends on curated metadata, vocabulary matching, and the available corpus. Multilingual and routing gaps remain visible in diagnostics.

Shared state, explicit limits

A common stream and UI support several workflows. Process-local sessions keep storage simple, but do not survive a Hub restart or synchronize across workers. Prompt instructions guide generation; they do not prove factual correctness.