How NarrateAI Keeps Executive AI Answers Accurate and Fast on Amazon Bedrock

How NarrateAI Keeps Executive AI Answers Accurate and Fast on Amazon Bedrock

According to Artificial Intelligence (AWS blog), NarrateAI reaches roughly 99 % numerical accuracy while streaming answers to executives in real time. The system does this by stitching together five engineering tricks, each covering a different failure mode that shows up when large language models (LLMs) are used in live business reviews.

Adaptive Pipeline Orchestration: One Size Doesn’t Fit All Queries

Executive questions range from “What’s my team’s quarterly attainment?” (a few document sections) to “Give me a full regional performance analysis” (hundreds of sections). A single processing strategy either truncates the context when it overflows the model’s token limit (about 200 K tokens) or wastes resources by always invoking multiple LLM calls. NarrateAI solves this by measuring the total retrieved text size ‑ |D| ‑ and routing the request.

Path Approx. % of queries Typical latency LLM invocations per query
Fast (single‑pass) ~90 % < 25 s (TTFT ≈ few seconds) 1
Normal (multi‑batch) ~10 % 50‑75 s 1 + N (N≈4 batches)

The first phase, Mode‑Aware Consolidation, greedily packs sections into the largest possible chunk that stays under the token limit. If the chunk fits, the query takes the fast path; otherwise it is split into N batches that respect document boundaries. The second phase runs either a single LLM call (fast) or N parallel calls (normal). The third phase merges the parallel outputs only for the normal path. In practice this reduces the average number of LLM invocations by about 72 % compared with a naïve always‑multi‑pass approach.

Cross‑Account Multi‑Model Failover: Turning Quotas into Capacity

Bedrock assigns a separate request quota to every (model, AWS account) pair. A single model in a single account can become a bottleneck during peak review periods, leading to throttling errors that force users back to spreadsheets. NarrateAI treats each pair as an independent capacity bucket. By configuring, for example, three models across three accounts, the system gains nine parallel quota spaces.

When a request arrives, the code tries the highest‑ranked model in the first account. If the call is throttled, a ThrottlingDetector immediately swaps in the same model in the next account, and only after exhausting all accounts does it fall back to the next‑best model. Randomly shuffling the account list before each attempt spreads load evenly without a central scheduler. The AWS Security Token Service (STS) can assume a new role in 100‑200 ms, far faster than the multi‑second back‑off loops traditionally used.

Over six months with more than 4 000 users, this N × M quota exploration kept throttling virtually invisible, delivering the same latency as a single‑quota system under normal load while preserving model quality.

Real‑Time Streaming Evaluation and Composite Evaluation Framework

Even when a model stays up, it can emit hallucinated numbers or vague language. NarrateAI plugs a streaming evaluator into the token generation pipeline. As each paragraph is produced, the evaluator runs three independent checks in parallel:

  1. Exact‑match filter – cheap string comparison against known metrics.
  2. Semantic verifier – a lighter‑weight model that scores factual consistency.
  3. Style guard – ensures professional tone and removes subjective phrasing.

If any check fails, the paragraph is either corrected on the fly or flagged for the consolidation step. Because the checks run concurrently with generation, they add only a few hundred milliseconds to overall latency.

Data Accuracy Verification: A Two‑Stage Guard Against Numeric Hallucinations

Numeric hallucination—fabricating a figure that looks plausible—is the most visible failure in executive dashboards. NarrateAI’s cascade starts with the cheap exact‑match filter. Only when a number is not found in the source documents does the system invoke the semantic verifier, which compares the generated figure’s meaning to the original data using a similarity model. This staged approach keeps costs low while catching the majority of errors before they reach the user.

The Hidden Trade‑Off: Managing Quota Complexity vs Simplicity

All five techniques work together, but they also introduce operational overhead. Running multiple accounts and models means you must maintain separate IAM roles, monitor quota consumption across nine cells, and keep model versions in sync. The benefit—near‑zero throttling and graceful degradation—is clear, yet the complexity can outweigh the gains for smaller teams that only need occasional executive queries.

In practice the trade‑off looks like this:

  • Small teams (≤ 100 daily queries) – a single high‑quota model often suffices; adding accounts may be overkill.
  • Mid‑size teams (≈ 1 000 daily queries) – cross‑account failover yields measurable latency improvements, especially during quarterly close periods.
  • Enterprise deployments (> 5 000 daily queries) – the full N × M setup becomes a cost‑effective way to guarantee SLA (service‑level agreement) compliance without over‑provisioning dedicated compute.

What you should watch next is quota drift. As usage patterns shift, the empirical threshold θ for fast‑path routing may need recalibration, and the model‑ranking matrix may require re‑ordering if a newer model becomes cheaper or more accurate.

How to Try This Today

  1. Measure your query size distribution. Dump a week of real requests, sum the retrieved token count per request, and plot the histogram. Identify a token‑count percentile that captures about 85‑90 % of queries; use that as an initial θ.
  2. Implement a three‑phase router. Start with a simple script that concatenates sections when total tokens ≤ θ and falls back to batch calls otherwise. Use Bedrock’s invoke_model API for the single call and parallel ThreadPoolExecutor for the batch calls.
  3. Add a cheap exact‑match filter. Keep a cache of known metric identifiers (e.g., “Q2 revenue”) and reject any generated number that isn’t in the cache before returning the answer.
  4. If you have multiple AWS accounts, test a failover stub. Wrap the Bedrock call in a try/except block that catches ThrottlingException and immediately retries with a different role ARN.

Even a lightweight version of these steps will give you a feel for the latency‑accuracy balance NarrateAI achieved, and you can iterate from there.

Sources

Read next

We count page views without cookies — no identifier, nothing stored on your device. Accept to allow cookies for analytics.