← Ask the Declaration
Under the Hood

How This Works

A complete retrieval pipeline with no servers, no API keys, and no tracking. Your question never leaves your browser. Here is the full engineering story.

§1 The Pipeline

Your Browser — Privacy Boundary

Your question never crosses this line

Step 1 — You ask
What protects free speech?
Can the government search my house?
What abolished slavery?
Step 2 — Browser converts to 384 numbers (all-MiniLM-L6-v2)
[ 0.23, −0.41, 0.87, 0.12, −0.53, 0.71, 0.09, −0.34 … 376 more ]
[ −0.18, 0.62, −0.33, 0.89, 0.14, −0.47, 0.55, 0.21 … 376 more ]
[ 0.44, 0.31, −0.67, 0.19, 0.82, −0.25, 0.38, −0.11 … 376 more ]
Step 3 — Dot product vs 108 passages (< 1ms)
Amendment I
91%
Amendment XIV
44%
Article I, §8
28%
Amendment IV
88%
Amendment III
51%
Amendment V
39%
Amendment XIII
95%
Declaration §2
34%
Amendment XIV
29%
Step 4 — Founders answer in their own words
Amendment I · 1791
“Congress shall make no law … abridging the freedom of speech, or of the press.”
Amendment IV · 1791
“The right of the people to be secure in their persons, houses, papers, and effects, against unreasonable searches and seizures, shall not be violated.”
Amendment XIII · 1865
“Neither slavery nor involuntary servitude … shall exist within the United States.”
No outbound API calls
OpenAI API
Vector DB
Your Server

§2 Three Signals, Not One

The diagram above is the first version of this site: one model, dense vectors, a dot product. It shipped, and it worked. But a single bi-encoder only retrieves well when the user phrases a question in the documents' own vocabulary, and real questions rarely do. So retrieval here now runs three signals that fail in different ways, so one covers for another.

Dense (all-MiniLM-L6-v2) catches meaning and sets the default order. BM25, a classic keyword score, catches the exact wording the small model misses; it does not reorder results, it only widens the candidate pool so an exact-match passage reaches the next stage. A cross-encoder reranker (ms-marco-MiniLM-L-6-v2) then decides the final order.

The difference between the two kinds of model is the whole point. A bi-encoder turns the query and each passage into a vector separately, ahead of time, and compares them with a dot product. That is why search is instant: the 108 passage vectors are precomputed, so a query costs one embedding plus 108 cheap multiplications. The price is precision, because a passage's vector has to summarize it for every possible question at once. A cross-encoder does the opposite: it reads the query and one passage together, so every word of the question can attend to every word of the passage, and returns a single relevance score. Far more accurate, but it cannot be precomputed. The standard resolution, used here, is retrieve then rerank: the bi-encoder and BM25 cheaply narrow 108 passages to about 20, and the cross-encoder scores just those 20 and reorders them. It loads in the background after the first search and degrades gracefully — until it is ready, results fall back to the dense order, so search never waits on it.

§3 The Numbers

None of that is worth claiming without measuring it. eval/queries.json is 45 labeled questions, each tagged with the passage a good answer should surface. node eval/run-eval.js replicates the browser's retrieval exactly and reports recall@5 (did the right passage land in the top five) and mean reciprocal rank (how high it landed) for each layer, so every stage's contribution is a number, not an adjective.

Configurationrecall@5MRR
Dense, text-only (v1)93.3%0.904
+ enriched embeddings95.6%0.897
+ BM25 candidate pool95.6%0.897
+ cross-encoder rerank97.8%0.952
Cross-document (multi-hop) slice100.0%1.000

Two honest readings. First, this corpus is easy: 108 canonical passages with distinctive vocabulary, so dense retrieval already hits 93%. Second, the reranker earns its keep on precision more than recall — recall@5 rises two points, but MRR jumps from 0.897 to 0.952 because the cross-encoder pulls the right passage higher. It recovered "which branch can declare war," which the bi-encoder had buried below the top five.

Is a reranker even justified at 108 passages? For latency, arguably not. It is here to demonstrate the retrieve-then-rerank pattern and to measure whether it helps, and the measurement said yes on ranking quality. Knowing that difference, and being willing to say it out loud, is the engineering. I also tested a larger embedder (bge-base, 768-dim, ~110 MB) against the small one: it had lower dense recall (91.1%) and only tied after reranking, so the site keeps the 5×-smaller model. The bigger model did not earn its bytes, and the eval is how I know.

§4 The Thread: Reading Across the Documents

The American founding is not one document but a 250-year conversation between several: a grievance in the 1776 Declaration becomes a power in the 1787 Constitution becomes a right in the 1791 Bill of Rights. The complaint about quartering troops (Grievance 14) becomes the Third Amendment. "Taxes without our consent" (Grievance 17) becomes Congress's taxing power, and later the Sixteenth Amendment. So each result here carries a lineage strip that traces its principle across the documents and across the years — some threads hand-curated for accuracy, the rest computed as each passage's nearest neighbors in a different source document, found by the same vectors that power search.

This is also the sharpest way this site differs from its sister project, Ask the Constitution of India. That corpus is a single document, so it optimizes for the opposite end of the same curve: minimize the download, stay CPU-only, because its readers may be on a weak connection asking privacy-sensitive questions. This one assumes broadband and an intertextual corpus, so it spends that budget on cross-document reasoning and a reranker. Same architecture, same $0 hosting, same nothing-leaves-your-browser — opposite operating points, each chosen for its audience.

§5 Why 108 Passages, Not 500-Token Windows

The default way to chunk documents for retrieval is mechanical: slice the text into fixed windows of a few hundred tokens, with some overlap so nothing falls between the cracks. It works, and it is also why so many retrieval systems cite "chunk 47" and hand you a passage that starts mid-sentence.

These documents deserve better, and they make it easy. The founding documents carry their own semantic boundaries: each grievance in the Declaration is one complete complaint against the king. Each section of the Constitution is one power or one limit. Each amendment is one right. Chunk along those seams and every unit of retrieval is a complete thought with a name a human already understands.

That is why a result here says Article I, Section 8 or Grievance 17 instead of a byte offset. A journalist can quote it. A teacher can assign it. A lawyer would recognize the citation. The retrieval system speaks the same language as the document it searches.

The transferable rule: before reaching for a token splitter, ask what the document's own atomic unit is. Contracts have clauses. API docs have endpoints. Runbooks have steps. Structure-aware chunking costs one afternoon of parsing and pays out every single query.

§6 The Economics of Zero

First, precision: this is not a "local LLM." It is a local embedding model, about 25 MB of quantized ONNX, running in your browser tab. There is no text generation anywhere in the system, which is why it cannot hallucinate and why the numbers below look the way they do.

Cost lineTypical hosted RAGThis site
Query embeddingAPI call, metered per token$0 — computed in your browser
Vector searchHosted vector DB, monthly fee$0 — a dot product over a static file
Answer generationLLM call, cents per query$0 — curated text, written once
Keys & rate limitsKeys to protect, quotas to hitNone exist
Cost if it goes viralScales with every visitorFlat. CDN serves static files

The only real cost is the one-time model download on a visitor's first search, cached by the browser afterward. That is the honest tradeoff, and for a public demo it is the right one: a demo that costs money per query dies the day it goes viral. This one cannot, because its marginal cost per query is zero at ten visitors and zero at ten million.

§7 When to Use This Pattern, and When Not To

Use it for small, public, read-heavy corpora: documentation sites, legal and policy texts, product manuals, FAQ knowledge bases. Anywhere under a few thousand chunks where the content is not secret and the questions repeat.

Do not use it for private data (the whole index ships to every visitor), for large corpora (the index download outgrows its welcome), or when users need generated prose rather than retrieved passages. Engineering maturity is matching the architecture to the problem, not to the hype cycle.

§8 What Is Curated vs. Computed

Honesty about the seams: by default the AI here does retrieval and question-matching only. The plain-words explainers under each passage, and the short answers for common questions, were written by a person at build time and shipped as static text. The founders' words are quoted exactly from the public domain Project Gutenberg editions. Generation exists but is optional, opt-in, and clearly labeled (next section). Each layer is labeled in the interface, so you always know who is talking: 1776, or 2026.

§9 Optional: Generation That Stays on Your Device

Retrieval always leads, and the quoted passages are the answer. But on a machine with WebGPU, you can opt in to a fuller written answer, generated on your device by a small open-weights instruct model (Llama-3.2-1B via WebLLM) from the retrieved passages only, with a system prompt that forbids adding facts and tells it to say when the documents are silent. It downloads once, runs entirely in the browser, and is labeled as generated and unverified. If your browser has no WebGPU, or you never opt in, the site behaves exactly as it always has.

How grounded is a generated answer? node eval/groundedness.js measures the share of an answer's content words that actually appear in the retrieved passages. The small fallback synthesizer scores about 61% on that proxy, and its worst cases are honest refusals the metric unfairly penalizes. That number is the reason the WebGPU option exists: a real instruct model, run on-device, is the upgrade for people whose hardware can afford it, at the same $0 and the same nothing-leaves-your-browser.

Two sister sites, one more time. This is the capability-driven half of the same tradeoff: where the India site minimizes for weak devices, this one spends a WebGPU-capable device's budget on generation quality. Retrieval is the floor both sites stand on; generation is a labeled convenience layered on top, never a source of new claims.
✦ ✦ ✦

Retrieval that citizens can quote, at a cost that virality cannot kill. The architecture is the product decision.

The entire build is open source: github.com/swapniltamse/ask-the-declaration