Neural web search, end to end

Inside an Exa Search

Exa runs two systems that never meet in real time. An offline pipeline turns the web into compact vectors and files them into clusters. An online path turns your query into a vector and scores only the nearest clusters. Press play, or pick a step. Click any box in the diagram to jump to it.

Offline · Step 1 of 12

Crawl the web

Exa’s crawler keeps discovering new URLs and fetches them across a distributed network of machines and IP addresses.

Source: Exa blog, Mar 2025

One page, three sizes

The same page at each stage of compression. Squares are drawn with area to scale.

8 KB4,096 × fp16 · full embedding
512 B256 × fp16 · Matryoshka cut
32 B256 bits · binary index

Per billion pages: 8 TB, then 512 GB, then 32 GB. Exa says the full index uses less memory than a gaming PC; the full vectors stay on disk for reranking.

In your setup

How Claude Code on this machine reaches Exa.

  • TransportHosted HTTP MCP server at mcp.exa.ai/mcp, registered in ~/.claude.json with your API key.
  • Toolsweb_search_exa (this diagram), web_fetch_exa (read a URL from Exa’s copy), agent_run.
  • RoutingE:\CLAUDE.md §3 already says entity and “things like X” questions go to Exa. Plain lookups go to Firecrawl, and single URLs go to Crawl4AI.

Sources

  1. How we built a web-scale vector database Exa, 17 Dec 2024. Steps 3–5 and 7–12: Matryoshka, binarization, lookup tables, clustering, filters, reranking, speed claims.
  2. How we’re building the next generation of search Exa, 11 Mar 2025. Steps 1–2: crawler, parser, S3, 144-GPU cluster.
  3. Beating Google at search with neural PageRank and $5M of H200s Latent Space podcast with Exa CEO Will Bryk, 10 Jan 2025. Step 2: how the model was trained by link prediction.
  4. I reverse-engineered Exa.ai infrastructure cost with napkin math K. Shivendu, 16 Sep 2025. Outside estimate of the rerank stages; Exa has not confirmed it.