The Index and the Vector (newsletter.dancohen.org)

🤖 AI Summary
Dan Cohen argues that large language models’ embedding vectors can bridge the gap between novices’ vague, lay descriptions and the precise metadata used in libraries, archives, and museums. Unlike traditional indexes that require exact keywords (artist names, movements, periods), vectors represent words and artifacts as points in a multidimensional space where proximity equals semantic similarity; aggregated cues (e.g., “long-haired,” “woodsy,” “mythical”) can combine to pull the correct concept (like “Pre‑Raphaelite”) toward the top. This makes conversational, iterative queries far better at mapping ambiguous human expression to curated collections, helping users discover and learn without prior specialist vocabulary. Cohen also reports a practical test: he’s added museum MCPs and article databases to a custom Claude rig with connectors to the Met and Art Institute of Chicago, allowing him to disable web retrieval (RAG) and rely solely on library- and museum-augmented responses. The setup returns curated articles and representative digitized works (e.g., for “Cubism”), demonstrating how domain-specific augmentation plus embeddings can enable robust, interpretation-friendly discovery. For the AI/ML community, this reinforces the value of embedding-based semantic search, domain augmentation, and conversational refinement as scalable ways to make curated collections accessible and pedagogically useful.
Loading comments...
loading comments...