Crazy Wisdom podcast show image

Crazy Wisdom

Stewart Alsop III | AI, Consciousness & Technology

Podcast

Episodes

Listen, download, subscribe

Episode #569: What the Romans Wrote Over, and What AI Is Erasing Now

In this episode of Crazy Wisdom, Stewart Alsop sits down with Charlie D. Becker, a second-generation bookseller whose family runs Houston's largest used and rare bookstore, to unpack the viral tweet that had people convinced AI companies were secretly buying up used books to train their models. Charlie walks through what he actually found in his own warehouse orders, why the more likely explanation is old-fashioned reseller arbitrage and FBA "bookjacking" rather than AI training data, and how that story connects to the real, well-documented case of Anthropic scanning and destroying physical books for legal reasons. From there the conversation moves into the difference between rare and valuable books, the discoverability problem in the used book market, historical parallels like palimpsests and lost texts, his own AI tool for used bookstores, and broader questions about wealthy patronage funding independent research and passion projects. For more, check out Charlie's personal site at charliedbecker.com and his Substack at charliebecker.substack.com.Timestamps00:00 Stewart introduces Charlie D. Becker, a second generation bookseller building an AI tool for used bookstores and discusses Anthropic's controversial book acquisition practices05:00 Charlie explains how Anthropic legally destroyed physical books by slicing spines to scan them, avoiding copyright violations while building training datasets for AI models10:00 Charlie describes receiving bizarre bulk book orders through non-Amazon platforms, initially suspecting AI companies but discovering evidence of sophisticated book arbitrage operations instead15:00 Discussion of the banal explanation for mysterious orders: algorithmic resellers buying cheap books from obscure platforms to flip on Amazon through FBA warehouses20:00 Stewart and Charlie explore historical parallels between the printing press era and today's digital transition, discussing the loss of archival records and palimpsests25:00 Charlie emphasizes the discoverability problem for rare obscure books and how profit-driven algorithms prevent people from finding books unless they know exactly what to search for30:00 Detailed explanation of Charlie's AI cataloging tool that creates bibliographic profiles from photos of pre-1970 books lacking ISBNs, addressing the hard problem of metadata creation35:00 Discussion of the technical challenges solved: archival-safe removable stickers, RFID systems, and creating canonical records that become definitive sources for rare books40:00 Charlie describes building copy-level databases beyond edition-level records, creating VIN numbers for books, and designing knowledge graphs linking works to editions and translations45:00 Vision for navigable work-edition hierarchies allowing researchers to explore translation genealogies and linguistic families, solving problems Amazon has no incentive to address50:00 Stewart raises the possibility of returning to gentleman's science and aristocratic private libraries in an age of AI abundance and accessible three-d printing technology55:00 Charlie reflects on supporting idiosyncratic passion projects regardless of profit, his fellowship from Jim O'Shaughnessy, and navigating economic inequality while promoting eccentric research pursuitsKey Insights1. Charlie D. Becker is a second generation bookseller whose family runs Houston's largest used and rare bookstore, and he is currently building an AI tool specifically designed for used bookstores. He became widely known after a tweet he wrote about AI companies potentially purchasing books went viral with approximately one and a half million impressions, though he emphasizes the importance of being careful about distinguishing between what he has directly observed, what is on public record, and what is his intuition or speculation about these events.2. The controversy around Anthropic and book destruction centers on how AI companies acquire training data from physical books. Court documents revealed that Anthropic acquired physical books in mass quantities to scan them, and they were industrially slicing the spines off to make scanning faster and easier. The legal justification for this practice was that because they destroyed the original physical book after scanning it, they were not violating copyright law since no duplicate copy existed alongside the original. A judge ruled this was technically legal, even though it appeared problematic to many observers, because the destruction of the original meant they were not running afoul of copyright provisions about making copies for distribution.3. Becker received unusual bulk orders for obscure books through a non-Amazon platform in late April, which led him to investigate whether AI companies were responsible for these purchases. However, after analyzing the pattern of purchases and where the books were being shipped, he concluded that a more mundane explanation was likely at work: sophisticated book arbitrage operations. These operations identify books selling cheaply on one platform that could be listed for higher prices on Amazon through Fulfillment by Amazon warehouses, and books that do not sell eventually get recycled or liquidated anyway, meaning rare books are being destroyed through normal commercial operations regardless of whether AI companies are involved.4. The main technical challenge Becker is solving with his AI tool relates to books printed before 1970, which lack ISBNs or International Standard Book Numbers. Modern book cataloging systems are built around ISBNs, which makes it extremely difficult and time-consuming to catalog older books for online sale since there is no automated way to populate bibliographic data for pre-1970 books. His tool uses computer vision and AI to analyze photographs of book covers, title pages, and copyright pages to automatically generate rich bibliographic metadata, and he has partnered with a PhD AI computer vision specialist to develop this technology over the past year.5. A surprising discovery during the development of this cataloging tool was that the existing data for many older books is extremely poor, inconsistent, or completely absent from major databases. For approximately a quarter of the books they process, the only existing records might be an incomplete eBay listing from years ago or a sparse entry in WorldCat, the interlibrary database. This means that rather than simply aggregating existing data, they are actually creating canonical records for many books that will become the authoritative source that others reference, essentially building new infrastructure for book metadata rather than just accessing what already exists.6. Becker advocates strongly for the preservation of obscure and seemingly unimportant books because while they may not have obvious value today, future researchers, tinkerers, or engineers might need them to solve problems we have not yet encountered. He compares this to historical palimpsests where important ancient texts were accidentally preserved when medieval scribes wrote over them, noting that the internet functions more like a palimpsest than an archive since we constantly overwrite and lose information rather than truly preserving it. The current system for deciding which books get preserved or destroyed is essentially random and driven purely by short-term profit motives rather than any thoughtful consideration of potential future value or historical significance.7. The long-term vision for the project extends beyond simple cataloging to creating a comprehensive knowledge graph that distinguishes between works, editions, and individual copies of books in ways that current commercial platforms do not adequately address. Unlike Goodreads which treats all editions of a book as a single work, or platforms like eBay that only show individual edition listings, Becker envisions a system where users can navigate between different organizational levels and explore the genealogy of works across translations, editions, and languages. The project also aims to create copy-level records similar to what libraries maintain, which would track provenance and availability of specific individual copies rather than just edition-level information, something no commercial platform currently does at scale.

Crazy Wisdom RSS Feed


Share: TwitterFacebook

Powered by Plink Plink icon plinkhq.com