I read a lot of investing newsletters, and it’s truly remarkable how similar they are. In one I just got, it was extolling the virtues of AI bots for web search, and the essential lesson was to not block (so-called) AI bots, because that bot traffic yields useful content. Setting aside many facets of that, wherein the LLM bot is simply harvesting content based on mathematic representations of words and is therefore only evaluating information quality on extremely narrow definitions. Any reader knows intuitively that this way of measuring and evaluating is extremely dubious, as it is unlikely to surface fundamentally important information. The fundamental assumption of LLM models used for language is that words in books are connected similarly. Essentially, if you’ve read 99 books, you know how the 100th goes. In part, I can get on board with this assumption — if you’ve read 99 books, you probably know all the words and letters used in the 100th — sure. Moreover, you likely generally know the story arc in broad terms. No objection here. However, the details beyond those broad strokes can vary quite significantly.
I think LLM bots can be quite useful for summarizing some content. For instance, let’s say I have 500 search results summarizing photosynthesis. Cool — give me the LLM-derived summary. I don’t necessarily think it will be fundamentally correct for all cases, as I suspect, the LLM is going to fail to convey important nuance along the way. I can live with that, in large part, because I am aware of what I’m getting. The levels of nuance and detail conveyed by LLMs are wholly insufficient for representing a whole, however, because that is simply not how they work. Like, when you say, “A person,” you are referring to a singular entity. That is a single unit. You are not referring to eight eighths of a person. However, in purely mathematical terms, 8/8 equals 1. But as a practical matter, a pile of pieces of a person would be _very_ different than one single person as a unit.
Back to the initial comment about wanting your website to be ingested for use in AI summaries. I trust you can see where this is going. I personally do not want my content to be used in “AI” training data, because the manner in which its used may be antithetical to my goals for the material. This may be a minor detail for a shit-for-brains stock blowhard, but I actually think intention is one of the most important things. If there is a key objectionable aspect of LLMs as AI, it’s that they split words into tokens to where the veracity of any particular claim is eliminated, leaving only residue in its place.