Raw Thoughts

Fair Use

September 5, 2026fair usecopyrighttraining datadata ownershipAGIopen sourcecreative commons

What I do not understand is how difficult it is for artificial intelligence companies to understand what these two basic words mean. This one phrase, a combination of two simple English words, defines when and how one can use a type of data. Extracting a passage of text from a web page, using it for training, to push an AGI narrative while claiming other nations are in an arms race to take over the world or something or the other. A.I. is not some nuclear technology, it is a means to an end, an end that Silicon Valley has been optimizing towards since DARPA connected the world. That knowledge should be accessible. But we forgot what it took to protect the creative commons in the process. Without that balance, the idea of an "internet" is not possible.

Figure

If a person puts in time to stitch together words onto a document that provides knowledge to a reader, the reader extracts that knowledge and is then able to contribute to society more effectively. Remixing that knowledge into a fiscal return of some kind, whether it be a new UX feature of their application or an addendum to their very own blog post. The idea of open source does not just apply to code, it applies to how the world shares intelligence with each other and evolves together. "Decentralizing intelligence centralizes growth for the whole."

If we take this knowledge, without consent, to train siloed and compressed machines of thought that are not able to accurately share the people involved in producing these paths of learning, the essence, the balance, the idea of the internet dies. I believe we can reverse "dead internet theory," but it starts with organizing and building solutions since courts are taking too long to adapt and/or there is not enough awareness of the reality of the situation and what A.I. competition truly looks like to mitigate the usual arguments to defend unlawful training data collection.

Figure

I do not believe "AGI" to be a goal to reach that matters that much. Humans evolve and A.I. will just prove that humans will scale and A.I. itself will always be a race to gather more and more knowledge, even from secretive sources, in order to always scale with how humans adapt to the current systems to remain relevant. It's a never-ending battle and I will always bet on the human to outmaneuver a machine. Thus, we are looking at the intelligence problem incorrectly. Data should be clearly owned by the individual, private and personal, and this data should be leased upon request into systems that aggregate it.

Large language models, at their core, are aggregates of sources within a corpus of data. This aggregate can be cited to truth, and that truth should be exposed and transparent rather than jumbled together on an S3 waiting to be pulled for another epoch.

Let's build that framework, the architecture, the proper storage units to do just that. The only way a country should be allowed to wave its flag in an arms race is to make sure the roads are paved along the way. Otherwise, the credibility of the very invention we pour billions into falters at the expense of the very people that helped build this technology in the first place. The creatives, the intellectuals, the commons.

Figure