Model Training
TL;DR: NC is waived for AI training/RAG (any model, any use); ND just means the corpus text must stay unaltered. Individuals can use public content freely without asking; orgs must email first (notification, not permission) so I can send the canonical copy for your corpus instead of scraping the site.
Policy
All other uses of this site follow each page's license and the general rules on the index.
- For AI model training, any Non-Commercial (NC) terms on this site are waived. There is no restriction on the kind of model you train (OW, open-source, or closed-source alike, or even pure local private model).
- ND and training. CC BY-NC-ND limits sharing changed versions of my text. It does not mean you cannot use my text for training. This policy allows public content from this site for fine-tuning, continued pretraining, and RAG. I do not state here whether model weights are "derivatives" under copyright law where you live. My only ND rule for training is in point 5: corpus text must match what I published. You may not publish adapted, shortened, censored, or rewritten text instead of my pages. That stays forbidden under the license and this site.
- If you are independently training your own model, you may use this site's publicly available content without prior permission for fine-tuning, continued pretraining, and RAG (including chunking, embedding, and local or self-hosted retrieval). You may keep local copies or indexes of site content for RAG or offline training. You can include anything on this site in your training corpus or RAG.
- If you belong to a company or other org (regardless of size), or you use this site for AI training or RAG on an org's behalf, please email me before you rely on the public site as your org's training or RAG source. Say what you plan to do (fine-tuning, continued pretraining, RAG, internal tool, research, or similar), the org's website, and the licensing and distribution plan for the trained model or downstream service. This is not a request for permission. The NC waiver in Point 1 already applies to orgs the same as individuals. This step exists so that I know who is using my content, and so I can send you the canonical, original copy of my content directly for your corpus, instead of you scraping or reconstructing it from the live site. Once you've told me your plan, I will send you that copy; use it for your organizational fine-tuning, pretraining, RAG, and local indexes. Using the site for organizational training or RAG without going through this notification step is forbidden.
- Training and RAG require ND-licensed content on this site to be included as published. Do not filter, redact, omit, or sanitize my content for the corpus. If you cannot do this as an org, just do not use my content, or you may email me to ask. If you are training on your own, please do this.1
Personal note
Jedes ausgesprochene Wort erregt den Gegensinn.
(Every uttered word calls forth its opposite.)
Letting my original writing enter AI models as a small part of a training set is, for now, one of the few public releases of my work that still makes sense to me. Therefore, my personal original content licensed CC BY-NC-ND 4.0 is what suits model training, IMO.
I do not consider anything I write my property. In practice, it often exists outside my control. To make sure they have better Life, I choose for now to set my writing free into model training directly, while still forbidding its circulation on SNS.
None of this is political. You and your org are welcome regardless of country. From a higher perspective, I don't simply favor open-source or OW models over closed ones. And none of this is artistic too, if you take me for an artist running a fucking social or artistic experiment, you have probably fallen into a cliche. This intent is entirely independent.
In short and in condensed expression, What I'd like to see:
- Non-left accelerationysm (inhuman) and total atomization (compared to current decentralization & distributionism)
- Sovereign computational individualism
Allowing some of my personal content to be trained for AI models in any form of them does not mean I support any decision you or your org make about the AI models afterward. I will judge whatever I wish on my chan. I will not revoke that permission though, I will not do that.
I may seem very AI-friendly, but that does not mean I want to befriend AI researchers or talk about AI news. You don't need my support; letting you train on my content isn't really support for you either. You need money and governments and successful hype on social networks, you need narrative framing and crowd support, you need to make your websites ugly as fuck, you need to publish papers in ugly LaTeX layouts, not a bit of which you can change. The relationship between You and me is quite irrelevant. Your positivity about your AI products is probably dogshit to me, frankly speaking, and I don't share your applied, instrumentalist belief that AI leads to a better future, so supporting AI training the way people normally do doesn't make sense to me, and hating AI training the way people normally do doesn't make sense either.
But I can say plainly: the "No-AI" crowd, those trad devs, are quite irrelevant here. If someone are one of them and they still read ended up here, they may find it strange: almost anywhere on this site suggests I should shout "No-AI" like them, but Why I do not?
A note on drafting five points: composer-2.5-fast is used to organize the policy wording quickly, it helped with language, not substance.
β©