Wikipedia:Village pump (idea lab)/Archive 77

From Wikipedia, the free encyclopedia

Better solutions for archiving web pages

Given the stance of Wikipedia on the matter of copyright, it is not possible for Wikipedia itself to host and manage a publicly viewable archive of webpages being cited. The situation with archive.today should be a wake-up call to better plan out how to deal with this matter. However, Wikipedia can develop, or at least endorse, a decentralized/distributed alternative archive: where the editors can use the developed software to store the DOM tree and/or screenshot of the webpage on IPFS or equivalent content-addressed cache, where we can use cryptographic hash to verify the authenticity of the cited material. This should also reduce dependence on Internet Archive moving forwards. Some volunteers can use copyright-havens (an allusion to tax-havens) to host a gateway to this decentralized archive, avoiding DMCA takedowns and such, in the manner of Sci-Hub, Libgen etc. (It is also of interest how the courts will decide on scraping of the internet to obtain training data for LLMs.)

What are the current state of the art in this direction? What do we need to do if we choose to go in this direction? What are some other viable directions to consider in this regard?

Smlckz (talk) 11:38, 12 February 2026 (UTC)

"What are some other viable directions to consider in this regard?" - I think SingleFile is worth having a look at. sapphaline (talk) 12:12, 12 February 2026 (UTC)
Sounds like a fun research project for someone who has the time for that. Would definitely encourage whoever wants to work on that. —TheDJ (talkcontribs) 14:54, 12 February 2026 (UTC)
I don't love the idea of Wikipedians setting up a massive copyright infringement site. Archive websites like .today are based in countries with lackluster copyright laws and use a variety of illegal methods for capturing copyrighted content. IA can get away with it since they're a non-profit which responds to takedown requests and respect the wishes of sites blocking it. The whole point of the Wikimedia movement is sharing free content, not coming up with ways to get around copyright law.
I'm not opposed to someone making a site like this, I don't really care. I just don't the idea of us developing or "endorsing" a project like this. IsCat (talk) 15:29, 12 February 2026 (UTC)
There's nothing wrong with copyright infringement as long as it's not hosted on WMF's servers. sapphaline (talk) 16:26, 12 February 2026 (UTC)
I mean, copyright infringement is illegal. I don't think it's smart to just ignore that.
Apart from that, WP:COPYVIOEL prohibits the addition of links that are knowingly in violation of copyright. To quote that page: If there is reason to believe that a website has a copy of a work in violation of its copyright, do not link to it. Linking to a page that illegally distributes someone else's work casts a bad light on Wikipedia and its editors. It's difficult to argue that the proposal here wouldn't fall under that. IA is acceptable because we can assume fair use, the same isn't true for this proposal (or even .today). IsCat (talk) 16:31, 12 February 2026 (UTC)
"copyright infringement is illegal" - "illegal" doesn't mean "unethical" or even "wrong". "IA is acceptable because we can assume fair use" - we can't. sapphaline (talk) 16:35, 12 February 2026 (UTC)
I never said copyright infringement is unethical. We live in a world full of laws and we have to follow them. IA themselves says that web archiving, at least the way they do it, is "broadly considered" to be fair use, see here. I don't think that's been challenged in court directly as IA has done takedowns and allows sites to exclude themselves. Perhaps we should keep it that way. IsCat (talk) 16:46, 12 February 2026 (UTC)
We live in a world full of laws and we have to follow them. If LLM scraper bots can get away with taking copyrighted materials as training data, using it for commercial purposes, does it make us a fool of ourselves for still following the law? By your own take, the use of .today links amount to violation. For not being sued for contributory copyright infringement by linking to .today links for so long, did we assume fair use in the manner of Internet Archive does? How do we then balance verifiability of cited matrials and still following law by not linking to archival sites? Smlckz (talk) 17:16, 12 February 2026 (UTC)
does it make us a fool of ourselves for still following the law? I don't think "someone else broke the law" is a good reason to break the law yourself. There are ongoing lawsuits relating to the copyright status of LLMs, and I'm sure we will have more clarity in that field soon.
By your own take, the use of .today links amount to violation. Maybe. By my understanding, .today processes takedowns of singular archives but not entire sites. They do, however, run ads on mobile. I'm not a copyright expert and I don't want to be seen as making a determination, but .today is at the very least in a gray area. You could say the same about IA, but their willingness to exclude sites and not intentionally get around paywalls, as well as being a non-profit, makes their argument of fair use much stronger.
I'm not here to be super strict on copyright and demand that we stop using archival sites, I just do not believe it is wise to make the project dependent on shady, gray area tools like .today. Web archivers are necessary for producing the encyclopedia, so we can't drop all of them. Sticking with the ones with the strongest copyright policy is our best bet. IsCat (talk) 17:31, 12 February 2026 (UTC)
Virtually everything on the wayback machine is a copyright violation and it meets none of the American requirements for fair use. Doing partial takedowns if anything makes their case worse because they are aware the rest is infringment but keeping the rest of fheir collection. Just because they obey takedown requests does not make it not infringement, it just means they may not face as severe legal repercussions. PARAKANYAA (talk) 18:28, 12 February 2026 (UTC)
I mean, yeah. There is a good argument that IA is copyright infringement. They don't legally have an obligation to patrol their site for copyright infringement, only to respond to it if they receive a takedown notice. This was a key issue in Viacom v. YouTube, where YouTube was found to not liable for violations of Viacom's copyright as they responded to takedown requests and was thus protected as a service provider under the DMCA. This doesn't meant they weren't hosting content that infringed copyright (they were), it was just that they weren't liable. IA is in a similar situation, they may be violating copyright but they respond to takedowns and there isn't any caselaw over the fair use status of web archiving. Thus, they can operate under the assumption that they are not violating copyright, even if they are. IsCat (talk) 18:38, 12 February 2026 (UTC)
Sure, IA will probably not crumble into dust tomorrow. But neither will .today. The issue people are raising is that the material we are linking to is a copyright violation. There's plenty of caselaw over both library archiving and internet copyright that makes this situation not ambiguous. Whether IA will be fine is irrelevant. PARAKANYAA (talk) 18:43, 12 February 2026 (UTC)
Yeah, I didn't raise copyright in my !vote on the .today RFC for this reason. My main concern here is making a potentially Wikipedia-affiliated archival platform that is not responsive to DMCA requests, which is what this thread is proposing. That site would be liable for copyright infringement under the DMCA, and the proposed solution is trying to make it impossible to enforce the law against the site. It's a bad idea all around. If someone out there wants to make that site and potentially assume liability for copyright infringement, they can do that. It would be a violation of COPYVIOEL to link to it on WP either way. IsCat (talk) 18:53, 12 February 2026 (UTC)
Well the WMF can't make it but linking it wouldn't be any worse than linking IA. Whether is a copyright violation is unaffected by the DMCA takedown, that is just the enforcement mechanism. PARAKANYAA (talk) 19:32, 12 February 2026 (UTC)
IA can operate under the assumption that web archiving is fair use because there isn't caselaw that clarifies either way. We can assume the same. IA responds to DMCA takedowns and is thus protected by section 512. The archival site herein proposed would not respond to DMCA takedowns and would not be protected by section 512. Linking Google Books is entirely different than linking Anna's Archive for much the same reason. IsCat (talk) 19:36, 12 February 2026 (UTC)
That they comply with DMCA takedowns shows awareness that the content on their site is copyrighted and not fair use. So, we can assume that it is copyrighted by the same. Again, whether someone is protected by section 512 is entirely about punishment or not punishment, it is not about whether the materials at issue are fair use. PARAKANYAA (talk) 19:41, 12 February 2026 (UTC)
They can accept DMCA takedowns while not admitting liability or that they host other infringing content. Section 512 does protect from certain liability, including from lawsuits. This limits the options of rightsholders and more or less forces them to use the takedown system. This has kept web archiving out of courts and has caused the lack of caselaw in the area. Without caselaw, IA can assume that they aren't violating any laws. IsCat (talk) 19:48, 12 February 2026 (UTC)
Yes, their liability is not the issue here. Whether they are liable or not is unrelated to whether the content itself is fair use/an infringement of the copyright holder's rights. And again: does this not go for any web archive, including the hypothetical one you take issue to? PARAKANYAA (talk) 21:05, 12 February 2026 (UTC)
I don't think you are understanding the point I'm trying to make. Liability matters a whole lot in lawsuits. The purpose of a suit is ostensibly to determine if an entity is liable or not for something. If IA is not liable for infringement, they will likely not get sued. If nobody is getting sued, there is no caselaw. If there is no caselaw, it's a gray area. IA claims fair use. Unless and until a court of law (or Congress) acts to determine if web archiving is fair use or not, they can continue to assume its legality. The presence of infringing content is entirely besides the point because we are assuming fair use.
As for the proposed archival service, they would be liable under section 512 as soon as they fail to comply with a takedown request. A defining feature of the proposal is hosting in foreign jurisdictions where copyright law cannot be enforced. IA and this proposal cannot be further apart in terms of liability. IsCat (talk) 22:00, 12 February 2026 (UTC)
That they comply with DMCA takedowns shows awareness that the content on their site is copyrighted and not fair use.
No, it shows awareness that some content on their site may be copyrighted, just like any site that allows user contributions. Legally that's a huge difference. It doesn't matter whether it's 99% copyrighted, or 99% original with only a single copyrighted item, as long as they do respond to takedowns. That's the legal obligation and the reason they aren't liable, the same way Facebook isn't liable for user content as long as they respond similarly. ChompyTheGogoat (talk) 12:57, 14 February 2026 (UTC)
From what I can understand from the discussion here, IA honoring DMCA takedown requests shields them from liability of copyright infringements. So is the case with IPFS gateways. Then, can we link to pages archived on IPFS without problems, right? While the IPFS gateways honor the request, the content might still exist in the network, so those with need can still use the content ID to find it anyway.. Those with vested interest in maintaining an archive of internet can continue to pin the content in copyright-havens. Can this situation be any better? Smlckz (talk) 02:42, 13 February 2026 (UTC)
I mean, true. Plagiarism is wrong, which is why I use websites like Anna's Archive when tracking copyright violations on Wikipedia, which are (a) usually plagariasm (b) pose a licensing issue, since our content should be in our own words and CC-BY, not copyrighted text. Cremastra (talk · contribs) 00:30, 13 February 2026 (UTC)
Internet Archive is still copyright infringement no matter how much they get away with it. PARAKANYAA (talk) 18:22, 12 February 2026 (UTC)
I'm not disputing this, for the record. IsCat (talk) 18:28, 12 February 2026 (UTC)
You did say we can assume fair use, which I do not think is true. PARAKANYAA (talk) 18:31, 12 February 2026 (UTC)
You are also a well known copyright maximalist, to be blunt, who puts almost all power with rights holders. This is a strain of thought you can find, particularly in Europe. But courts in America tend to be less severe and more generous with public benefit. -- GreenC 19:16, 12 February 2026 (UTC)
I do not like copyright. I am, if anything, a copyright abolitionist. But that I want it to be a different way doesn't mean it is. PARAKANYAA (talk) 19:43, 12 February 2026 (UTC)
Copyright abolishment is an extremist position outside the mainstream. If that is your basis you might (consciously or not) try to pit copyright law against itself playing into the divisions to stoke and inflame. But please don't use Wikipedia for that end it is disruptive. -- GreenC 20:02, 12 February 2026 (UTC)
I am in favor of presenting reality as it is. PARAKANYAA (talk) 21:02, 12 February 2026 (UTC)
I don't think there is any need for endorsement by Wikipedia, just the need for some team to set such up. on IPFS or equivalent content-addressed cache sounds like a great idea. Anna's Archive already has links to IPFS-stored content, maybe some IPFS discussion place would be a good place to propose this. One could include in the proposal the idea for the archival site to archive all external links found/crawled on Wikipedia (that aren't yet archived in the Internet Archive and ideally these as well). Prototyperspective (talk) 15:52, 12 February 2026 (UTC)
One such place is another and maybe also . I would not continue the discussion about legality or ethicality above (or the assumptions underlying the various claims for and against such) as there is no need for Wikimedia officially starting or backing this. Also forgot to mention that archive today is not just used for standard Web archiving but also for making paywalled content accessible. If you would like to debate this subject or to see more data & claims relating to it, see the structured argument map Do paywalls around academic publications slow scientific progress?. Prototyperspective (talk) 17:35, 12 February 2026 (UTC)
I asked in the IPFS Matrix channel, where they pointed me to , where reading through the documentation of I found that it has "experimental support" for IPFS:
I also found this project (which began a decade ago!): ..
Well.. it seems some solutions are already there, but are really obscured. Smlckz (talk) 04:41, 17 February 2026 (UTC)
Thanks for doing that – interesting. However, those aren't yet full solutions – rather potential components of potential solutions. A full solution would for example archive all external links used anywhere in Wikipedia using IPFS and then make these available using some website as well as regularly scan for new links. I think IPFS would probably be the best technology to use here but maybe there's some similar alternatives like ZeroNet(?) I wonder whether one could make a wish in the m:Community Wishlist for some tool that fetches all external links used in Wikipedia maybe from the dumps so that these can be all archived with a tool like the one(s) you linked which also would probably need to developed further for this to work and then be implemented in the project to archive all the links (except video media embedded in them) and make it possible for users to access them and store them decentralized. Prototyperspective (talk) 19:07, 17 February 2026 (UTC)
  • Who says web archiving is a copyright violation? That's not totally true. It misunderstands how copyright works on the Internet. It also misunderstands library law and service provider law, which is not the same law as you and I operate under. There are specific statues in the law that allows for web archiving, when done a certain way, by certain parties. I have yet to see a single person on Enwiki who can explain why archive providers have operated for 30+ years and are not immediately shut for copyright violation. I understand why, but most people can not explain it. -- GreenC 17:42, 12 February 2026 (UTC)
    I'm not aware of any federal statutes that specifically allow for web archiving, could you cite what you're referring to?
    Your comment intrigued me so I found a law review article from Prof. Paul D. Callister, a professor of copyright and library law at the University of Missouri-Kansas City. You can read it here. In it, he claims that copyright law is fundamentally dissonant with archival practices. He notes that in a review of caselaw, he was unable to find any case were "preservation" were found to be transformative or fair use. He criticizes the claim that IA is somehow free from copyright infringement just because they've been around for so long, and he also notes that he was able to find any litigation against IA over web archiving. He argues that this is likely due to IA's strong takedown procedures instead of the law protecting IA. He concludes that it would be beneficial if web archiving was made to be fair use under law, but without a court deciding on the issue or Congress acting, web archiving will continue to be at the very least a gray area.
    If you could point me to differing viewpoints or caselaw/statutes that contradict this, that would be great. IsCat (talk) 18:12, 12 February 2026 (UTC)
    Also: I have yet to see a single person on Enwiki who can explain why archive providers have operated for 30+ years and are not immediately shut for copyright violation. I understand why, but most people can not explain it. Archival sites are understood to be service providers under section 512 of 17 U.S.C. and are thus not liable for monetary damages in a copyright lawsuit unless they fail to remove the content upon a takedown notice. This removes the incentive to file and has generally prevented courts from ruling on if web archiving is fair use or not. This allows IA to operate under the assumption of fair use, even if it does not actually exist under law. IsCat (talk) 18:24, 12 February 2026 (UTC)
    @IsCat I think you have a typo there: "he also notes that he was able to find any litigation against IA over web archiving". You probably mean "unable to find"... David10244 (talk) 06:12, 17 February 2026 (UTC)
    See, for example: "The general problem for libraries, under the code section, is that they are only allowed to make archival copies of what they actually have in their collections.Footnote164 Neither Harvard’s law library nor registrant libraries are likely to own the Web materials that users contribute to the Perma.cc archive (although it is possible). Also problematic is several of § 108’s provisions require that the “copy … becomes the property of the user… .”"
    "The article concludes that Perma.cc’s archival use is neither firmly grounded in existing fair use nor library exemptions; that Perma.cc, its “registrar” library and institutional affiliates, and its contributors (see ) have some (at least theoretical) exposure to risk; and that current copyright doctrines and law do not adequately address Web archival storage for scholarly purposes." PARAKANYAA (talk) 18:23, 12 February 2026 (UTC)
    Those are legal arguments, and there are also counter arguments. The fact is, there is no law that strictly allows or strictly disallows web archiving. Web archives currently operate under a number of statues including: Fair Use, Section 512 DMCA Safe Harbor, and Section 108. Courts look at loss of commercial potential, the nature of the organization (non-profit library), the nature of the users (Wikipedia researchers and educational purposes). Honoring take down requests for DMCA. Courts exist to protect rights holders and public benefit. We see it in every case: Google Books was permitted to do mass scanning; Hatchet Case which gave Internet Archive a bunch of wins with out of print books and limited previews. Courts will protect interests beyond the rights holders, they find a balance, and the outcomes are complex and detailed. You can't make broad sweeping statements or ridged pronouncements that web archiving is an outright copyright violation. Even when it goes to court, it never turns out so simple because public benefit weighs heavily in court decisions. -- GreenC 19:04, 12 February 2026 (UTC)
    None of the four fair use criteria in the United States apply. They are reproducing the whole work, replacing the exact commercial market for the work, typically of commercial works. Being a non-profit might be considered in the balance of these factors but when the rest of the criteria are like this it is not a get out of jail free card. Per section 108, libraries can only archive materials they have in their collections, which is not The Web. DMCA safe harbor means they are unlikely to face significant penalty, sure, but that doesn't mean what is there is not a copyright violation.
    Google Books had the books they were scanning in their collection and did not allow readers to read the books in their entirety (e.g. it did not replace the commercial work and used a limited portion, unlike web archiving). The IA lost the Hachette case at every step of the way, "public benefit" or not. PARAKANYAA (talk) 19:38, 12 February 2026 (UTC)
    Your are repeating the most negative articles and opinions. And no they didn't loose every step of the way, again, you take extreme simplified black and white positions. -- GreenC 19:48, 12 February 2026 (UTC)
    How are they not replacing the market of the work or reproducing the whole work? PARAKANYAA (talk) 21:02, 12 February 2026 (UTC)
    I'm not even remotely well versed in the fair use doctrine, but surely if the web page is no longer accessible, the rightsholder isn't monetising it, so it's not competing? (Which still doesn't excuse copying extant webpages, though I'm sure a lawyer would argue that accessing a page through the internet archive is significantly less convenient than just going to the original, so it's not a substitute.) JustARandomSquid (talk) 09:15, 13 February 2026 (UTC)
    What is the commercial value of something you are not charging for? Hawkeye7 (discuss) 21:35, 13 February 2026 (UTC)
    Ads, exposure, a following, a good lawyer will come up with whatever's necessary, I'm sure. JustARandomSquid (talk) 22:27, 13 February 2026 (UTC)
    A copyright owner isn't required to be earning revenue from every way possible in order to protect its rights to use an approach in future. Also there are other ways to provide access to content than via the web. And if a rights holder chose to provide free access to their content, that doesn't mean it has relinquished its rights to control how that access is provided. isaacl (talk) 06:15, 15 February 2026 (UTC)
    Exactly. "Free as in beer" is not nearly the level of freedom needed to make a publicly accessible archive, and orphan works still enjoy full protection even though they are ridiculously problematic for potential researchers. DMacks (talk) 06:39, 15 February 2026 (UTC)
    Well, no, but it makes the respect for commercial opportunities part of the fair use argument much stronger. JustARandomSquid (talk) 07:34, 15 February 2026 (UTC)
    I disagree. For example, just because a book is out of print, or not being offered in an e-book format, doesn't mean it's fair use for someone else to create an e-book version. isaacl (talk) 09:26, 15 February 2026 (UTC)
    @GreenC "Hatchet case" --> "Hachette case"? Hachette is a book publisher. David10244 (talk) 06:14, 17 February 2026 (UTC)
    See Hachette v. Internet Archive. WhatamIdoing (talk) 20:08, 18 February 2026 (UTC)
I think a restricted Wikimedia archiving solution for webpages, accessible only to Wikipedians/Wikimedians - similar to the Wikipedia Library - would be a good idea. It would be harder for the general public to check a web reference archived in this way, of course, but I think it would be a good compromise between copyright concerns and verifiability needs. This also given that the long-term sustainability of at least one of the two large archive services we currently rely upon seems very questionable - even if we should decide to retain archive.today/archive.is links, I'm pretty sure they will be a giant heap of dead links not too far in the future anyway. And the Internet Archive, though unquestionably more reputable as a nonprofit enterprise that truly wants to achieve something good for the internet, is under heavy pressure too, technically, politically (given the current political climate in the USA), and copyright-wise. Gestumblindi (talk) 09:36, 13 February 2026 (UTC)
Yes, restricted access would be helpful evidence that the archive does not have a detrimental effect on the market for the content in question. There are other potential compromises, such as archiving textual content only (since that's probably what we want to cite), which would further distance the archive from any suggestion of intent to infringe. Barnards.tar.gz (talk) 13:22, 13 February 2026 (UTC)
It's a bad idea for WP:V. Making it so the only way to verify what Wikipedia is saying is to be an active Wikipedia editor harms our readers. TWL is acceptable because all the resources it has are available through others means, including other libraries. IsCat (talk) 16:07, 13 February 2026 (UTC)
Well, imagine archive.today and the Internet Archive going offline, which is absolutely possible. Then, an archive restricted to Wikipedia editors would still be better than nothing, if copyright issues are otherwise preventing it? That's what I would call a compromise. Maybe access could be made somewhat easier than for the Wikipedia Library, and maybe some way for external users to have access on request could be implemented. Gestumblindi (talk) 19:52, 13 February 2026 (UTC)
I have a somewhat wackier suggestion: an archive accessible to the public, but new pages can only be archived by autoconfirmed users, like Perma.cc.
The reason I'd want to do this is because a "web archive" can imply multiple use cases. A lot of people use archive.is to hop paywalls, which is obviously a copyright violation of the kind that can incur risk for Wikimedia. If we limit the number of users who can archive pages, we thereby limit bandwidth usage, copyright risk, and scope creep. Furthermore, an archive produced in this way seems to me to have a stronger claim to be a DMCA Safe Harbor than one which resembles a paywall-hopper in function.
If we take measures to limit the scope, perhaps we could open this to the public without incurring as serious a risk. NotBartEhrman (talk) 00:00, 14 February 2026 (UTC)
It would probably create a headache for us if Wikipedia becomes known as a provider of a paywall bypassing tool ("just make 10 junk edits to get access!"). Also, I think archiving sites that successfully bypass paywalls are using dark magic that a reputable site would struggle to replicate. Barnards.tar.gz (talk) 13:35, 14 February 2026 (UTC)
Bypassing paywalls should never be the goal anyway. The purpose of web archiving is not to access material for free that is still perfectly accessible online, though requiring payment (a printed book is also a legitimate source, and books aren't free as well, still you might be able to get them on loan through a library - just as some online content through the Wikipedia Library), but to save material to keep it accessible when it's otherwise offline and not available through any other means. So, a hypothetical Wikimedia archiving service would, IMHO, best be set up in a way that you can preemptively save web pages that are used as references, but the archived copy is made accessible only if the original source goes offline, and probably with some further restrictions (only for Wikipedia editors or something like that). Gestumblindi (talk) 14:11, 14 February 2026 (UTC)
Jinx! ChompyTheGogoat (talk) 14:24, 14 February 2026 (UTC)
This is more or less the kind of "dark archive" that I have been thinking about. Every page used in a citation would get automatically saved, and the saved copy goes into a lockbox that can only be opened if the original source is offline and the reader has WPL access. Stepwise Continuous Dysfunction (talk) 02:55, 15 February 2026 (UTC)
I think archiving sites that successfully bypass paywalls are using dark magic Well, you can consider the design/development process behind youtube-dl, Invidious and alternative web front-ends.. Few are like Wikipedia to provide good public APIs, most are pretty hostile to the development and continued operation of alternative front-ends, devolving into a cat-and-mouse game.
For websites, apart from paywalls, there's also the hurdle of captchas, or more recently Anubis and so on, that have been deployed to protect the sites from recent waves of DDoS from LLM scrapers: If web crawlers used to have some sense of propriety back in the days, when crawlers actually followed instructions of robots.txt and set apprpriate user-agent string, it appears to be quite lacking in these days of LLM scrapers: with residential botnets and user agents of mildly older browsers being common occurrence. Even if a FOSS solution is developed that can overcome the restrictions of websites, and then gets used by these LLM crawlers, it'll be a rather disgraceful situation, not considering that mitigations that will also be developed accordingly, leading to a protracted stand-off. A non-FOSS solution will create a dependance and might lead to a situation like that of .today. Smlckz (talk) 18:11, 14 February 2026 (UTC)
I'd take it one step further: the complete archive only accessible to users with special permissions, whose responsibility is to set individual page archives to public access if and when the original becomes unavailable - similar to how the deadlink status works now, but with additional protections on the archive to prevent abuse. Deadlink flags by other editors would add it to the approved editor log for confirmation. IANAL but I believe that would be along similar lines to current Commons policy on non-free images, which are only allowed if no suitable free alternative is available - the archived version would only be publicly accessible if the original copyrighted source is not. I don't think it matters quite as much who triggers the archival process (anyone can on IA) if it can't be viewed unless those conditions are met. The archiving could even be done by bots crawling wiki citations to ensure they're backed up.

I also agree that limiting the archives to text would help strengthen the fair use argument. If and when media qualifies for archiving is a discussion that should probably take place on Commons, amongst people more familiar with those policies, but afaik most media is not used as sources to substantiate claims, right - if it proves a claim we'd look for a text source explaining that? ChompyTheGogoat (talk) 14:23, 14 February 2026 (UTC)
As a Commons admin, I just have to point out that there is no "current Commons policy on non-free images, which are only allowed if no suitable free alternative is available"; Commons doesn't allow any non-free images, I think you're mixing this up with English-language Wikipedia's policy for non-free images locally hosted as "fair use". Gestumblindi (talk) 15:26, 14 February 2026 (UTC)
Quite possibly! Still new here - I thought I'd seen links to Commons in discussions on the subject, but I might be misrerembering.
m ChompyTheGogoat (talk) 15:30, 14 February 2026 (UTC)
(The point still stands however.) ChompyTheGogoat (talk) 15:31, 14 February 2026 (UTC)
Yes, a reasonable point otherwise, I agree. Gestumblindi (talk) 15:44, 14 February 2026 (UTC)
You're probably remembering the local Wikipedia:Non-free content criteria. WhatamIdoing (talk) 03:21, 17 February 2026 (UTC)

General Background Color

Hello. [I have not seen much of the inner workings of Wikipedia previously, so please excuse me if my letter format, including the compliments preceding my idea, is thought of as superfluous.]

Thank you all for all of your efforts collectively for 25 years.

It's apparent that you've made several advancements in the past couple of years regarding your site's appearance, specifically in the ability to manage content organization.

I have a suggestion just regarding the overall visual, and it should be fairly simple, because of the CSS layer of website code. It might take even a single change effecting all of your pages.

Wikipedia has always conveyed with its bare appearance that it was set up in someone's spare time, even though the content elements have become very sophisticated.

But given that you're touting your 25 years, and drawing attention to your extensive processes and the many people involved, a slightly fancier presentation appears long overdue.

Please make the background of your pages a faint beige. It adds a touch of class and generally makes text easier to read, given that it provides a tiny bit less contrast, as black on white glows fairly quickly. Color #FFFEF5 is an example.

With the various options that you've been adding, including the dark option (which takes contrast to its own level), maybe you'll feel it best to make this an option.

Thanks for your time and consideration. ~2026-90420-1 (talk) 04:31, 10 February 2026 (UTC)

Making it an option doesn't seem too bad to me, however I wonder what others think about this. Cdr. Erwin Smith (talk) 07:00, 10 February 2026 (UTC)
Personally, I've never quite understood why people want everything to be grey-on-grey rather than adjusting the brightness of their display if things are too bright for them. Anomie 12:47, 10 February 2026 (UTC)
Just to keep thoughts on the actual suggestion, it's all of the existing colors on beige, not grey on grey. And darkening your display lessens all of the colors, so it actually makes things more grey on grey.
Also, this isn't about every other website or every other application, most of which don't have this problem. This is about Wikipedia. So, it's inappropriate to adjust your display for just Wikipedia windows.
And this isn't a random suggestion; it's an actual step that has improved readability in various presentations. A hint of beige just makes everything less harsh and helps people to keep reading, which is the intent. ~2026-90420-1 (talk) 15:48, 10 February 2026 (UTC)
If we want to make it an option, we should probably take this to Meta-Wiki or MediaWiki wiki so that it would be on all Wikimedia projects. ✨ΩmegaMantis✨(he/him) ❦blather | ☞spy on me 21:00, 10 February 2026 (UTC)
m:Community Wishlist/Wishes is the place to go. Ponor (talk) 19:14, 11 February 2026 (UTC)
Old people here remember the times when #ffffec was the background color for non-mainspace pages. sapphaline (talk) 16:10, 10 February 2026 (UTC)
I'd prefer a #FFF8E7 myself to gain street cred with the astronomy nerds. :) ✨ΩmegaMantis✨(he/him) ❦blather | ☞spy on me 21:11, 10 February 2026 (UTC)
The Wikipedia app (at least on iOS) has a sepia tone option. It could be smart to add it to the web version. ✨ΩmegaMantis✨(he/him) ❦blather | ☞spy on me 20:55, 10 February 2026 (UTC)
One idea I have toyed with in the past is going the other way, applying a very light tint for non article pages - a pale green for draft pages, a pale red for old revisions, or something like that. Enough to make it visible at a glance that a reader is not looking at an actual live article. Never got around to mocking up some CSS for testing it, though. Andrew Gray (talk) 21:02, 10 February 2026 (UTC)
That would be quite useful, though I could see people getting annoyed about it. Interestingly, when I use the Wikipedia app, I have mode set to sepia tone, but since pages in the Wikipedia namespaces and some other namespaces aren't exactly supported as in app mode and just open a mobile version, the background turns white and so I get a version of what you are describing, in reverse. ✨ΩmegaMantis✨(he/him) ❦blather | ☞spy on me 21:07, 10 February 2026 (UTC)
@Andrew Gray I could have sworn that back in 2006 or 2007, non-mainspace pages had a light blue or green background. ~ ONUnicorn(Talk|Contribs)problem solving 21:30, 10 February 2026 (UTC)
@ONUnicorn ...you know, it wouldn't surprise me if I have been working off a hazy memory of that! It doesn't seem to be the case for any of the main ones listed at Wikipedia:Skin (I tested them all on WP, WPT, and mainspace) but possibly it was there for a period and off again. Andrew Gray (talk) 23:25, 10 February 2026 (UTC)
Monobook still does. Anomie 00:15, 11 February 2026 (UTC)
"Chick", "Standard", "Nostalgia" and "Simple" used to change the color for non-mainspace pages to #ffffec by default until their deprecation in 2013; Cologne Blue still does this, also by default (for example). Monobook doesn't do this by default, it's an old on-wiki customization. sapphaline (talk) 10:32, 11 February 2026 (UTC)
Wikipedia has always conveyed with its bare appearance that it was set up in someone's spare time. Nothing wrong with that. Phil Bridger (talk) 23:27, 10 February 2026 (UTC)
Thank you very much for showing all of the colors that have been mentioned. And it's interesting to see that the text color isn't quite black. Regarding the exact shade of beige, it should just be kept in mind that the color as an entire background seems darker than a box sample, and I think that subtlety might be most effective for extended reading. --- Also, to help get this implemented, for the most common situation, is it best to keep it as simple as possible? The other ideas that are building on the idea to just change the standard background appear to be helpful for more-specific situations, so they might be best as separate options. --- For the vast majority of the situations, the other ideas might be overwhelming to users. This includes that users might not understand why they're seeing differing backgrounds. For that purpose, it actually might be most helpful to have a clearly worded (and distinctly colored?) banner at the top of the content for unusual types of content. ~2026-90420-1 (talk) 14:48, 11 February 2026 (UTC)
So people can see the colors being mentioned:
  •   #FFFFFF (actual white)
  •   #FFFEF5
  •   #FFFFEC
  •   #FFF8E7
  •   #F8FCFF (MonoBook)
  •   #202122 (normal text on wiki)
  •   #000000 (actual black)
WhatamIdoing (talk) 23:46, 10 February 2026 (UTC)
I appear to have replied to the wrong message, above. It might be helpful for this message board to have a horizontal separator line after each non-indented (top-level?) message. In any case, I've pasted my message here: --- Thank you very much for showing all of the colors that have been mentioned. And it's interesting to see that the text color isn't quite black. Regarding the exact shade of beige, it should just be kept in mind that the color as an entire background seems darker than a box sample, and I think that subtlety might be most effective for extended reading. --- Also, to help get this implemented, for the most common situation, is it best to keep it as simple as possible? The other ideas that are building on the idea to just change the standard background appear to be helpful for more-specific situations, so they might be best as separate options. --- For the vast majority of the situations, the other ideas might be overwhelming to users. This includes that users might not understand why they're seeing differing backgrounds. For that purpose, it actually might be most helpful to have a clearly worded (and distinctly colored?) banner at the top of the content for unusual types of content. ~2026-90420-1 (talk) 14:53, 11 February 2026 (UTC)
Android app sepia theme
Iff this was to be implemented, it should definitely be an opt-in setting just like the Dark mode.
Bikeshedding on the precise background colour, it would make sense to use the same as the Sepia theme in the Wikipedia apps: currently #F8F1E3  . the wub "?!" 17:54, 11 February 2026 (UTC)
Thanks for the exact Sepia shade sample as well. Does everyone like the slightly darker Sepia or a slightly lighter beige? What shade seems to make the text easiest to read? --- And hey, let's talk about the birds of Europe! ~2026-90420-1 (talk) 20:53, 11 February 2026 (UTC)
  is too subtle, it would make almost no difference at all.
  looks better than the former, but a bit too dark.
  is a lighter version of the former, and looks ideal to me. But hey, it's just me. Cdr. Erwin Smith (talk) 08:25, 12 February 2026 (UTC)
Okay, I'm sorry about this. I originally copied the wrong color code. I've actually been looking at EEEDDC in another window the whole time. Its tint is more brownish than the Sepia, which is a little bit peachy. And it actually seems close to the third color in the above message. (For anyone who doesn't have access to web code, the color code can be put in Word's Page Color > More Colors > RGB option to see it at scale.) ~2026-90420-1 (talk) 12:43, 12 February 2026 (UTC)
That's all right.
Since we have had a good discussion, I would now recommend you to open an RfC on the Proposals Tab.
Be sure to add all 4 of the colour options, 3 here and 1 here, for a comprehensive debate. Cdr. Erwin Smith (talk) 07:10, 14 February 2026 (UTC)
This is really interesting - I had been looking at these on my desktop and thinking, huh, they're all the same blank white, is it a really imperceptible shade? But on the phone they're all very distinct. I wonder what configuration weirdness I have going on here. Andrew Gray (talk) 21:13, 11 February 2026 (UTC)
Multiple reasons. First of all.. by default a lot of LCDs for home computing are calibrated to emit WAY too much light and are also not adaptive to your surrounding light situation. It can pay off to spend some time to actually calibrate your screen (like it tells you to in the setup manual that no one ever reads). Mobile phones and often TVs as well are adaptive and will change their light intensity on the fly. This makes sense because they have to deal with a lot of changes during the day due to changing light conditions. In a work environment, you often want a more consistent color profile, instead of something that changes all the time.

Secondly: LED vs LCD. Expensive TVs and phones often have LED screens. LCD screens have to illuminate the entire back of the screen with light, even in the black areas. Combined with oversaturation of light LCDs will be more effected by not well calibrated screens.

Thirdly: Desktop and laptop screens are often pretty bad. Mobile phones are, due to their size (easier manufacturing) and the competitiveness in the market often higher quality, with shorter lifetime cycles. While in the laptop and desktop market the screen is an expensive part that can easily be saved on by manufacturers, with most consumers not really realizing what they are missing in return.

For instance, my home Desktop screen is a 600+ euro 27" LCD (when I bought it 6 years ago), which will show these color differences. But my (newer) workplace 27" LCD is a much cheaper LCD and unless I do a lot of calibration on it, is often not able to show the difference between these colors.

If you wonder why all designers use MacBooks... this is part of the reason. MacBooks by default, tend to come with better screens than most other consumer computers. If you order online and know nothing about screens, with Apple you sorta know what you are getting, whereas with most other brands it will be a complete toss up.
This is also why I always tell everyone over 40 to buy a better monitor, because using a better screen will literally change your life when your eyes start getting worse. —TheDJ (talkcontribs) 14:44, 12 February 2026 (UTC)
@TheDJ Thanks - I think this will be this evening's rabbit hole to go down! I am indeed over 40, and I strongly suspect this monitor is at least old enough to be thinking about its exams... Andrew Gray (talk) 18:45, 12 February 2026 (UTC)
For the record, I support such customization for logged-out users, and honestly Vector-2022 maintainers should provide an option for them to set any background and on-text color they want. sapphaline (talk) 10:38, 11 February 2026 (UTC)
@Sapphaline Such customization is of course possible with CSS variables. There are some 50+ of them that you can override and mess with if you want. This is precisely how dark mode works for instance. Just install an extension that modifies them, or use the browsers local styles, something like Greasemonkey etc. People can do whatever they want by following a few YouTube tutorials. However compiling and maintaining multiple of such sets of 50+ variables is a pretty expensive operation and its is very easy to introduce problems with accessibility etc. That's why it is left up to users to do this if they want to. —TheDJ (talkcontribs) 14:52, 12 February 2026 (UTC)
Please do no such god damned thing. The clean white background is far classier. --User:Khajidha (talk) (contributions) 22:38, 18 February 2026 (UTC)
It'll be an option : ) Cdr. Erwin Smith (talk) 05:50, 19 February 2026 (UTC)

Planning help for AI workflows

A discussion has been started about what help page on using AI on Wikipedia to draft first. See Wikipedia talk:Help Project#Planning help for AI workflows.    The Transhumanist   13:35, 19 February 2026 (UTC)

a addition of a blogs to wikipedia

by what i mean is a separate part of wikipedia that is purely for sharing interesting facts like average probability of a certain leaf, there is no widespread interesting fact sharing platform to my knowledge besides quora posts bug this would be more things you know, whole quora is more lessons you learn from what i've read and quite alot of people would like to fill that feeling, the best place for this would be to do is specific tumblr communities and those are more about fiction to my knowledge, this would be more about real life topics, it would feel out something akin to watching those interesting videos on relatively random topics, it would be separated into many sections, the main thing against this idea is that "why would this encyclopedia have unsourced content?" and to that i say because i think a majority of wikipedias users would also like this feature because they probably know alot of things they would like to share, the actual reason this shouldn't get added is because it would probably be better as separate wiki project, very different from anything to do with wikipedia and basically any other wide stream wikiproject. Misterpotatoman (talk) 09:20, 3 February 2026 (UTC)

There are plenty of sites that do what you want. Why do you feel that it's appropriate for an encyclopedia to have blogs? Phil Bridger (talk) 09:27, 3 February 2026 (UTC)
@Misterpotatoman can you clarify your idea a bit further. I think having a shared blogs (what editors are working on now) as part of a project could be interesting and build community, especially if we could somehow grab project members edit summaries to do with articles that are part of the project Not certain about on article as we try and avoid diverting people, but diverting high conflict editors from protected pages could be good.. (With fun, I find wiki meetups and pages to do with Wikipedia on FB and Reddit and mastadon helped.)  Preceding unsigned comment added by Wakelamp (talkcontribs) 10:51, 3 February 2026 (UTC)
i think there could be something be blogs purely for sharing interesting things and facts, like those interesting youtube videos but text, it would be separated into alot of topics, as far as i know theres no widespread interesting fact sharing platform anywhere and i good chunk of people want to share interesting facts, Misterpotatoman (talk) 05:54, 5 February 2026 (UTC)
Why would an encyclopedia want to associate itself with something that would have no sourcing or quality standards? Why would a person reading an encyclopedia want to read your unsourced gibberish? --User:Khajidha (talk) (contributions) 15:55, 4 February 2026 (UTC)
"unsourced gibberish" - this is Ideas - blasting people you don't agree with is unhelpful ~``Wakelamp (talk) d[@-@]b
People posting random stuff that they think they know, without anything to back it up beyond "trust me, bro" is far more unhelpful. --User:Khajidha (talk) (contributions) 13:15, 9 February 2026 (UTC)
well, i guess we could also implement a safeguard which doesn't allow you to publish until there are a specified number of citations per para/heading (to prevent loopholes). the rest of it has to be handled by the editor publishing the blog.
in any case, calling it "unsourced gibberish" is really just overextending. why can't you call it "uncited"? that would be far more acceptable/ woaharang (talk) 14:44, 19 February 2026 (UTC)
  • That’s the problem… I have a feeling that any such project (hosting unsourced content) would quickly devolve into little more than “everyone blasting people they don’t agree with”. Blueboar (talk) 01:07, 6 February 2026 (UTC)
    Wikipedia already has that – the Village Pump pages, for example. If you look at Top 1000 edited pages in last 30 days, you see that there's lots of activity on non-article pages. Andrew🐉(talk) 08:09, 6 February 2026 (UTC)
  • Bad idea, in my view. Jusdafax (talk) 08:26, 6 February 2026 (UTC)

Coloured WP:RSP

Automatic categories via WikiData

better calculation where to suggest on a zero return search.

LLM harm reduction policy

Process wikipedia into tree knowledge dependency data structure using AI for human use in MMU RAG chats via api (et cetera)

Following talk page only?

Limiting Temporary Accounts to Article Edits and Article Talk Pages

Language redirect idea

Five strikes down to three

Baby Globe Fear

I have an idea

Preventing abuse of the DEPROD system

Searching by category with a higher depth

Thanking for things other than edits.

book reading mode

A lower-contrast dark mode?

Other areas of Wikipedia

Creating the "Disambiguation" and "Disambiguation talk" namespaces

A dataset containing 200,000 architectural objects: is it of interest to the community?

Move Switcher Gadget Buttons to the Top

award for collaborating well in CTOPs

Template for downed features

Collapsing infoboxes for mobile users

Show new users what an encyclopedia is!

edit notice for FAs

Make non-free use rationale template fill out "n.a." values with some boilerplate fluff

Better support for LaTeX markup

Related Articles

Wikiwand AI