Wikipedia talk:WikiProject Citation cleanup

From Wikipedia, the free encyclopedia

Pages with duplicate reference names

This is the first WikiProject I'm joining. I forgot where I originally found the Category:Pages with duplicate reference names link from, I think it was from a technical discussion related to adding warnings and automatically categorizing articles with redundant reference names. As I started slowly going through it, I wondered if there was a related WikiProject, and this seems to be it. The backlog is so huge, that I'm inviting others to also join in . I'm not sure who started, but when I did, apart from recently changed articles, the A-B was already done.

It's a mostly systematic and boring task, but apart from exceptions, per-article work is not too long, so it's easy to just put in some available time, and stop anytime. I so far found a few sometimes challenging aspects, maybe some of this could eventually be part of the WikiProject documentation:

  • Infoboxes may already have reference tags, so duplicates may still show up for some sources when the deduplicating work is fully done. Checking the infobox template, then fixing the article source to use that reference may be necessary. I have lost half an hour fiddling with references on an article before discovering this. Common examples are country census data.
  • Some climate articles include chart templates causing some citations to clash.
  • It may be difficult to quickly distinguish identical and redundant inline citations. When an article is especially messy, copying all or most citations to the reflist (Help:Footnotes#WP:LDR), and then replacing inline references with tags seems to help.
  • It is unfortunately common for articles to have multiple citations with a common redundant tag, as well as a number of inline references to that tag. This means that precise information, if it once existed, was already lost. Checking external references to fix these may take longer, or may not be possible. When those point to multiple parts of a single site, or multiple pages of a single book, a strategy may be to consolidate those into one generic link to that resource (i.e. a range of pages, or sometimes just the whole book, etc).
  • I come across various citation templates which I'm not always familiar with, so some articles require me to read the Wikipedia citations documentation.
    • Song articles often use the singlechart or Certification_Table_Entry templates which support the optional refname parameter.
    • Schools using the NCES School ID template support the optional ref_name parameter.
  • The red error message shown about the redundancy shows a tag name. This name is not always obvious to locate in the article source, because various characters are substituted. Moreover, some citation templates automatically aggregate information into tags.
  • And of course one encounters many other types of citation related bugs when doing this (incomplete or obsolete accessdate, external links in website, red links, etc). It can also be the opportunity to add relevant top-page tags, categories or stub-template whenever necessary.

That's about it for now I think. PaleoNeonate (talk) 19:21, 27 March 2017 (UTC)

Bare URL drive?

I got this idea from Tails Wx, who proposed a {{citation needed}} drive. I doubt anybody watches this page, but here's my proposal anyways:

  • There are currently 65 000+ articles tagged with bare URL citations
  • This has only decreased by 4 000 since 4 Jun.
  • Bare URLs are for the most part easy to fix.

Perhaps we should initiate a Bare URL drive, like GOCE drives, to reduce the backlog and increase editor participation. Edward-Woodrow :) [talk] 14:23, 21 July 2023 (UTC)

I would still be interested in this. I would be happy to help coordinate it if that is needed. C F A 💬 21:47, 19 July 2024 (UTC)
That would be great! Cremastra (talk) 07:45, 20 July 2024 (UTC)
I've started a draft here with some ideas based on previous drives. C F A 💬 19:33, 20 July 2024 (UTC)

This seems to be a really slow-moving thread, but fwiw, there is a user warning template called {{Uw-bareurl}} that you can add to a user talk page as advice about bare urls. Pinging @Edward-Woodrow, CFA, and Cremastra:. Mathglot (talk) 23:24, 1 May 2025 (UTC)

The backlog is down to ~17,000, but AFAIK ~3 weeks ago it was around 15,000; thus, it seems to be fluctuating as editors both fill in bare URLs and tag articles for bare URLs. I would weakly support a drive at this time.
One caveat however: I have seen various editors using reFill and similar scripts without reviewing the output, inadvertently breaking citations or adding nasty stuff. The backlog drive page should include a warning to prevent WP:MEATBOT violations under penalty of disqualification. OutsideNormality (talk) 03:21, 3 August 2025 (UTC)
I am willing to help coordinate a drive, provided I'm not the only person doing so. I suggest we read through our instructions veeeeery carefully before launching it. I learned on WP:JUN24 that it's very easy to write rules you don't realize are vague or subject to grey areas until you have eight people on the talk page anxiously asking about it. Cremastra (talk · contribs) 03:50, 3 August 2025 (UTC)
Pinging CFA, who has expressed interest. FYI: I am Edward-Woodrow. Cremastra (talk · contribs) 03:51, 3 August 2025 (UTC)
Tagging articles could, and should be done by bot, in my opinion. Why waste our most precious resource (editor time) on that? How are people currently finding them, using AWB? That would probably be fastest, but I would just use an Advanced search.
More information Sample search queries to elicit a list of articles with Bareurls in a given subject area ...
Close
How are people finding them now? Mathglot (talk) 04:48, 3 August 2025 (UTC)
@Mathglot I think we're talking about different things. ReFill is a script used to fill in bare url citations which sometimes produces scrambled references. I have no objections to using a bot to tag articles. Cremastra (talk · contribs) 15:00, 3 August 2025 (UTC)
My response was aimed at OutsideNormality's post, in particular the part about a fluctuating backlog, with some increase in bareurl total count due to recent tagging. The gist of my comment was: let's get *all* of the tagging done right now, before starting the drive. That has several advantages:
  • we get an idea of the total scope of the project;
  • volunteers previously involved with tagging can shift their effort entirely to the resolution phase instead;
  • we will get that satisfying feeling of accomplishment, as we watch the monotonic decrease in the backlog numbers while we work on it (especially if we graph it, so we can actually see it);
  • as it gets lower and lower, it motivates volunteers and spurs us on; maybe new editors see measurable progress, and join in to help.
If we don't do that, than we may see the numbers go up and down as ON reported, where the actual progress of volunteers resolving bareurls is hidden or at least murky, due to sporadic additional tagging. That is a discouraging factor, and may impact the effort. If accurate, that would argue for requesting a tagging bot to be created, and within a day or two of starting it, every bareurl should be tagged. That will then facilitate the backlog drive.
In the meanwhile, there will always be a trickle of new bareurls, and the bot can run continuously, tagging them as they appear. To help motivate and track progress, the bot could save stats in Commons which would then be picked up by embedded charts at this WikiProject, graphed automatically by the new Charts extension. (I have been fiddling with this, and I believe the following graphing proposal is feasible.) Picture two time-series graphs:
  1. Line graph: number of bareurls left; and
  2. two stats: green line (or bar): number of BUs resolved this week (or whatever unit time); red: Number of BUs added this week.
It will be fun and motivating, to watch graph #1 go down and down until it reaches a kind of stasis at some level, with BUs-added ~= BUs-resolved, while graph #2 gives us a good idea of the ongoing maintenance situation, whether incoming tends to hold at a steady rate, or increase over time or seasonally or what (we don't really know that, at this point), and whether we are keeping up, or need a mini-drive down the road to catch up again. Pinging Ahecht, who may have ideas based on their experience with bots and charting. Mathglot (talk) 20:40, 3 August 2025 (UTC)
I was wondering how the backlogs got created, because there certainly aren't THAT many concerned editors adding new tags. It wouldn't surprise me if there was a significant portion of bare URLs that have gone unnoticed, given that the main tagging instances were ran in March/August 2022/2024 and I've already stumbled across multiple pages that have untagged bare URLs just through browsing the site. Might be ideal to run such a bot on the final day of each month if it's too intensive to keep going once a day. Corsaka (talk) 12:32, 5 August 2025 (UTC)

Web refs lacking URLs at Lampsilis abrupta

I've already tagged the page with Template:Full citations needed, but figured I'd flag the issue here as well. Lampsilis abrupta has a large number of bare text references to web pages that lack URLs, all added by @Leahsherise (who has made no edits since) in 2009. If anyone has the time and inclination to attempt to track down URLs for these refs so that the text can be properly verified, I'd be very thankful. Cheers, Ethmostigmus 🌿 (talk | contribs) 01:26, 24 January 2026 (UTC)

 Done. I also dumped some reference ideas on the talk page. Kodning 🌸 (talk) 08:56, 18 May 2026 (UTC)

RfC notice: VisualEditor automatic reference names

Hi, I’m Johannes from Wikimedia Deutschland’s Technical Wishes team. We are considering to work on Community Wishlist/W17: Improve VE references' automatic names and reuse. This has been a long-term issue for wikitext editors (see e.g. WP:VisualEditor/Named references) which has been among the top-voted wishes in several Community Wishlist Surveys, e.g. 2017, 2019, 2022 or 2023.

We would like your input on the solutions proposed on our project page. We are considering several options, which can be combined if desired by the community.

  • Changing the default pattern for automatically generated reference names (currently ":n", e.g. ":0", ":1"...) to use the reference type instead (e.g. "book_reference-1").
  • Providing a simple mechanism for communities to configure a different default name.
  • Generating automatic reference names based on the domain name (if it’s a web citation).
  • Generating automatic reference names based on template parameters (e.g. "title" or "last"+"first") – defined by the community.

Feedback

Visit our project page to read about our proposal in detail and share your thoughts on metawiki.

Please note: We will only implement a solution if there’s clear consensus among the global community. Our intention is not to build the perfect solution, but to find a simple and lean one that alleviates the pain caused by auto generated names. We are aware that some experienced VisualEditor users might prefer an option to manually change reference names in VisualEditor, but such a UX intervention is difficult to achieve across reference types and thus out of scope for our team, we can only improve the auto-naming mechanism. We are happy about suggestions for improving certain details of the proposed solutions. Any other feedback and alternative proposals are also welcome – even though it’s out of scope for us, it might still be relevant for future work on this topic.

Please support us interpreting consensus by clearly indicating your opinion (e.g. by using support/neutral/oppose templates). We are aware of WP:NOTVOTE, but given that we are facilitating this discussion with users from different wikis, potentially commenting in their native language, clearly indicating your position helps us avoid misunderstandings.

Thank you for participating! --Johannes Richter (WMDE) (talk) 11:55, 19 March 2026 (UTC)

Moving on to next stage of the ReferenceExpander clean-up project

Hi all, I've finished working through batch 2 of the RefExpander cleanup project. If you are unfamiliar, the quick summary is that for several years a script was employed that was intended to convert bare url refs into nicely templated citations. The script had little to no safety controls or error handling and a few power-editors employed it quite widely breaking a lot of references. A common example would be a reference that looked liked <ref>[example.com/article Article Title]</ref> being turned into <ref>{{cite web|url=example.com/404 |title=404 |author=website name for some reason}}</ref>. Batch 1 was around 2500 articles and took like 2 years to sift through. Batch 2 was around 800 articles and I have finished it in around a year with on and off efforts at completing it. I am embarking on batch 3 which I think will be the final batch, which is around another 1900 diffs. I assembled this final batch based on what I believe my previous comrades on the project did to make the tables for batch 1 and batch 2. I have low key developed a serious addiction to working my way through these diffs.

Thank you for any assistance on this endeavor.  Preceding unsigned comment added by Gnisacc (talkcontribs) 23:18, 8 May 2026 (UTC)

Related Articles

Wikiwand AI