Wikiwand AI

Wikipedia:Village pump (WMF)

Discussion page for matters concerning the Wikimedia Foundation From Wikipedia, the free encyclopedia

 Policy Technical Proposals Idea lab WMF Miscellaneous 
The Wikimedia Foundation (WMF) section of the village pump is a community-managed page. Editors or Wikimedia Foundation staff may post and discuss information, proposals, feedback requests, or other matters of significance to both the community and the Foundation. It is intended to aid communication, understanding, and coordination between the community and the foundation, though Wikimedia Foundation currently does not consider this page to be a communication venue.

Threads may be automatically archived after 14 days of inactivity.

Behaviour on this page: This page is for engaging with and discussing the Wikimedia Foundation. Editors commenting here are required to act with appropriate decorum. While grievances, complaints, or criticism of the foundation are frequently posted here, you are expected to present them without being rude or hostile. Comments that are uncivil may be removed without warning. Personal attacks against other users, including employees of the Wikimedia Foundation, will be met with sanctions.

« Archives, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17

AI-generated edit suggestions

When Simple Summaries rolled out, the overwhelming response from the community was that we do not want AI features. One year later, we are now getting more AI features, e.g., AI-generated edit suggestions. This has been in the works for about 2 months and is due to be announced any time now, so you heard it here first. This is going to be long, sorry, there is a lot of ground to cover.

First: These appear to be the suggestions, because based on how hard that URL was to dig up I assume the announcement wasn't going to link to them in one place. (It's unclear which of these lists if any is the actual list used in production, but there do not seem to be major quality differences based on spot-checking several of them, the newer lists are not obviously better and the larger lists are not obviously spottier). There are three broad problems here, besides the fact that adding new AI features is the opposite of what people have asked for:

  1. Besides the most obvious low-hanging fruit (typos), the suggestions do not contain any concrete fixes. I am pretty sure this is due to not wanting to create a vector for adding AI-generated text (from the study: We don’t have any plans nor desires to use models to create edits directly), and I agree with that. The problem is that, based on the kinds of suggested edits we've seen in the past (e.g., Newcomer Tasks), people will resolve this ambiguity by using AI anyway.
  2. To that point, it is unclear who this is for. A lot of the suggestions are obvious low-hanging fruit that competent copy editors would have noticed on their own anyway. People who are not fluent in English will not be able to do much with a suggestion like "Rephrase the sentence to present the information in a neutral tone and qualify the superlative with a time reference." Newcomers who might have been overwhelmed will probably be even more overwhelmed because of the lack of direction.
  3. Several of the suggestions are just bad. Sorry, but there's no kinder way to put that: they are bad suggestions that encourage bad edits. This is just a spot check and is incomplete, but I've identified a few common categories of bad suggestions (drawn from multiple lists at the link):

Political/geopolitical nightmares: I strongly suspect the geographical name suggestions are going to or have run into some geopolitical snarl at some point. But the NPOV suggestions already have, and demonstrate the reason why using AI to "correct" non-neutral point of tone is the opposite of "low-risk" (as described in the writeup) and generally a bad idea.

  • FreedomWorks: This article about a conservative organization contain(ed?) the sentence "During the 2020 election campaign, FreedomWorks pushed false and misleading claims about mail-in-voting, targeting ad campaigns on swing states with high concentrations of minority voters." It is cited to a Washington Post article describing clearly false and misleading claims. The AI's suggestion, however, is The original wording presents FreedomWorks' actions as definitively false and misleading, which is non‑neutral. It should be rephrased to a neutral description of the disputed nature of the claims.
  • Yishuv: The original article contains the following sentence: "The League of Nations codified support for the eventual 'establishment in Palestine of a national home for the Jewish people' into the foundational document of the British Mandate in Palestine, thereby facilitating what detractors later regarded not as aliyah, or the flight of refugees from Nazi and Fascist atrocities, but as the Zionist colonization of Palestine." Obviously that's not perfect, but this suggestion seems to be a misinterpretation: The passage uses loaded language that frames the British Mandate support for a Jewish national home as ""Zionist colonization"", reflecting a partisan perspective. It should be rephrased to present the differing interpretations without endorsing one.
  • The LLM is oversensitive to even the most factual descriptions of political views or affiliations. Example: DeAndrea G. Benjamin, re: the sentence "During her confirmation hearing, Republican senators questioned her decisions granting bond and early release of defendants": The phrase ""Republican senators"" introduces a partisan label that is unnecessary for a factual description of the hearing. The senators who raised these questions are factually members of the Republican Party, all three sources frame it as such, and the partisan breakdown is an inherent fact of the situation.

Suggestions to introduce factual errors:

  • Frost Children, regarding a sentence about the album called Smile! :D: The sentence contains an emoticon, which is non‑neutral and informal. If someone followed this suggestion they would turn a correct statement into a wrong one.
  • 1797 in Denmark: The word "pyusician" is a misspelling; it should be corrected to "musician". The actual correct spelling here would be "physician".

Parsing errors producing nonsense: This is the same problem Simple Summaries had. The text parsing breaks in many ways, and the LLM generates suggestions based on the broken version.

  • The LLM has trouble with wikitables, and will often directly reference a JSON snippet it received, resulting in bizarre suggestions like Remove the nonsensical JSON list and replace it with a brief, readable statement or omit it entirely. (Markazi Jamiat Ahle Hadith)
  • The tool seems to assume that excerpts are one full sentence long, even when they're not. This results in several suggestions to "break up" sentences that already are. This happens a lot, but an illustrative example is Santa Fe Place (the "Stores" paragraph), where the actual prose problem is the opposite as Overly long sentence with many clauses; needs to be broken up for simplicity): the sentences are choppy and some could be combined. (also there's an obvious comma splice the AI fails to mention)
  • Sometimes titles, sidebars, etc. get interpreted as article text: 2019–20 Philadelphia 76ers season: The lead repeats ""NBA professional basketball team season"" twice, creating redundancy. Obviously, it does not; the culprit is that the sentence the LLM interpreted was "NBA professional basketball team season NBA professional basketball team season The 2019–20 Philadelphia 76ers season was the 71st season of the franchise in the National Basketball Association (NBA)." (See also Boadicea Haranguing the Britons, where it does this for the template)
  • This also happens with templates, as in Chen Lijun (actress): Redundant repetition of the subject’s occupation; the phrase “Chinese” appears twice. The first "Chinese" comes from the "lang-zh" template.
  • Cross-wiki links get it too, as in Switzerland in the Eurovision Song Contest 1966: Extraneous ""[it]"" markup after Mascia Cantoni; it is an artifact of extraction and should be deleted.
  • So do blockquotes, as in Nacht und Nebel: The passage contains a run‑on sentence and an incomplete citation phrase ""According to historian Wolfgang Sofsky:"" that leaves the reader expecting a quotation.
  • So do stub templates: The stub template line includes an unnecessary ""vte"" fragment and could be phrased more cleanly. (the "vte" shows up a lot) or Remove the redundant stub messages that appear as ordinary text at the end of the article.
  • Some suggestions are already fixed. For instance, Bomb-making instructions on the Internet was (very obviously) vandalized in Special:Diff/1358132635, and the vandalism got reverted by ClueBot basically immediately. The suggestion nevertheless refers to the vandalized version (Remove the non‑encyclopedic, opinionated rant). I don't know how the LLM got hold of a revision that existed for only a few seconds.

Suggestions to conceal the symptoms of a larger problem:

  • As seen above in the bomb-making example, the LLM doesn't seem to know about vandalism and will never describe it as such, regardless of how obvious it is. This seems likely if not certain to encourage someone to "fix" the tone of vandalism without addressing the actual claim, which is how we get years-long hoaxes.
  • One thing that happens fairly frequently is that the suggestion feature will flag an article that, in context, is clearly AI-generated. (Examples: Jeremy Coller, Vimbuza). It never picks up on this and suggests minor tweaks to wording that would just put a band-aid on the issue. (The csv is actually fairly useful for this, but only in full searchable list form.)

LLM-specific tics: Obviously the suggestion text itself are AISIGNS overload but what I mean here is that LLM edit suggestions/summaries have some consistent quirks. I haven't done an in-depth look for them but two known ones do show up frequently:

  • For some reason LLMs have a fixation on "superlatives" being inherently non-neutral, when sometimes they are just true. I don't know where this comes from -- WP:NPOV doesn't mention anything about superlatives -- but it shows up all the time in LLM-generated revision suggestions, and also all the time here. For instance, in List of Hercules: The Legendary Journeys and Xena: Warrior Princess characters: The description of Hercules is too lengthy and includes biased phrasing such as ""strongest man in the world"" [...] Or on Trinidad, California, a sentence starting "On December 31, 1914, the largest recorded ocean wave ever to hit the United States West Coast" (in a paragraph with 5 citations) is criticized with The sentence makes an unqualified superlative claim about the wave’s size, which could be seen as non‑neutral. Adding an attribution phrase mitigates this.
  • LLMs will over-justify anything. The following reads like a parody but it is actually a suggestion from these: The word 'ithe' is a typo; it should be 'the'. Fixing this error improves the sentence's correctness.

I don't really know what to say at this point. Based on the convoluted phabricator issue snarl it seems that there was some human review done at some point, but nevertheless it took me only about ~1-2 hours to find the above issues, and that was only a spot check. This seems like a reasonable amount of human QA to expect before pushing an editorial feature to prod. It also seems reasonable to expect a full audit of potentially controversial subject matter (by "full audit" here I mean even just the results of CTRL-F "Israel", "Palestine", etc.), and ideally a review by people who are familiar with AI-generated revision suggestions and the general areas in which they get things wrong. (None of this deviates much from the thousands of similar justifications I've seen in AI-generated edit summaries.)

I also think the whole premise is just flawed. LLM and machine learning tools are not only going to have high rates of false positives, but false positives that take time to evaluate. (For instance, there are various LLM tools to scan articles/edit summaries for possible issues, but the point is that they generate lists to be manually reviewed later.) They do not scale to a scenario like automated edit suggestions, where the assumption is that the suggestions are pre-vetted and can be evaluated quickly. Gnomingstuff (talk) 17:54, 15 August 2026 (UTC)

@Gnomingstuff I get your frustration, and agree with some of your points, but I still think this is worth a try. I will often ask the LLM-bots to proofread my articles. Some of the suggestions are obviously good (mostly low-level stuff like spelling, repeated words, etc). That's the kind of stuff that once you've read your own writing 100 times, you read right past and don't notice, so I find it an invaluable service. The higher-level suggestions (tone, phrasing, flow) I'm much more likely to reject, but I accept them often enough that it's worth doing. But that's really no different from when I'm working with a human reviewer. I'll often push back and say "Nah, I think the way I've got it now is fine".
So I think the trick here is to figure out how to educate people that these really are just suggestions and they need to apply their human judgement about whether to accept them or not. You are correct that for new editors, that may be problematic. Still, I think this is something worth trying as long as we monitor how well it's working out and be willing to pull the plug if it turns out to not be useful.
LLMs are just the most recent technology step between scribes writing on clay tablets and where we are today. We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" Or we can experiment with them (inevitably with some failures) and learn how to take the best advantage of them to improve our product. I vote for the latter. RoySmith (talk) 18:17, 15 August 2026 (UTC)
Please don't ping me to a discussion that I started less than an hour ago and am clearly aware of.
I don't think that We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" is a fair assessment of something that I spent actual time looking into. Gnomingstuff (talk) 18:30, 15 August 2026 (UTC)
The question is whether this will help new editors to develop good judgement. LittlePuppers (talk) 04:55, 16 August 2026 (UTC)
Please just throw me in a ditch. Polygnotus (talk) 18:48, 15 August 2026 (UTC)
Per WP:TALK and the header of this talk page, please avoid fact-free rants and aggressive exclamations, and try to contribute actual arguments instead. Regards, HaeB (talk) 07:01, 16 August 2026 (UTC)
Fixing this error improves the comment's correctness.  — Hex • talk 15:46, 16 August 2026 (UTC)
The above exclamation was not nearly aggressive enough. Let me rephrase for clarity: This is one of the worst features I have ever seen proposed for anything, ever. Enabling a feature like this is literally insane. –jacobolus (t) 20:10, 29 August 2026 (UTC)
At the very least this specific feature will be community configurable via MediaWiki:Editcheck-config.json/Special:EditChecks once it goes live so we can turn it off. But yeah, my first reaction is the same as Polygnotus'. * Pppery * (alt) in solidarity 18:58, 15 August 2026 (UTC)
Kill this with fire, and fire whoever wanted to impose this upon us. Dishraceful and going against clearly expressed community sentiment. The Wmf should not produce any tools that make content suggestions ever, this is not what they zxist for. Fram (talk) 20:08, 15 August 2026 (UTC)
Kill it with fire. The only good thing about this is that it appears we have the ability to turn it off. Tazerdadog (talk) 20:54, 15 August 2026 (UTC)
I fear the suggested edit feature has shown that new editors have a bad tendency to follow suggestions blindly. This isn't the fault of new editors, but poor explanations of what is being suggested and that they are only suggestions.
Looking at the example above make me think this will only make the situation worse, especially as the LLM shows that it doesn't understand policy (a common problem for ever LLM). -- LCU ActivelyDisinterested «@» °∆t° 20:59, 15 August 2026 (UTC)
@ActivelyDisinterested, @Kowal2701 and Gnomingstuff, and others in this thread, you are looking at a feature in it's pre-pre-pre-alpha stage. What the ticket tells you is that the Editing team is preparing to deploy a very very early experimental version of the feature to experienced editors who have opted into enabling a suggestion mode beta and append a specific parameter to the URL (i.e. basically nobody will get this feature unless they specifically click a link and have a very specific beta preference enabled). The code is being enabled so that it can be demoed to folks, used to perform rudimentary qualitative A/B tests and gain very preliminary feedback from Wikipedians at conferences (which occurs before wider consultations with the community). This is nowhere close to being deployed anytime soon without significant bug fixes and the call-to-action in the thread, "is due to be announced any time now" is just patently false. For what it's worth, I personally haven't made my mind up about this specific feature, but I'm willing (and would strongly urge other folks) to provide the team with the ability to spend some more time atleast trying to iterate and experiment on the feature to see if some variation of it could be made useful to some Wikipedians. Sohom (talk) 22:50, 15 August 2026 (UTC)
I realize that this is an experimental version of the feature, but based on the actual content that exists, this isn't the "show to editors and assume they like it" stage, it's the "internal minimum-viable-product demo" stage -- and a MVP you'd need to very carefully babysit to make sure something like The current content is a series of JSON objects that do not convey readable information to the reader. doesn't pop up onscreen. Arguably it's not even that, but the stage of "go back to square one and rethink because the premise is inherently flawed."
I think that "due to be announced any time now" is a fair interpretation of there being an August ticket called "Announce availability of 'experimental' suggestions" with the description The announcement we publish ought to equip volunteers with the info. they need to answer the following questions.... Like... it's due to be announced. That's... what the ticket... says.....
  • "Assume" is not my wording, it's directly from the ticket: In T428311 and T431376, we – staff, in collaboration with experienced volunteers across a range of Wikipedias – will have assumedly determined the initial batch of LLM-generated MoS suggestions to be reliable.
Gnomingstuff (talk) 06:46, 16 August 2026 (UTC)
Let's nip this in the bud please before it becomes a fait accompli (if it hasn't already). Incredible that there's been no community consultation about this AFAICT Kowal2701 (talk, contribs) 22:35, 15 August 2026 (UTC)
I'm somewhat interested in what an llm could dig up in a widespread analysis of MOS:GEO, but unfortunately the suggestions for MOS:GEO are not really about MOS:GEO but are normal typos and (misunderstood) context suggestions. This may be in pre-alpha, but it is frustrating to read "we’ve been successful in developing bespoke, one-off models that surface specific kinds of editing suggestions in a reliable way. For example, we use the Add-a-Link model to suggest relevant inline links between articles" when there have been deep flaws to the add-a-link model that have been unaddressed since its implementation. It is also known that the revert metric used is flawed, so it is disappointing to see it still being referred to. (I recently provided an example to WMF devs of a tone check edit making the article more promotional, but I don't know if that's an edge case or a more widespread issue like add-a-link has.) The "tools that show promise" user story is also quite cheeky; I can't decide where that lies on the amusement to annoying scale, it could be seen as endearing.
On the current suggestions, there is a mix of "valid and useful" and quite wrong. The valid and useful ones I saw were mostly typo suggestions. The MOS:GEO ones that went beyond that were sometimes nonsensical. The NPOV ones I checked I would avoid suggesting. The metric being used to assess readiness, in this case "The size of this vetted set of suggestions has given the team the confidence", needs to be relooked at. A consideration that seems to be lacking from all these suggestion ideas is that putting any of these into a formal structure gives them an imprimatur of authority. That's tricky to work around, but the "accelerate the speed with which we can surface meaningful signals" language suggests it isn't a strong consideration. "low-risk edit suggestions (as defined above)" is another metric that needs to be reassessed, if you're trying to touch upon NPOV you have left low-risk behind. If it is true that "it can take more than a year to produce a single type of suggestion", it does seem like far too much effort given the quality of the results. I hope the experimental team will have another think about the fundamental assumptions here and the metrics used for assessment, to help shape future development. CMD (talk) 23:42, 15 August 2026 (UTC)
The idea itself is interesting, and I wouldn't be against experimenting with AI as a way to surface article quality issues (e.g., what EditCheck is doing). This seems to go much further, with the model presenting specific, actionable suggestions, although it stops short of ready-to-post edits.
Is this still experimental? Absolutely. As @Sohom Datta points out, this is about to be deployed as "experimental suggestions" open for further feedback, to a testing audience clearly distinct from its target audience. I doubt the kind of editor knowledgeable about beta features and actively seeking out to test this will be misled by the AI's suggestions.
I will concede that there is an ambiguity in the way the ticket presents the matter, which is not ideal: the purpose of these edit suggestions is to both edit more effectively with tools that show promise and contribute to making those suggestions more reliable by using and evaluating them in real editing contexts. This, while technically true, should be clarified to shift the emphasis towards the latter.
Now, what gives? Of course, no one wants the current version to be shown to newcomers, given the major flaws pointed by Gnomingstuff and others above. What we can do, however, is twofold. Now, discuss whether this feature could be developed to provide constructive help in theory, not considering its current lack of readiness. This is an open question. And later, once the feature is sufficiently mature and ready for a rollout to newcomers, discuss whether that current state is worth rolling out, whether it needs further development, or if the project should be cut short. Chaotic Enby (in solidarity · talk · contribs) 23:42, 15 August 2026 (UTC)
But we've done this so many times where features seem to never get cut short after a certain point in development regardless of feedback, it's only the rare occasion when enough people kick and scream Kowal2701 (talk, contribs) 00:29, 16 August 2026 (UTC)
Agree, and this is why the sunk cost fallacy has to be considered. In fact, I can see the opposite outcome from this initial discussion, namely the developers being left with the impression that all the issues the community has are with the current state of the project, and that investing more will make it worthwhile and gain community acceptance. This is in fact far from obvious, and why we should, I believe, center the discussion on the viability of the project as a whole. Chaotic Enby (in solidarity · talk · contribs) 00:40, 16 August 2026 (UTC)
It's just very difficult to have a constructive discussion/approach when most are worried/anxious they're going to end up ignored and powerless Kowal2701 (talk, contribs) 01:14, 16 August 2026 (UTC)
Pointing down to Peter's comment below, but basically: these specific suggestions were done in a way to try to avoid sunk costs. The development effort on Editing's side has been quite low (gerrit:1299651 + gerrit:1320209) and has mostly been about setting up a framework for a suggestion-type that can ask an API to hand us fairly arbitrary suggestions on an article and then for us to log feedback from a user about whether they seem valid. The broad idea was to make them available to people who knew how to toggle past multiple layers of "are you sure? this is a beta / experimental", and gather data about which specific instances were considered helpful and which weren't. (Screenshot below as well, but we're also not showing the content-specific suggestions you can see in the raw data file; we just show the static_description field for each one, so you just get the "this might violate MOS:GEO" level of prompting about it...) DLynch (WMF) (talk) 13:30, 16 August 2026 (UTC)
I think the first wave of responses show useful feedback on what instances are helpful, what aren't, and what need extensive tuning to filter out tricky cases and focus on the most useful / confident / low-risk suggestions. Glad that NPOV is being dropped; that's extremely complex and contextual.
A good recurring issue that won't show up in spot checks of individual suggestions, is that articles overall deserve article-level checks before spending time fixing small details. @Gnomingstuff put this well above. I would be interested to see a version that allows article-level suggestions (e.g., for tags that might apply to entire articles or sections), especially for new or single-author articles. – SJ + 22:53, 18 August 2026 (UTC)
I am curious what you mean by whether this feature could be developed to provide constructive help in theory. My experience as an engineer is that these sorts of discussions are not productive. You simply make things, you learn a lot along the way, some of the stuff you scrap, other stuff ends up being very useful, but not towards the goal you originally set out to achieve, and on occasion you actually end up producing what you originally set out to do. Discussions about what sort of engineering efforts would theoretically produce good products are simply a waste of time. Czarking0 (talk) 15:52, 16 August 2026 (UTC)
Hey all -- I'm Marshall Miller, director of product at WMF (this project is with the teams that I work with). I'm commenting to let you know that we see this and that members of the Editing team will be able to comment with more detail, background, and clarifications this week.
But yes, let me first say that this is the very earliest stage of testing/trying/experimenting with this idea, and just for experienced editors. As we have done for all the edit check and suggestion features so far, we will only advance this feature farther if communities are supportive, if the suggestions are reliable, and if the data shows that they make a positive difference for the wiki. And all these checks and suggestions are configurable by communities at Special:EditChecks.
About why we're pursuing this: we have seen good success and community support with edit checks and edit suggestions, and the idea of suggestion mode. So far, these have run off of either simple logic ("this blob of text was pasted from ChatGPT") or small machine learning models ("this sentence uses peacock words"). What they all essentially do is point out to human editors when they are violating a wiki's policies in some way -- and they have been shown to reduce revert rates and make newcomers more successful. But there are a lot of wiki policies that could be useful to point out to people. LLMs are a new technology, and they can and do make mistakes. It takes careful testing and tuning and evaluation to get them to perform reliably and may not even always work out (and it may not in this case either). But they may make it possible to produce more of these useful checks that help newcomers make better edits, and may help experienced editors notice things that need fixing. We will 100% need the input of all of you to help us figure out together whether we're on to something or not.
Our approach to all this is that editing decisions should be made by humans (except for the very simple kinds done by things like ClueBot, etc). And that features like these try to help humans notice/find places where they could apply their judgment.
Okay, anyway -- more to come from team members who are deeper in the details. MMiller (WMF) (talk) 04:03, 16 August 2026 (UTC)
WMF A/B tests are notoriously unreliable, and invariably interpreted in the most positive light possible. See e.g the image viewer disaster, or the initial claims about edit check where the posted positive results turned out to be false. Why should we trust whatever results will be posted this time? More importantly, why is such a tool created when the WMF should know by now the massive pushback they would get against AI content suggestions? Aren´t there enough other improvements requested (often for many years already?). Fram (talk) 06:51, 16 August 2026 (UTC)
So far EditCheck has produced amazing results (e.g. vastly increased the share of newcomer edits that contain citations) – and the Editing team listened a lot to community members while developing new checks/suggestions. But you don’t have to trust anyone given that all suggestions and edit checks can be enabled and disabled by local admins. Johannnes89 (talk) 16:56, 16 August 2026 (UTC)
Yeah, I think I meant referencecheck (or whatever it is called), not editchecks (which I haven't checked), should have been more careful in what I said. They claimed a serious number f added references, but it turned out that a lot of edits were tagged as "reference added" when this wasn't true, and a lot of other "references" were completely invalid but counted as a success anyway. I posted this with clear examples, but the WMF ignored this completely. But that's about reference adder AB tests, not edit check, so again, I should have checked before posting. Fram (talk) 09:41, 17 August 2026 (UTC)
From glancing at the list linked above (totaling 349 suggestions - 22 labeled MOS:GEO, 116 labeled NPOV, and 211 labeled "simplify language"), here are my thoughts:
MOS:GEO - most seem fine at a glance. Lots of suggestions to fix diacritics and spelling. Several which are well outside the purview of MOS:GEO. Seems lacking in nuance in some edge cases (but saying more confidently would require some fact-checking).
NPOV - lots of issues. It is too timid to say anything forceful, even when it's warrented and supported in RS, and wants to add qualifiers (e.g. "reportedly") for simple statements of fact, such as "improved quality of air" or "top of the chart". It seems generally opposed to any words which are not incrediby boring, and even some which are: perilous, successful, unreliable, transparent, vocal critic. It is often very unclear in what it is referring to ("the evaluative phrase", "the promotional claim", "subjective description", "the evaluative language"). Multiple suggestions are to remove language which is not there. It also has a terrible time recognizing attribution, and suggests several times (probably a dozen+) that it be added when it's already there. I've skimmed through maybe half of these, and a majority have issues.
Simplify language - meh. Some are fine. It seems to want incrediby short sentences. Ironically, I also disagree with the one I see where it suggests combining sentences. Most suggestions are pretty vauge. One suggestion is "make this neutral". Also says "American English is preferred on Wikipedia" on a British biography.
Summary of my views: GEO is mostly decent but may lack nuance, NPOV isn't really useful because it lacks the understanding to know when a strong viewpoint is neutral and generally dislikes big words, and simplify language is overzealous in suggesting short sentences. LittlePuppers (talk) 06:16, 16 August 2026 (UTC)
Okay, I was looking at a different file from Gnomingstuff and one which is at least a few weeks old. Take that how you will. LittlePuppers (talk) 06:19, 16 August 2026 (UTC)
Glancing through what is (I think) the latest (and much longer) version, there may be some improvement, but most of my thoughts still apply. LittlePuppers (talk) 06:29, 16 August 2026 (UTC)
The MOS:GEO ones are often not fine. For a start, being outside the purview of MOS:GEO suggests some underlying flaw in the model. "Update the country name to conform with Wikipedia’s geographical naming conventions", I have no idea what that is meant to refer to. "The sentence is amended to specify that the Grand Canal is in Venice, providing clearer geographic information" lacks understanding that the Venice location was established in the prior sentence. "The parenthetical abbreviation after “Guantanamo Bay detention camp” is incorrect and should be removed" is simply wrong, although it is perhaps an unnecessary abbreviation a reader may also not understand. "The place name "South Island of New Zealand" is not formatted according to MOS:GEO; it should use commas between the island and the country" is again just wrong. "The original text mentions "the Atlantic" without specifying that it refers to the Atlantic Ocean, which may cause confusion", not sure what to say about that one, Atlantic Ocean is even written out explicitly earlier on the page. CMD (talk) 06:35, 16 August 2026 (UTC)
Yeah, the set I was looking at initially had a very limited list for GEO. A lot of what you mention reflects broader issues with all the categories as well. LittlePuppers (talk) 06:56, 16 August 2026 (UTC)
Most of these are known issues with LLM-suggested edits, or at least the kind of thing that has certainly been possible to know about for at least a year:
  • The "promotional claim"/"evaluative language" stuff is a 2025-era LLM tic. Very specific verbiage, especially the "evaluative" part, that shows up over and over again in AI edit suggestions and basically nowhere else. Here's a bunch of examples.
  • The "original text mentions the Atlantic" suggestion is another consequence of the isolated-sentences approach; the LLM is responding to the sentence "According to Herodotus they dwelt geographically along the sea south of Libya on the Atlantic," and so it doesn't have the context of first reference/subsequent reference.
Gnomingstuff (talk) 06:56, 16 August 2026 (UTC)
Further context or not, saying "the Atlantic" is not going to cause confusion. CMD (talk) 07:06, 16 August 2026 (UTC)
"American English is preferred on Wikipedia", great so it's not just wrong but will make a bad situation worse. -- LCU ActivelyDisinterested «@» °∆t° 09:07, 16 August 2026 (UTC)
The issue with language suggestions in regard to NPOV is that LLM are not neutral, and do not give neutral suggestions. So using them to make these kind of suggestions is a way of creating a fake consensus. -- LCU ActivelyDisinterested «@» °∆t° 09:11, 16 August 2026 (UTC)
Thanks Gnomingstuff for this thorough demonstration of why LLMs are systems for producing text-like slop. No matter how many attempts are made to patch these behaviors, they will keep happening because no comprehension or intelligence is involved, and never will be. Just guess after guess, a fountain of hot slop staining our precious reputation as one of the few uncontaminated places online.
This shameful and embarrassing effort needs to be canceled immediately and the donation money wasted on it so far written off. I'm not even going to start getting into the unethical nature of using LLMs in the first place, which should have been sufficient on its own to rule out even considering something like this.  — Hex • talk 15:58, 16 August 2026 (UTC)
The sheer quantity of bullshit from the WMF is exhausting at this point. Cremastra (talk · contribs) 04:41, 17 August 2026 (UTC)
This is yet another reason to not trust the WMF and it shows how it's impossible to assume good faith on their part. The WMF at this point is an active threat to the very existence of Wikipedia. Ita140188 (talk) 07:58, 17 August 2026 (UTC)
Dealing with WMF is like living through Groundhog Day. They have too many employees so they bureaucratically create "jobs" building crap that nobody asked to solve "problems" that don't really exist, creating a bigger set of unforseen consequences (because WMF is composed of many software engineers and few Wikipedians and is always and forever tone-deaf to community desires). The volunteers who make the project run are all "power users" to them... Well, here's what the "power users" are saying, "tech bros"...... NO AI ON WIKIPEDIA. Didja get that? Carrite (talk) 14:57, 17 August 2026 (UTC)
Wikipedia has a reputation as one of the last bastions of information, in an era of hallucinated, enshittified LLM-generated slop. Any embrace of AI-powered anything on the platform constitutes a plan to throw that into the bin. ser! (chat to me - see my edits) 11:51, 18 August 2026 (UTC)
HELL NO. AI slop is already a massive problem on Wikipedia and many editors have to waste their precious time getting rid of it. The literal WMF coming in and force-feeding AI slop into the mouths of editors is, to put mildly, disgraceful and shameful. AI is not reliable for anything, especially for use on a perceived beacon of human, somewhat trustworthy (even if it really isn't) content like Wikipedia. We should never have even thought of such a horrible decision, let alone create a bare-bones prototype of it and make it accessible to editors. 🪐Kepler-1229b | talk | contribs🪐 17:17, 13 September 2026 (UTC)
@Kepler-1229b, You should read the rest of the discussion? The literal WMF coming in and force-feeding AI slop into the mouths of editors is, to put mildly, disgraceful and shameful. is not what is happening here based on reading the rest of the discussion. Sohom (talk) 17:21, 13 September 2026 (UTC)
It's a possible outcome. I know that it's supposed to be opt-in right now, but it could very well be possible later on. We need to nip this in the bud before it gets there. 🪐Kepler-1229b | talk | contribs🪐 17:37, 13 September 2026 (UTC)
Any installation on en-WP without consensus from the community sure feels that way, because we'd have to deal with the results even if we choose not to utilize it ourselves. We're already having that problem with the AI-based Revise Tone newcomer tasks. The rest of us have to come through and clean up after the slop. ChompyTheGogoat [ Bleat | Munched ] 01:40, 14 September 2026 (UTC)
I oppose the use of LLMs to write any article content (with exceptions for translations, grammar, adjustments to wiki syntax, or the like; but never for writing any original text), but I think they can be far more useful for the information search part. MGeog2022 (talk) 17:55, 18 September 2026 (UTC)
Which is already fully allowed, because you don't need internal Wikipedia tools to search Wikipedia. Everything is accessible to external programs. Given how difficult navigation is here I do utilize it occasionally to help me find guidelines, as well as sources. ChompyTheGogoat [ Bleat | Munched ] 20:27, 18 September 2026 (UTC)
Oh god no, I hope this can still be stopped. This is just a vector for introducing misinformation to wikipedia articles. TietoTeekkari (talk) 19:43, 20 September 2026 (UTC)

Taking a step back: could AI suggestions be beneficial?

As pointed out above, the current development stage is way too early for a broad rollout. This should have been better clarified, both to reassure the community about the experiment, and to provide clearer development goals.

However, taking a step back, a discussion can still be held regarding the potential of this whole endeavor. Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way, assuming the current pitfalls could be smoothed out in development? More concretely, are these expectations realistic, and is it worth investing further resources in this project?

These are questions I don't, personally, hold the answers to. However, we should be discussing them together, alongside members of the Editing team involved in its development (courtesy ping to @Quiddity (WMF)), if we want them to work in sync with community sentiment, and avoid investing resources in dead ends. Chaotic Enby (in solidarity · talk · contribs) 23:57, 15 August 2026 (UTC)

Would the community, in theory, agree to an AI model surfacing suggestions to newcomers. I think a better way to explore this tool would be to make it available to established editors first. People who have the experience and policy knowledge to be able to properly evaluate the suggestions. Maybe the people will say "The suggestions were all spot-on and incredibly valuable". Maybe they will say "Nothing this thing suggested made any sense at all, it's total garbage". More likely, somewhere in between. But let's do the experiment rather than pre-judging it. RoySmith (talk) 00:08, 16 August 2026 (UTC)
I think a better way to explore this tool would be to make it available to established editors first. While it might not have been clear at first, this is, in fact, exactly what the experiment is planning to do. The tool is still in development, and we can't say, in advance, how it will end up in terms of quality. For now, I'm just trying to figure out the proportion of editors who either will find it a non-starter in principle (regardless of the suggestion quality) or are opposed to investing further resources in its development for any other reason. Chaotic Enby (in solidarity · talk · contribs) 00:20, 16 August 2026 (UTC)
Mark me down as "opposed to investing further resources in its development for any other reason". Wikipedia has a large amount of technical debt that would be easy for a WMF developer to fix, but instead they are doing this?
As one example, Arbcom is a very important function, and many people who participate get stressed over whether they are under the word count limit. But the tool that puts a banner at the top of your comment fails to accurately count your words! Worse, you can't invoke it while composing -- you have to post and hope that you didn't go over. And if an arb replies to you inline, that increases your count! This is the sort of thing that a WMF developer could fix in an afternoon. It's important, but making an accurate arbcom word counter will never hit the top of any survey of things many editors want to see fixed.
Another example: what happens if when I sign this comment I accidentally hit the "~" key three times instead of four? How about 5, 6 or 7? This is a typo that happens again and again. How hard would it be for a WMF developer to make it so that you get a "are you sure" message before accepting a signature that is almost always a typo?
There are hundreds and hundreds of these easy to fix things that depend on old scripts written by volunteers and all too often no longer maintained.
I think the WMF should spend a significant amount of developer effort -- 75% or 80% -- fixing these small, non-sexy quality of life issues and only then devote the other 20% - 25% to fun things like AI suggestions. --Guy Macon (talk) 01:02, 16 August 2026 (UTC)
Or, if you don't like the above 4-tilde signature, --Guy Macon (talk) (3 tildes),  --01:13, 16 August 2026 (UTC) (5 tildes), or --01:13, 16 August 2026 (UTC)Guy Macon (talk) (7 tildes).
Guy, on that last one, you don't need to add a signature at all anymore. Discussiontools will do it for you. I haven't added one to the end of this post, for example. In solidarity, asilvering (talk) 01:04, 16 August 2026 (UTC)
Will I (Guy Macon) get an error message for this unsigned post or do I have to make a preferences change to get the autosign goodness? Show preview says it will post the unsigned comment with no error message.
Discussiontools is a nice tool, but it is no substitute for software baked into Wikipedia that checks for common errors and throws up an "Are you sure?" message. I could give you a hundred examples of places where the WMF is depending on unpaid volunteers to maintain basic functions that keep the Encyclopedia running smoothly while focusing on the exciting new stuff --01:30, 16 August 2026 (UTC) Guy Macon (talk)
Guy would need to be using DiscussionTools for that, but unfortunately they are using classic full page source editing where none of those helpful things will occur. DLynch (talk) 13:13, 16 August 2026 (UTC)
FYI I made Module:Word count, I can refine this further if you think it would be useful. It does not 100% solve the issue you mention Czarking0 (talk) 15:57, 16 August 2026 (UTC)
Although I agree with a lot of this sentiment, I think an organizational strategy which places say 20% of the engineering resources on long term tech rather than present issues is reasonable. As long as the development of this feature counts under that I do not see the problem. Czarking0 (talk) 16:00, 16 August 2026 (UTC)
+1, that seems to be what is happening here (see the comment below about 'avoiding sunk costs' and getting feedback early and often from established editors). I can see individual categories of suggestion being useful to experienced editors, focus on making something that works for them before considering anything that might be visible to newcomers. (They already have Special:Homepage) – SJ + 23:14, 18 August 2026 (UTC)
I agree with Fram above in that this feels like a way (though a much more subtle way than Simple Summaries) for the WMF to influence editorial decisions which, with the exceptions of legal reasons, they just should not be a part of. ♠JCW555 (talk)♠ 01:55, 16 August 2026 (UTC)
There isn't really any plans for the WMF to get involved in editorial descision (and I say that as somebody who has through m:PTAC reviewed the Annual Plan). The plan for Edit Suggestions is purely meant as a assistive tool to help editors and kinda comes from the idea of being able to surface gadget/userscript suggestions to everyone without having to have coding knowledge. Sohom (talk) 02:57, 16 August 2026 (UTC)
But the fact that the AI is making these suggestions at all in the first place is the WMF having a subtle hand on editorial decisions in my mind. The examples Gnomingstuff lists above, like the FreedomWorks example, is an example of the AI making an editorial judgement that it should not be doing. Some of the others are more subtle like the DeAndrea G. Benjamin example, but they're still editorial decisions that the WMF shouldn't be engaging in period. ♠JCW555 (talk)♠ 03:23, 16 August 2026 (UTC)
They are using generic models without fine tuning for these suggestions, so out of all the parties that could be said to have made editorial decisions in this scenario I would say OpenAI and Google would rank above the WMF, and I really wouldn't consider either of those companies to have made any editorial decisions... the problem IMO is potentially encouraging uh... not really making any editorial decisions, and vibing through things without considering the context of, e.g. the contents of the sources as we are supposed to. Alpha3031 (t • c) 08:28, 16 August 2026 (UTC)
It is not worth creating AI prose suggestions for newcomers (especially if it takes a year for each model!). Putting aside quality questions, editors here are expected to be competent and able to contribute in English. While there are different ways to be confident, we expect the ability to read and write in English, and thus to some extent to have the tools to be able to copyedit themselves. We also expect editors to be able to read and comprehend our guidelines and policies (pre-emptive clarification, read, not memorise). The best way for us to be able to understand competence in these areas, and thus to be able to assess and offer advice if needed, is to see their edits. Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors, and harmful to the new editors in giving them a misleading picture of how things work, and more harmful when the llm doesn't understand our policies and practices (as this one does not). Apologies to Sohom, but "There isn't really any plans for the WMF to get involved in editorial descision" just isn't true if this sort of system is being set up. This sort of suggestion task is the WMF making editorial decisions, even if they're making it through an llm. There are many reasons time would be better spent anywhere else. (I distinguish "prose suggestions" from say the add-a-link task, as that at least teaches a technical competence that we do not expect new editors to have. Such teaching does seem useful, although as mentioned above it would be nice if it was developed further.) CMD (talk) 03:45, 16 August 2026 (UTC)
Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors I just want to emphasize this -- the English Wikipedia community is not good at welcoming newbies who make mistakes. That is a problem.
The English Wikipedia community is actively hostile to editors making mistakes with large language models.
If a newbie puts a poor LLM-based reword into an article, an experienced editor is almost certainly going to WP:BITE them off. And a newbie, operating in a system they're unfamiliar with, with a human-sounding voice telling them "This copyedit is right", is never going to have the knowledge and is almost certainly not going to have the confidence to challenge the AI when it presents them with a bad suggestion. That is going to set them up for failure when a grumpy human editor, burnt out from dealing with LLM edits, callously reverts them.
I think using machine learning to help with encyclopedia maintenance is a wonderful thing! And I think large language models are really cool -- but they require a high degree of skill to use correctly in the manner that looks like it's being explored here. And, again, the community is so burnt out from dealing with poor quality LLM content that any editor who uses this tool and (inevitably) makes a mistake will be attacked by the community. Does the team working on this understand that, @MMiller (WMF)? GreenLipstickLesbian💌🧸 05:10, 16 August 2026 (UTC)
Seconding that this could be helpful for all sorts of maintenance but should be aimed at experienced editors while working out kinks.
I would personally like to see rubrics for highlighting potential vandalism and LLM edits. And I want much less text in my sidebar and zero suggested text: just a few words indicating the kind of issue to look for / the kind of style guidelines to check.
And I'd like to see explicit self-evals of each rubric for suitability to the task and for false positive/negative rates, which could also be compiled and developed by community maintainers. – SJ + 23:52, 18 August 2026 (UTC)
@GreenLipstickLesbian -- I think that the most important thing that would prevent against the situation you're describing is that the suggestions wouldn't propose text for the newbie to accept/reject. It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?" -- it would not give them a re-written sentence to add. I know that in that situation, the AI may be wrong about whether the sentence needs to be rewritten, and the sentence may be perfectly fine -- but the newbie might be like, "Hmmm, well I guess I'll reword it?" and they may make a pointless edit, or may make the sentence worse.
Suggestion tasks like "add a link" and "add an image" actually do give newbies specific edits to accept or reject, and we see them generally apply good judgment and be constructive. Yes, many of them mess up and get reverted, but that might have happened to them anyway if they were left to their own devices. So these tasks cause a bunch of things to happen at once, and the question is whether it all adds up to a net positive or net negative, you know?
What do you think? MMiller (WMF) (talk) 21:13, 19 August 2026 (UTC)
@MMiller (WMF) Thanks for the response! Yes, I think these sound like reasonable precautions. I do really want to emphasize the point about making it clear to experienced editors what the newbies actually are seeing is very important.
Having had the beta experimental version of suggested edits enabled for a few days, I definitely see the vision. And I see how it could be very useful. (My response to many of the "consider adding a citation" suggestions has, admittedly, been a)"lol no, I'm removing the unsourced text for other PAG issues", b)"this is cited, it just needs an inline citation", c) "... yes, i see how citogenesis is made", or, d)"... we need an easily accessible essay on 'how to source a statement' that we can link to the newbies here" because wow, i'm having trouble". ) But I also see the biting that happens when newbies do make pointless edits/less than ideal edits ( has some which took more than a few years to rectify), and so I am worried about adding any "they're just thoughtlessly adding machine suggestions to articles"-esque ammunition, even it's it's not strictly true. GreenLipstickLesbian💌🧸 21:33, 19 August 2026 (UTC)
This is why I suggested some basic FAQ type popups upon first edit, including an agreement to not use LLMs, to prevent true good faith mistakes and remove plausible deniability for the rest. ChompyTheGogoat (talk) 17:37, 26 August 2026 (UTC)
RE: especially if it takes a year for each model! to be fair to the teams involved the stated idea behind this specific experiment seems to be wanting to try something that won't take a year per task on what the editing and ML teams think are relatively low risk tasks... I'm just not sure that the community and the teams involved have a sufficiently compatible idea of what tasks are low risk. I think it would be possible to develop something acceptable to the community, but unless very sure about it, it may be best to get a vibe check from the community before thinking something is "low risk", and treating things as "high risk" otherwise. Alpha3031 (t • c) 08:36, 16 August 2026 (UTC)
Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way – I won't, and if that ever happens I'm gone. Models are bias black boxes, and even if suggestions were 100% accurate this would still be an issue. fifteen thousand two hundred twenty four (talk) 04:37, 16 August 2026 (UTC)
If AI had far more quality control, I wouldn’t be opposed to this. Except it doesn’t yet. LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician.
This could probably work in theory. But it would assume an AI with a reasonable knowledge of the given article subject, some common sense, and a degree of fluency in English (by which I mean not doing things like pyusician/musician), and, given the above examples provided by Gnomingstuff, this is definitively not that.
I oppose AI-based features being integrated into Wikipedia software, and will continue to do so unless a day comes when LLMs equal or surpass the common sense of a human. Perhaps in a few years LLMs will be reliable enough for this to work smoothly and without issue. With the current state of AI, I don’t think we’re there yet. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 04:59, 16 August 2026 (UTC)
The AI could work 100% of the time and I'd still oppose it because these edit suggestions are being directed by an AI that's controlled by the WMF, which outside of legal purposes, should never have any editorial control on Wikipedia in the first place. ♠JCW555 (talk)♠ 05:17, 16 August 2026 (UTC)
That’s a fair point; I hadn’t thought of that. (Yet another reason why this is a terrible idea.) Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 05:39, 16 August 2026 (UTC)
The solution for this is making prompts publicly available and allowing each community to tweak them. Alaexis¿question? 06:36, 16 August 2026 (UTC)
LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician. - that's a rather odd example to get hung up on, given that automated spell checkers and their admittedly sometimes absurd correction suggestions have been around for decades (long before the term "hallucinations" came to be used in this context). Many or even most professional writers - journalists, book authors, scholars - use them routinely, and have learned to live with the occasional fail of this type (pyusician —> musician).
Indeed, in stark contrast to your logic here, WP:SPELLCHECK has long described them as potentially useful for Wikipedia editors, too, as long as they don't blindly rely on such tools:

Spellchecking software and online tools can be helpful when copyediting Wikipedia articles. [...]

  • No spellchecker is completely accurate. You must check the output of any tool you use. [...]
  • You are responsible for all spelling and grammar changes you make, even if the changes are suggested by an error-checking tool, such as Grammarly or ChatGPT.
Maybe it is time to conceive of Wikipedia editors a bit more as adults who are generally capable enough of using such tools even though their suggestions are not 100% error-free (very few things are).
That said, I do generally agree with your point that AI needs quality control, and folks should definitely ask if WMF has done enough here yet (I'm not sure it has). It's just that - as the spell checker example shows - it's not realistic to demand 100.000% accuracy, or to point to isolated failure cases without assessing how frequent they are. (To be fair, User:Gnomingstuff did already make an informal heuristical effort at the latter with regard to their list above: it took me only about ~1-2 hours to find the above issues, and that was only a spot check, i.e. there is informal evidence that these are not very rare at this point. But ultimately I'd be interested in more concrete assessments of how likely editors using this tool will be to encounter each of those failure cases in practice.)
Regards, HaeB (talk) 06:54, 16 August 2026 (UTC)
Most adults have a lot more experience with spelling than they do with Wikipedia's guidelines. LittlePuppers (talk) 06:58, 16 August 2026 (UTC)
Sure, but that doesn't mean that we keep the edit button away from them.
See Wikipedia:Competence is required, which is perhaps a better known version of this principle that we generally expect editors to be competent adults (metaphorically, with apologies to all the very smart and capable teenage editors among us) that do not require special restrictions and safeguards to protect them against their own mistakes.
Regards, HaeB (talk) 07:30, 16 August 2026 (UTC)
The thing I’m concerned about is everyone knows that spellcheck is obviously not infallible, as the Internet often likes to humorously point out. But to a newcomer, anything that comes from ‘Wikipedia’ seems naturally correct and reliable.
Certainly not everyone would fall into this trap, since, as you point out, it isn’t like all newcomers are to be treated as children who don’t know what they’re doing. But it’s likely that some would; perhaps enough that the consequences of Wikipedia’s reputation for factual accuracy is something we should take into account on this. Given all of that, I don’t think we can draw a one-to-one comparison between run-of-the-mill spellcheck and an edit suggestion feature built into Wikipedia itself.
That’s why I argue that we should hold ourselves to a higher standard; after I’ve given your points some thought, however, I think you are correct that my own request for equal [to] or surpass[ing] human-level reliability is asking a bit much of a literal artificial intelligence. Simultaneously, I think spellcheck-level absurdity is far too low a bar for this feature.
Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 09:11, 16 August 2026 (UTC)
Sounds like something that could be addressed with a user level disclaimer about LLMs before the feature is enabled on their account. Czarking0 (talk) 16:02, 16 August 2026 (UTC)
That would definitely fix that issue (I’m surprised I didn’t think of that, to be honest). Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 00:11, 17 August 2026 (UTC)
I agree with @RoySmith that some suggestions are more suitable for experienced editors. I believe that checking citations is a good use case with an LLM flagging potentially problematic citations and editors verifying them manually (we have a proof-of-concept that I've been maintaining but as long as it's a userscript it won't make a dent in the sourcing problem). This is a task that is not too complex but requires familiarity with WP:RS and some experience. Alaexis¿question? 06:35, 16 August 2026 (UTC)
Based on the suggestions here, my stance is basically the same as what WP:LLM already says: typos and punctuation suggestions seem unnecessary but mostly fine (except when they're not, see the "physician" example), anything beyond that is unlikely to be fine, and suggestions related to NPOV are light years away from fine. 100% accuracy isn't possible but, realistically speaking, people are going to rubber-stamp whatever suggestions they receive.
As far as the spot-check goes, the timeframe is not really all that scientific since I actually saw this last night (not even while doing AI stuff, I was doing Commons new image patrolling and saw one of the screenshots from it), slept on it, and did the full writeup the next day. I also have seen tens of thousands of AI edit suggestions, as well as thousands of snippets of parsed Simple Summary text, so I already knew the places where this was likely to run into issues. This is also why I think the "experienced editor"/"new editor" dichotomy is not helpful, in general. Experienced editors are not necessarily experienced in copyediting and/or the specific problems that crop up with AI, and newcomers are not fragile baby birds who have never edited a piece of writing in their lives. Gnomingstuff (talk) 07:10, 16 August 2026 (UTC)
Automated suggestions could be helpful in very limited circumstances. I can imagine an experienced editor who has chosen to opt in to them assessing each proposal carefully, skilfully rejecting those with the sorts of problem listed above, and improving Wikipedia by acting only on the good ideas. I'm sceptical that AI is the best way to produce such input but I'm willing to be proven wrong. Methods such as this are totally unsuitable for new editors, many of whom will blindly obey The System, make bad edits in good faith, and get reprimanded or blocked. There is also the unoriginal but valid argument that editors would prefer the WMF to devote its resources to other activities, even if we don't always agree exactly on which ones. We're seeing a lot of these "Here's a new toy no one asked for that we've secretly blown this year's development budget on" announcements. I do honestly try to hope for the best, but I'm afraid my initial reaction has become "Oh no, what have they broken now?" Certes (talk) 09:57, 16 August 2026 (UTC)
No, LLM suggestions would be the opposite of useful. They would only reduce the amount of thinking an editor does, leading inexperienced editors to potentially make an edit because the LLM said so. People trust tool output, so if wikipedias own tooling gives them a suggestion, however ridiculous, they might easily do it.
This should be stopped at all costs. TietoTeekkari (talk) 19:48, 20 September 2026 (UTC)

Wouldn't it make more sense to focus on basic copyediting suggestions? I just went to a random article, Mircea II of Wallachia, and requested some copyedits from chatgpt.

More information suggested copyedits ...
Close

Most of these seem like pretty decent suggestions, don't dip into npov/tone danger zones, and are simple enough for new editors to review and handle. ScottishFinnishRadish (talk) 00:14, 20 August 2026 (UTC)

Also, that might be the first time my first random article click wasn't about a sports ball player. ScottishFinnishRadish (talk) 00:15, 20 August 2026 (UTC)
the nice thing about having an insite spell/grammar checker would be less people would use Grammarly (we could tell people using Grammarly to use the insite one instead, similar to how PasteCheck works). I've found asking a chatbot for corrections gives decent results, but obv you can't defer to them (esp. on content) Kowal2701 (talk, contribs) 19:59, 20 August 2026 (UTC)
Why is it a problem that people use Grammarly? RoySmith (talk) 20:02, 20 August 2026 (UTC)
it goes beyond its scope as a spell/grammar checker and rewrites content, and a lot of people using it don't know it's LLM-powered. It's the worst because people will put time and effort into writing, then put it through Grammarly which turns it into slop. We see it at AINB a fair bit Kowal2701 (talk, contribs) 20:10, 20 August 2026 (UTC)

Context from Editing Team

Hi y'all – I'm Peter Pelberg, product manager of the Editing Team. Together with the Machine Learning Team, we've been working on the experimental model-generated edit suggestions. You have raised a range of valid concerns/questions that warrant responses. You can expect those when I'm back online in earnest next week.

In the meantime, I'd like to clarify some aspects of the work and the thinking that's informing it.

First and most importantly: there are no plans right now to deploy LLM-generated edit suggestions as default-on to anyone. The closest thing to a deployment that we've talked about is a potential A/B experiment that would not move forward until we, volunteers and staff, have deemed these suggestions reliable and promising. This commitment is in Phabricator by way of the experiment (T431377) needing T428311 to happen first. For context, T428311 states, "Learn whether experienced volunteers at en.wiki think the experimental suggestions are sufficiently reliable to be shown to newcomers so that we can decide: Will we invest the effort to scale this initial set of LLM-generated MoS suggestions across languages via T431376?"

Now, with regard to what this initial set of 30,000 LLM-generated suggestions is and is not…

  1. Who has access to these experimental suggestions? In this initial, experimental phase, these suggestions will only be available to people who A) enable the Suggestion Mode beta feature and B) install this user script. Note: We will soon introduce a setting within Special:Preferences so people interested in trying out experimental suggestions don't need to install a user script to access them.
  2. What do these experimental suggestions do? This initial batch contains experimental suggestions for three types of improvements/issues (listed below). Each suggestion will highlight the span of text it is relevant to, offer a generic description of the issue, and crucially, ask the experienced volunteers encountering it to indicate whether you think the suggestion itself is valid/useful or not. And if not, offer an explanation as to why. Said another way: These suggestions do not offer fixes or ask experienced editors to make any.
  3. What are "experimental" suggestions? "Experimental" suggestions are a set of – as I think @Sohom Datta put it well – pre-alpha suggestions. The purpose of them is to evaluate their reliability. Currently, there are a total of 6 experimental suggestions available at en.wiki to people who opt-in to seeing them. 2 of these 6 suggestions are powered by machine learning models; the other 4 are based on deterministic heuristics. You can see the full list here: Special:EditChecks#Experimental checks.
  4. What types of suggestions is this LLM generating? This initial batch contains suggestions for simplifying language, rewriting language in a neutral point of view, and adjusting place-based names to follow Wikipedia's geographical guidelines. We are by no means committed to this set of suggestions or to using LLMs in general. We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk.

Last thing for now: we are aware of the sunk cost fallacy. Thank you for raising it, @Kowal2701. In fact, as I hope the above demonstrates, the whole point of developing experimental suggestions, and making them available in production to experienced volunteers who have explicitly opted into them, is as @Chaotic Enby described: to assess whether they're reliable. Ones that aren't, we will abandon. The ones that are, we'll work together to figure out how and to whom to make them available.PPelberg (WMF) (talk) 05:48, 16 August 2026 (UTC)

as […] “1.” and “3.” demonstrate. This is perhaps not the most important point, but it is bugging me: were the numbers meant to all be “1.”? Comment struck due to this being fixed; apologies. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 06:08, 16 August 2026 (UTC)
Thanks for responding now and letting us know you don't have time to fully engage immediately, hopefully you will have time to go through all the comments above later in the week as you say. As the process of contingent experiments has been brought up again, perhaps it would help to answer the experimental question plainly, because it can be very easily answered without a complicated multi-month experiment. The suggestions are not reliable, and definitely not reliable enough to show to newcomers (not that this is the only metric that could be considered). To look further, re We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk, the linked section lacks an assessment of accuracy or risk. Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers. There is a mismatch in understanding somewhere. CMD (talk) 06:18, 16 August 2026 (UTC)
Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers.
Honestly can't put it any better than that. I do appreciate the response and understand it is the weekend. Gnomingstuff (talk) 07:18, 16 August 2026 (UTC)
Tbh this is a systemic issue to do with WMF governance rather than a reflection on anyone here, where developers and idea labs seem to typically be distant from the projects and communities, and everyone at both ends is left straining to compensate for this. Phab being largely open to the public goes a long way, but I wonder if there could be something like a WP:VPI on meta for developing/screening ideas (obv staff's medium is meetings and private convos etc. and idrk what currently happens, but something like "John and I were talking yesterday about this, X is a potential solution" would be good to get some feedback at the earliest stage). I've previously suggested that staff should be encouraged to do a little bit of editing each week at any wiki they choose (and reduce hours/week a bit for this while keeping salaries the same), but it risks the WMF's status as a 501(c)(3) organization (I was told) Kowal2701 (talk, contribs) 07:57, 16 August 2026 (UTC)
I don't really know that it has anything to do with governance in this case. The issue would be the same regardless of whether the development process was public or semi-public or completely closed off; it's a classic, perennial issue of workplaces. The issue being -- and I am really trying to be as polite as possible here -- that QA is either not being done, or not being done correctly.
From what I understand, the suggestions were assessed for quality via AI, and then someone seems to have reviewed 15 suggestions manually. Unless there's more internal discussion (there probably is, this stuff is nearly impossible to follow), that seems to be the whole QA process.
Meanwhile, I actually have worked in QA. For a task like this, we would be given a spreadsheet like the ones here and would review every single line individually. We caught a lot of errors this way, the final product was better as a result, and it only took a day or two of work at most. I do not feel this is an unreasonable amount of work to be expected from a professional organization. Gnomingstuff (talk) 17:52, 16 August 2026 (UTC)
A stage of basic competitive fault-finding, annotating a large sample by hand to flag and classify problems, goes a long way. Your immediate reflex to do this as the first response above demonstrated this nicely.
If an originating team works in public to catch and patch issues in that way, before others run into them, with high standards for accuracy, that would set discussions like this off on a better foot. – SJ + 00:23, 19 August 2026 (UTC)
Thank you. Sorry, I didn't know who was in the team working on it, I trust all these people. Kowal2701 (talk, contribs) 06:51, 16 August 2026 (UTC)
I oppose any A/B tests of this feature. It is antithetical to what Wikipedia should be, and is not something the Wmf should attempt. Suggestions sent by some central, biased, singleminded entity diminishes the wide variety of input, style, viewpoint, ... that makes the richness of the community. Wikipedia ia a beacon of human diversity, an LLM is the opposite of this. Fram (talk) 06:58, 16 August 2026 (UTC)
In contrast to such hyperbolical rhetorical flourishes, the community has already been widely using AI suggestions sent by some central, biased, singleminded entity controlled by WMF for over a decade, in form of ORES (or now the "revert risk" models). These AI suggestions (on whether to revert another editor's changes as vandalism) are made available to any editor at Special:RecentChanges even though they might be seen as more consequential on average than those of the new tool under development here, and have a fairly high error (false positive) rate. I have personally implemented many thousands (probably tens of thousands) of those AI suggestions over the years, and survived skipping many more mistaken suggestions.
I don't want to entirely dismiss concerns about control by the Foundation though. In case of these vandalism detection AI models, the more recent work on them seems to have been dominated by decisions and aims of the WMF Research department that may not entirely align with the community here (for example, they seem to have foregone possibly substantial quality improvements for English Wikipedia in favor of language equity, and implemented their own conceptions of "fairness" with regard to IP editors).
The main employee behind the success and broad community acceptance of the original ORES left WMF years ago, and his parting recommendations for "Community-centered Evaluation of AI Models on Wikipedia" do by and large not seem to have been taken up by WMF.
Regards, HaeB (talk) 08:36, 16 August 2026 (UTC)
As I said earlier, keeping the prompts (and the pipeline in general) publicly available and enabling each community to tweak them would go a long way in assuaging these concerns. Some competence is required to maintain these tools, but there is no reason not to make prompts and benchmarks transparent. Alaexis¿question? 10:50, 16 August 2026 (UTC)
The system prompts used for the model appear to be published on gitlab here according to the mw:VisualEditor/Suggestion Mode/Model-generated editing suggestions#Research findings. Not sure which size Gemma gemma4:latest is, maybe the E4B? I would be interested to know if other models were evaluated, for example Nemotron 3 Super is a similarly sized model to gpt-oss:120b, but a bit newer. Mistral is dense so might be a bit slow to run. In any case, there are many newer models in the same weight class, surely it wouldn't be that hard to generate say ~1000 from each of them to see if anything is clearly better, given the only thing that has been modified is the system prompt? Alpha3031 (t • c) 12:22, 16 August 2026 (UTC)
Thanks. Going off of that, it appears that no reference information is given, which is not surprising looking at the NPOV suggestions, and that no actual Wikipedia policies/guidelines are included in the prompt. LittlePuppers (talk) 15:24, 16 August 2026 (UTC)
A successful model would likely go beyond a prompt alone and be equipped, at the very least, with retrieval-augmented generation capabilities in order to access the text of the policies/guidelines themselves, and, in the case of NPOV, ground its answer in available reliable sources. I don't think that would be enough for NPOV specifically (there is still too much of a black-box effect in the model's weights, that won't be fully offset by prompting or sourcing), but this is to say that the current approach is far from optimal. Chaotic Enby (in solidarity · talk · contribs) 15:46, 16 August 2026 (UTC)
The particular challenge is, I think, that there seems to be a trend to use increasingly sophisticated models to push increasingly challenging tasks increasingly to newer editors. LittlePuppers (talk) 15:33, 16 August 2026 (UTC)
Regarding point 4., community feedback makes it pretty clear that anything regarding NPOV or tone will not be seen as low-risk, and might be interpreted as an attempt by the WMF to influence editorial decisions. I believe it would help the prospects of the experiment to commit to follow the emerging consensus, which is clearest on this specific aspect.
On a broader level, I will reiterate my suggestions on Phabricator on working with the community:

[...] the scope and timeline of any planned community rollout, and the importance of working with the editor community and respecting a future consensus on whether to deploy it, should be made as clear as possible to avoid a repeat of Simple Summaries.

Chaotic Enby (in solidarity · talk · contribs) 12:31, 16 August 2026 (UTC)
I'm with the rest of the group in saying that the NPOV suggestions are high risk and difficult to get right. I'm excited about trying to figure out if a good suggestion model can be developed for simplifying language. There is a large body of research showing that Wikipedia's science content is often way and way too complicated and linguistic complexity is an important part of that. We've had discussions on Wikipedia about using heuristics (sentence lenght, number of syllables per word), which is contentious. Perhaps an AI model can do this better than a heuristic, as it can distinguish between a long sentence with an easy sentence structure and an impenetrable long sentence. In solidarity, —Femke (talk) 🐦 15:27, 16 August 2026 (UTC)
@Femke I have tried both AI models and conventional readability tests like Flesch–Kincaid and this is an area in which current AI tech cannot and should not be used. And my conclusion was that conventional readability tests also suck.
NPOV is another area where AIs cannot and should not be used.
We can use AI to find typos, but fixing them actually requires an experienced Wikipedian. Polygnotus (talk) 15:33, 16 August 2026 (UTC)
I too have plenty of experience using AI for this purpose and find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. As these edit checks are only about identifying overly complicated text and because we have an abundance of low-hanging fruit in this area, I'm confident that a tool can be developed. The big question is if within the limits of a certain budget, it can become good enough, and to what extent it attracts the right editors to fix it. In solidarity, —Femke (talk) 🐦 15:40, 16 August 2026 (UTC)
One of my common complaints at FAC (especially for scientific articles) is choppy writing style (sort of the opposite problem from overly difficult text). I often advise authors that they should try using some of the LLM tools to make suggestions for how their writing could be improved. I don't know how well this advice is received. I also don't know how much it is used in the intended way of offering suggestions vs just copy-pasting the output into their article; that would be an abuse of the tool, but people abuse tools all the time and the fact that they do so doesn't mean the tool is bad. RoySmith (talk) 15:43, 16 August 2026 (UTC)
find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. Subtle changes in meaning convert text supported by reliable source(s) to text not supported by reliable source(s).
I'm confident that a tool can be developed Yeah, developing a tool is the easy part. The hard part is ensuring it is a net positive.
The big question is if within the limits of a certain budget The WMF is not the right party to create such a tool because they are unfamiliar with the challenges Wikipedia editors face and because they have a tendency to waste a lot of time and effort on creating shiny new tools while neglecting decades of tech debt.
The good news is that volunteer devs can do it, but they will run into the problem that it won't work as I explained above. Polygnotus (talk) 15:46, 16 August 2026 (UTC)
Some teams are unfamiliar with the editing challenges. Other teams have a track record of listening, or have editors in their team. The editing team is an example of the latter. For instance, when people at enwiki and dewiki asked for better VE performance recently, the team managed to fix this tech debt really rapidly.
I'm very willing to help test this part of the tool. I have experience with editors doing this well and less well. In solidarity, —Femke (talk) 🐦 15:58, 16 August 2026 (UTC)
@Femke You can try WP:SCRIPTREQ, the people there are a lot more agile than the WMF is.
Please ping me if you have something I may be able to help test it. Polygnotus (talk) 16:02, 16 August 2026 (UTC)
Tbh I still think we should look at incorporating SEWP articles here as 'simple summaries', but that's a different discussion Kowal2701 (talk, contribs) 15:47, 16 August 2026 (UTC)
I wouldn't assume they'd want us to. Simplewiki is a tiny wiki with just a few active people. If we incorporate their content their vandalism rates will go through the roof. Polygnotus (talk) 15:49, 16 August 2026 (UTC)
They'd also get more good-faith editors though! What I'm thinking is that SEWP still exists as a separate wiki, we just have an opt-in feature for displaying one of their articles behind a button. IIRC @Ferien was sort of open to the general idea (may be wrong) Kowal2701 (talk, contribs) 15:53, 16 August 2026 (UTC)
Adding a CTA, if we get consent from the Simplewiki regulars, may be a good idea. Polygnotus (talk) 16:04, 16 August 2026 (UTC)
Yeah, I am personally quite open to the idea, though I am not too sure what our community as a whole would make of it at the minute. --Ferien (talk) 21:04, 17 August 2026 (UTC)
What kind of typos is AI fit to resolve that WP:AWB is not? Czarking0 (talk) 20:13, 16 August 2026 (UTC)
One thing I'm having in mind are typos that depend on the context to make sense of them (or to whether there is a typo to begin with). Chaotic Enby (in solidarity · talk · contribs) 20:19, 16 August 2026 (UTC)
FWIW, the fixes in Special:Diff/1369578064 were all suggested by Claude. I imagine most of them would have been caught by other tools, but I was impressed by the flagging of Freilberg. I don't remember exactly what it said, but the gist was that it spotted that I had "Peter Freiberg" in one place and "Peter Freilberg" in another. It figured out that these were probably referring to the same person and while it didn't know which was wrong, it assumed one of them was, and left it up to me to figure out which. RoySmith (talk) 20:24, 16 August 2026 (UTC)
And I note the firsts of those fixes is changing a direct quote in a way that is inconsistent with the source. Sure, you can probably justify that (MOS:QUOTE allows minor typographic things to be silently corrected) but I don't think that's something an AI should be recommending in any way. * Pppery * (alt) in solidarity 20:49, 16 August 2026 (UTC)
So, you're saying the fix is correct, and if a human had suggested it you would agree with it, but since an AI suggested it there's a problem? RoySmith (talk) 20:56, 16 August 2026 (UTC)
I'm saying it might be correct (not that it is correct), but determining whether it is is a judgement call I don't want an AI to make. * Pppery * (alt) in solidarity 21:22, 16 August 2026 (UTC)
The AI didn't make the judgment call. It just brought this to my attention and I made the judgement call. That's why my name is on the diff. I'm really not seeing the issue here. I made an error, a tool alerted me to it, and I fixed it. How is this a problem? RoySmith (talk) 21:36, 16 August 2026 (UTC)
The funny thing about using LLMs here is that they start working against each other. The default behavior of AI trying to "rewrite articles into formal encyclopedic tone" is to undo any text simplification (and to do so poorly, for instance replace "is" with "serves as", "uses" with "utilizes", etc). The "text simplification" category here, however, seems to be basically a catch-all. "Simplification" suggestions here range from stuff like Correct the misspelling of the player's first name from "Russel" to the proper "Russell" (which isn't even correct!) to "break up these sentences."
There's also the problem that telling someone to Rewrite the sentence for clarity and smoother flow is not actionable to the majority of people: if someone doesn't know how to copyedit then they don't know how to do that, and if someone does know how to copyedit they don't need those vague directions. It's like prompting people as AIs. Gnomingstuff (talk) 18:23, 16 August 2026 (UTC)
(Ironically, the LLM prompt that judged suggestions reads in part You are a strict reviewer. Your job is to find flaws, not to be nice. Really weird feeling to be envious of an LLM, its opinion certainly seems to be taken more seriously and its tone is given much more leeway.) Gnomingstuff (talk) 20:42, 16 August 2026 (UTC)
Who on earth outside the group of people responsible for this is going to read Gnomingstuff's report and consider those suggestions to be "reasonably accurate"?  — Hex • talk 16:02, 16 August 2026 (UTC)
Pre-alpha or not, the WMF shouldn't be experimenting with ways to funnel freeform model suggestions to editors to begin with. Machine models should never be allowed to so directly influence the contents of the project, as even if the suggestions are individually found to be valid and reliable, there will still exist overall biases. An LLM will favor certain sources, certain topics, certain sides. This will be reflected in what suggestions are and are not made, and editors evaluating and implementing individual suggestions will be entirely blind to any larger systemic issues they would be enabling.
Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact, and on a group level by the diversity of our views. A monolithic model has neither, it predicts tokens. fifteen thousand two hundred twenty four (talk) 21:45, 16 August 2026 (UTC)
"Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact - Citation needed: "making people aware of their bias doesn't do anything to mitigate it." Levivich (talk) 22:45, 16 August 2026 (UTC)
Citation: I made it up from first principals and subjective experience, the preprint will be out soon.[Humor]
I feel entirely comfortable claiming that editors who operate with awareness of their own potential biases will take steps to mitigate them in this structured environment where WP:NPOV serves as a strong guiding force. If you find this unpersuasive, so be it, it is human to disagree. fifteen thousand two hundred twenty four (talk) 23:58, 16 August 2026 (UTC)
The operative word here is can. As human beings that are capable of independent thought and self awareness, we can choose to examine our own biases and seek out information and experiences to change how we think about the world around us and the assumptions we make. Large language models are capable of exactly none of those things, because of the very simple fact that they are computer programs. Claude is exactly as capable of herself awareness as MS Paint is of having an independent thought.
Not every person makes the conscious choice of examining their baises, or is even fortunate enough to exist in a socioeconomic situation to even be able to, but that is beside the point ‑‑gurkubondinn 00:26, 17 August 2026 (UTC)
"Research from Harvard found that the effects of personal interventions such as awareness raising at a personal level are positive, but short-lived."
"And the worst method, the one that actually has no effect at all, is to tell people to be good people, to be egalitarian, and so on. ... It is easy in the sense that it last for a short period of time, but it won't last very long. ... Now, when young people encounter this result, when they see that, yes, they were able to make change, but the change doesn't last, they get very sad, because they want a better world. And I'm not at all sad about that. ... So our brains change, our minds change, associations move around, but they always will gravitate to whatever is your cultural default. And so to bring about actual change, society around us has to change, and then we will move, and then the default will be a new default."
LLMs are biased because humans are biased. Because they're trained by humans, and humans train their biases into the machines. We are no more able to cure bias in machines than we are able to cure it in ourselves. There are plenty of ways in which human intelligence is superior to machine intelligence, but lack of bias isn't one of them (in either direction). Levivich (talk) 02:57, 17 August 2026 (UTC)
Absolutely, I agree with you. That's why the models can never be "neutral" or "unbiased". My point was just that I think you misunderstood 15224's reply, people have the capability to examine and be aware of their biases, but computer programs don't. It's not automatic and it doesn't happen for all people (for various reasons), but we have the ability to examine our own biases because (unlike computer programs) we are conscious beings. It's a bad comparison is what I'm saying, and it anthropomorphises computer programs (that were created by biased people). I have a beef with whoever it was that decided to wrap LLMs in chatbot interfaces. ‑‑gurkubondinn 11:06, 17 August 2026 (UTC)
We're looking forward to getting more deeply into this with you all this week. Before that, I wanted to express gratitude to y'all for the perspectives you are sharing. From the concerns about transparency and community control over models of this sort to issues with specific suggestions you're encountering, please keep the feedback/questions/concerns/etc. coming.
And for anyone who is interested in seeing what these suggestions look like in practice, please do the following...
TRYING EXPERIMENTAL SUGGESTIONS
  1. Ensure you have the Suggestion Mode beta feature enabled
  2. Enable "experimental" suggestions by pasting the following snippet into your common.js: mw.loader.load( 'https://meta.wikimedia.org/w/index.php?title=User:DLynch_(WMF)/alwaysbesuggesting.js&action=raw&ctype=text/javascript' );
  3. Open a page in VE that has one of the 30,000 suggestions available. E.g. https://en.wikipedia.org/w/index.php?title=List_of_examples_of_Stigler%27s_law&veaction=edit
Note: this week, I'm going to see if we can share a spreadsheet so you can see all of the model-generated suggestions in one place rather than having to tap around the wiki looking for them. PPelberg (WMF) (talk) 00:59, 17 August 2026 (UTC)
I wrote a small script so you can view the list of suggested edits on an article without needing to enable the beta feature, switch to VisualEditor, or add that line to common.js to enable "experimental" suggestions: User:DVRTed/sandbox/edit-suggestions.js. — DVRTed (Talk) 03:06, 17 August 2026 (UTC)
The spreadsheet would be helpful, if only because I don't know which of the csvs is the "real" one. In general I think this kind of thing is much easier to review in spreadsheet form than article-by-article. Gnomingstuff (talk) 04:18, 17 August 2026 (UTC)
@Gnomingstuff: what you described makes total sense to me.[i][ii] I've checked in with engineering and it turns out that compiling this list will take a bit of time. Assuming nothing unexpected turns up, you can expect me to return here with a link to a CSV you (and everyone else here!) can review before this week is over.
---
i. Knowing, definitively, what suggestions we need y'alls expertise in reviewing and being able to differentiate those from earlier iterations that we've since discarded.
ii. Seeing all of the suggestions in one place so that you don't have to hunt them down yourself. PPelberg (WMF) (talk) 22:49, 17 August 2026 (UTC)
Is the intent that the U/I would just present these suggestions to the user and let them edit the article themselves if they opt to accept the suggestion? Or is the idea to have a "Make this edit" button that the user could just click? The reason I ask is that if it's the later, it would make sense to include some machine-readable marker in the wikitext (I'm thinking an HTML comment) identifying the source of the inserted text. And/or have a log of such changes. The idea is to make it easier for somebody to audit the performance afterwards. RoySmith (talk) 23:04, 17 August 2026 (UTC)
The latter would fail WP:NOLLM, so that would be a complete nonstarter. The former is not great either. ‑‑gurkubondinn 23:15, 17 August 2026 (UTC)
WP:NOLLM says "Editors are permitted to use LLMs to suggest corrections to their own writing, and to incorporate them after human review. This is limited to spelling, punctuation, capitalisation, grammar, and other simple mistakes." So this would indeed be a starter in those cases. RoySmith (talk) 23:22, 17 August 2026 (UTC)
The suggestions that Gnomingstuff went through go far beyond the very narrow exception in NOLLM. And this exception only applies to [an efitor's] own writing. ‑‑gurkubondinn 23:44, 17 August 2026 (UTC)
The latter would be impossible in the current implementation, anyway, the LLM is not prompted to provide an actual change and the suggestions are usually just stuff like "The sentence is long, contains a grammatical error, and could be expressed more clearly." Gnomingstuff (talk) 05:09, 18 August 2026 (UTC)
@RoySmith: Good question. If/when we (staff + volunteers) come to think these suggestions are reliable and useful, the intention would be to offer a suggestion that would contain the following:
1. A description of the issue and type of fix that is needed. Both of which will need to be generic enough to make sense across the contexts it might appear within while at the same time being concrete enough for the people encountering it to know how to start on the path of acting on it. So for the "Simplify language" suggestion, it might be something like, "Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language."
2. A link to the local policy/guideline the suggestion originates from. This serves both as an opportunity for people encountering a suggestion to learn more and also a way to ensure that suggestions are grounded in project consensuses and conventions.
3. Two actions: one action to Dismiss the suggestion and along with it, a way to express why someone has elected that choice. And a second action – maybe we'd label it Rewerite? – that when tapped would A) focus someone's cursor into the span of text the suggestion thinks there is an issue within and B) cause the article text in question to enter a "revising state" so people are clear about where exactly their focus is needed (see screenshot below). From there, the responsibility would be on the person acting on the suggestion to decide what to fix and how to fix it. Said another way: there are NO plans for these suggestions to make fixes with the click of a button let alone to describe specific solutions.
Screenshot showing the revising text state within Suggestion Mode
If you'd appreciate a more concise answer, what @Gnomingstuff described here is spot-on. PPelberg (WMF) (talk) 05:32, 18 August 2026 (UTC)
@PPelberg (WMF): So how would you cram enough information in such a tiny area to give the person all the information they need? Are you aware that many guidelines are like 7k words? If you give a newcomer a link to WP:NPOV, that is obviously not enough to have them actually be able to judge the neutrality of an article if they do not have relevant experience and knowledge of the field. And what will happen when inevitably people complain that edits are not improvements? Will this just WP:BITE newcomers even more? Polygnotus (talk) 05:40, 18 August 2026 (UTC)
@Polygnotus: great spot. The questions you're asking sit at the very core of this project.[i]
We seem to be aligned in thinking[ii] that some suggestions may be more harmful than helpful to show to newcomers. They're likely, as you alluded to, too complex and experience-dependent to distill down into a relatively small piece of guidance. If we don't account for this, we could lead newer folks into publishing edits that experienced volunteers revert or respond to with hostility. This could in turn drive these potential contributors away.
And while I don't think we can know for certain which suggestions those are in advance, I think we (staff and volunteers) are develpoiong a pretty good sense that conversations like this one are helping us refine.
I also think we'll learn a lot over time by trying things out, and there are a few features we've put in place to help:
Experimental suggestions: suggestions can be enabled as either default-on or experimental. The latter means that people who have explicitly enabled the soon-to-be-available setting have the space to safely experiment with suggestions. Through that testing, they can decide who—if anyone—an experimental suggestion should be shown to by default.
Tags: all edits in which someone sees and/or acts on a suggestion are tagged. You can see these in action by filtering Special:RecentChanges for Edit Suggestion seen or Edit Suggestion used.
On-wiki configuration: volunteers can independently decide the minimum number of edits someone must have published in order to see a suggestion. If you visit Special:EditChecks and look at the link suggestion, you'll see that volunteers have set the minimumEditCount value to 1000. This means the suggestion will only be shown to editors who have made ≥1,000 edits.
The idea is that, together, the above will let us see the kinds of edits these suggestions lead folks to make in practice and, with that, decide whether, how, where, and to whom they're shown.
How does this sound to you? What, if anything, do you think we might be missing or misunderstanding?
---
i. I hear the questions you're asking as something like: "How might we translate a great deal of nuance and complexity into a format that is simultaneously 1) succinct and simple enough that newcomers will engage with it and 2) explanatory enough that in doing so newcomers will be equipped with the information and know-how they need to act on them in ways experienced volunteers see as constructive?"
ii. Please correct me if I've misinterpreted what you've said PPelberg (WMF) (talk) 05:32, 19 August 2026 (UTC)
I would urge you to please stop developing this feature. It seems like any version of this feature will be detrimental to wikipedia.
A suggestion from an LLM, a suggestion from an official wikipedia tool, will easily be read by an editor as trustworthy, when we know it is not. LLM usage offloads critical thinking to software, meaning the editor is doing less of that themselves. This alone will have a negative impact on quality. And even grammar edits can change the meaning of a sentence, so this is just a non-starter in general.
Please do not pursue this track. TietoTeekkari (talk) 20:00, 20 September 2026 (UTC)

I'm pessimistic about this, but I have a very high bar for pessimism/hopelessness to stop me from supporting a low-stakes experiment. Giving this tool to experienced users to try out seems like one of those low-stakes experiments worth trying, in case there's a way to make it work (again, I'm pessimistic, but possibly with certain constraints regarding task and topic?).
But also, just to put a finer point on something, because I think it does good to repeat it now and then: there's the worry about the quality of these suggestions, but there's also the worry about, for lack of a better word, branding. The branding that the WMF has begun using, "knowledge is human", is an echo of a popular sentiment here. At a time when absolutely every company and every project is cramming in as much AI as possible, Wikipedia is mostly headed in the other direction. That's a good thing IMO, but as a result, any pitch someone has for an LLM-based tool is evaluated with a handicap applied. If you'd otherwise be graded on a scale of 1-10 where 1 is harmful and 10 is helpful, start by subtracting 2 or 3 for "we don't want to be associated with that" (or, alternatively, "we see these as detrimental by default"), and it needs to be really useful to get a critical mass of people behind it. But no complex tool starts its life as an 8+ on that scale, I don't think, so super-low-stakes tests are IMO a good way to start working towards it. FWIW. — Rhododendrites talk \\ 22:06, 16 August 2026 (UTC)

I really appreciate and agree with the branding note Rhododendrite gives here. Best, Barkeep49 (talk) 22:48, 16 August 2026 (UTC)

A humorous interlude

I know people love to make fun of AI hallucinations, so I couldn't resist posting this fun map (4:38 in the video). Strange spellings aside, Hoboken, NJ has gotten transported to the midwest, Chicago is on the Pacific coast, and Boston has been relocated to Colorado. I want some of whatever it's smoking. RoySmith (talk) 02:12, 17 August 2026 (UTC)

Also available in Africa. I think I'll stick to the Commons maps. Certes (talk) 14:30, 17 August 2026 (UTC)
These LLM map-fails are legion. My take away from these examples is that they are not only vivid examples of LLM hallucination, and thus LLM unreliability, but they are also vivid examples of the unreliability of humans, because each one of these published hallucinations is only possible because humans obviously failed to check the work--to even look at the maps--before publishing. Humans are unreliable: an important lesson for any crowdsourced project. Levivich (talk) 14:57, 17 August 2026 (UTC)
I don't know about this particular channel but loads of them are (almost) fully automated with no human in the loop. Polygnotus (talk) 15:39, 17 August 2026 (UTC)
Then again, aside from quite a few mistakes, have you all seen the "Chloe vs History" channel on youtube? Quite the advancements in AI lately. Some of the advancements have been made on this channel, which should probably have a Wikipedia page. Randy Kryn (talk) 15:42, 17 August 2026 (UTC)
That Africa one is from a US State Dept presentation. That ain't no YouTube clickbait, that's supposedly professional humans at work. You'd think they'd have looked at the slides before making their presentation. And, I'm speculating here, but I bet multiple humans were involved in that, because when the US State Dept gives a presentation at an int'l conference, I don't think it's just one person creating and making the presentation all by themselves. Maybe. But either way: evidence that at least sometimes, even professional humans presenting at an int'l conference obviously don't bother to check their work. I guess my point is: what makes LLMs unreliable isn't just that they hallucinate, it's that humans won't catch it because sometimes they don't even bother to look. Another well-known example is lawyers submitting briefs to courts with hallucinated citations -- literally licensed professionals in the performance of their profession. So this isn't a problem limited to youtube clickbaiters or "kids on the internet," even trained and licensed professionals, even on a global stage, succumb to laziness. At least the lawyers get fined for it -- now there's a new fundraising channel for the WMF: fine editors for publishing hallucinations on-wiki! Levivich (talk) 16:15, 17 August 2026 (UTC)
At least with hallucinations, they're usually immediately obvious to anybody who bothers to look. This particular example is clearly YouTube clickbait. The bottom-feeders who produce these things don't give a whit about accuracy, just that they can churn out some mildly entertaining garbage videos that collect likes and shares and other revenue-producing metrics. But the fact that people are willing to abuse a tool for their own commercial benefit doesn't mean that the tool is inherently worthless. People abuse wikipedia in all sorts of ways (spam, SEO, reputation management, advertising, etc). Does that mean we should shut Wikipedia down? RoySmith (talk) 15:49, 17 August 2026 (UTC)
PS, it doesn't take AI to generate garbage maps. Some people are able to do it all by themselves with nothing more technologically advanced than a sharpie. RoySmith (talk) 16:22, 17 August 2026 (UTC)
Sure, and wikipedia had hoaxes and plenty of innocent misinformation long before LLMs became popular, but the problem, of course, is one of scale: LLMs make this sort of thing like 100x more common than before. It used to take hours to write a good hoax on wikipedia, now it takes minutes. Levivich (talk) 16:30, 17 August 2026 (UTC)
"Chiiicago". AAND it's in multiple places at once. Ladies,Gentlemen and enbies, The Chicago hivemind. Starlet 01:43, 25 August 2026 (UTC)

Back to business after the humorous interlude, a full ban on AI?

  • No AI on Wikipedia and this should be extended to article talk pages, Wikipedia discussions, visible editing suggestions, or draft pages. Too many work-arounds and exceptions have been allowed and these issues and concerns could go on for years. There should be some good faith edit reverts for AI edits made in mainspace while the confusion existed but maybe end the confusion from here-on-in with a blanket ban on AI. Thanks (darn it, the word "thanks" was aided by AI). Randy Kryn (talk) 15:06, 17 August 2026 (UTC)
    I'd support that. I've had enough with AI turning everything to shit, and clearly WMF can't take no for an answer. Tercer (talk) 15:55, 17 August 2026 (UTC)
    I'd say that Polygnotus should be allowed to play around with AI because Polygnotus actually understands its limitations and assumes full responsibility for any edit made with their account.
    The WMF is a different story unfortunately. Polygnotus (talk) 15:59, 17 August 2026 (UTC)
    Does Polygnotus always refer to himself in the third person? RoySmith (talk) 16:03, 17 August 2026 (UTC)
    Polygnotus does not. But Polygnotus could! Polygnotus (talk) 16:05, 17 August 2026 (UTC)
    AI, or LLM? Cluebot, for example uses machine learning (AI but not LLM) to great benefit with remarkably few errors. Certes (talk) 16:36, 17 August 2026 (UTC)
    Generative AI is what bothers me, irrespectively of the underlying technology. I'm fine with cluebot because it doesn't generate content, so there's no risk of poisoning Wikipedia. I would still be fine with it if it were LLM-powered (even if that would be completely pointless). Tercer (talk) 17:46, 17 August 2026 (UTC)
    I would absolutely support a ban on AI (LLMs) for (direct or indirect) content generation or suggestion. Ita140188 (talk) 19:12, 17 August 2026 (UTC)
    I would definitely support this. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 19:14, 17 August 2026 (UTC)
    I would support this on principal, at least until the AI industry stops being complete garbage. —Leaf.Sheap ⇖ /.°°.\ ⇗ (They•Them) 21:18, 17 August 2026 (UTC)
    I would support aswell. I don't think (generative) AI has a place on Earth at all, let alone on wikipedia (which has always been guided by the principles of human knowledge and effort). TheDowningStreetCat (talk) 23:41, 17 August 2026 (UTC)
    100% support. I fail to see how AI slop would actually help the project in any conceivable way. All I’m seeing is negatives… lots of them. 296cherry (talk) 18:29, 20 August 2026 (UTC)
    What would you think of suggestions more along the lines of what I outlined here? ScottishFinnishRadish (talk) 18:32, 20 August 2026 (UTC)
    I would support a ban on LLMs to stop similar things from happening in the future. TietoTeekkari (talk) 20:02, 20 September 2026 (UTC)
  • Right now one of our biggest strengths is that we are a large, human written, generative AI free chunk of knowledge. My crystal ball isn't working, so I can't tell you if AI will be good enough to do something cool with Wikipedia in the future. That said, Wikipedia is one of the last places which should incorporate AI because it will give up our unique advantage. Proposals like this from the WMF are not sufficient, are not close, and make me want to use blanket statements to convince the WMF not to keep trying to shoehorn AI into things over and over again. Tazerdadog (talk) 16:59, 17 August 2026 (UTC)
    I feel like it should be pointed out, again, that this is not a proposal. It's an experiment. "Learn whether experienced volunteers at en.wiki think the experimental suggestions are sufficiently reliable" is what they're doing. I think Gnoming helpfully answered that question. Levivich (talk) 17:58, 17 August 2026 (UTC)
  • This would be an overreaction. User:Polygnotus and many others have been building various LLM-powered tools, including ones that are used to detect LLM edits (User:Fermiboson/AIlog) and to fight vandalism (Wikishield - LLM is optional). There are more than one hundred editors using these tools. We do want WMF to experiment and implement the best ideas. Alaexis¿question? 07:06, 18 August 2026 (UTC)
  • A full ban does not make sense. We already have a range of community tools that do cool things with Wikipedia using AI. In particular I want the best available tools for review, including those that take advantage of AI for trainable pattern matching and classification. That includes anything that helps with slop-detection, edit checks, citation checks, page linting, and vandal-fighting. This experiment falls under review... with a higher bar for quality I can see a range of suggestions being useful, particularly after a few rounds of focusing on quality. – SJ + 01:07, 19 August 2026 (UTC)
  • I would support a ban on generative AI touching any Wiki content regardless of whether LLMs improve. JoelleJay (talk) 11:30, 22 August 2026 (UTC)

Continued discussion

Thank you all for keeping the conversation going, and thank you to Gnomingstuff and others for taking the time to look through outputs. @PPelberg (WMF), @SSalgaonkar-WMF, and I have read the whole thread and tried to pull out some of the most important things we heard and questions being asked. Peter and Sucheta, please add if I missed anything (well, for that matter -- anyone please tell me if I missed anything). I'm just going to list these questions/topics for now, and we'll keep working through them and hopefully you'll keep discussing with us (there are many of you and not as many of us!) This is to build on what I posted above and what Peter posted above.

Firstly, I think we're on the same page about something especially important: we're not going to deploy LLM-backed suggestions to anyone unless this community is supportive of it. And if they do become deployed, they'll be configurable via Special:EditChecks the way existing checks and suggestions are (i.e. community could decide who gets to see them, change what they say, what articles they show up on, or turn them off). I hope that in projects like these, we at WMF are bringing capabilities to volunteers so that they can produce the kinds of checks and suggestions that help both newcomers and experienced editors get the most good wiki work done with the least amount of drudgery. There is also a question of whether these suggestions would be a vehicle for content suggestions -- the answer there is no, we are not designing these to propose text to the user; rather they point out places the user should look and what they should look for. The LLM explanations that Gnomingstuff pointed out initially are in those files as a way for us to understand internally why the LLM is making the suggestion; they would not be shown to editors. (But I know there is a more subtle question here -- when just pointing out a sentence that needs attention in some way exerts that sort of influence).

In that vein, I wanted to say that this current project is more about figuring out what it might be like for checks/suggestions to be generated via LLMs, than about what those exact types of suggestions are. If NPOV is not a good one to pursue (volunteers here have given many good reasons why NPOV is particularly tricky), then we should put that one down in favor of simpler ones to try out. (Worth noting that checks and suggestions are being generated via a bunch of other ways, too, like simple rules, text match, simple models, etc.)

Also just a quick nomenclature thing:

  • Checks: this refers to edit checks that react to what the editor is doing at that moment in the editor. e.g. they paste in a blob of text from ChatGPT, and the check pops up right then and says, "Please avoid copying text from other sources".
  • Suggestions: this refers to suggestions for improvement that have been pre-calculated and are waiting to show up when someone clicks Edit, e.g. they open up the editor, and there is a box that says, "This link appears more than once in this section." Right now, Suggestion Mode is available on English Wikipedia as a beta feature, and suggestions are available in the feed on Special:Homepage.

You can see all the checks and suggestions and their statuses at Special:EditChecks (note that "experimental" checks are only available to people who have a specific user script installed).

Sorry -- this post is getting long. But on to the list of questions we want to be able to talk about here, raised in the thread above. This list of questions is not something that we just want to provide answers to and then expect that everyone will agree with us. This is a complex area, and we don't have all the answers, but we're trying to figure them out with you.

  1. What kind of communication with communities has there been on this project so far?
  2. Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing?
  3. How did we / do we QA lists of suggestions like these?
  4. What if communities don't want certain suggestions on their wiki?
  5. What determines whether suggestions are good enough to go beyond testing?
  6. Is NPOV appropriate to point an LLM at, given its nuance and complexity?
  7. If an LLM provides suggestions, how do we make sure that the editor isn't swayed/biased just by virtue of it coming from Wikipedia (a trusted source)?
  8. Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things?
  9. How could the models we use be transparent, open, and have community controls/auditing?

Alright -- more to come as we get into those questions. -- MMiller (WMF) (talk) 23:53, 17 August 2026 (UTC)

@MMiller (WMF): Hi Marshall!
Because of a long history the relation between the community and the WMF is, let's say, far from perfect. And arguably worse than ever.
This is obviously not your fault, but its important to set the scene.
Because Wikipedians are smart and informed they are, generally, skeptical and wary of AI.
This is the correct position in 2026, because we do not yet know the impact of AI on the environment, jobs, and the upcoming resource wars.
The community generates the value and people think they donate to support it. The WMF is terrible at doing the community actually wants and needs. MediaWiki's tech debt is a sad joke. Community members who try to explain what we need from the WMF are routinely ignored. For decades.
Instead of working on the things that are important, the WMF nerds build shiny new toys. Debugging old code sucks. Doing something fun with AI/ML is fun. Because WMF leadership is terrible it seems (from the outside) that no one is working on the important stuff while we get an endless stream of halfbaked projects that waste a lot of time and money. The WMF basically never finishes a project it starts, overcommits at an early stage, and only gets feedback when its too late.
The WMF has a tendency to drop in and reveal they spent a lot of time working on a terrible idea, without asking input from the community, and then are surprised when the community rejects it.
The WMF is terrible at communicating its wishes and goals and what its working on.
The previous WMF attempt to "do something with AI" was a terrible idea and everyone hated it. It proved yet again that the WMF does not understand the Wikipedia community, what writing an encyclopedia means, and (the limitations of) AI.
The WMF did not learn from this and is again presenting a half-baked plan they spent a lot and time and resources on.
The community desperately wants and needs the WMF to succeed and produce great software at a reasonable cost, but that has yet to happen.
Quite a few members of the community are nerds who have spent a lot of time playing around with AI so they know what its limitations are in the context of writing an encyclopedia.
As the person behind the AI Proofreader, AI Source Verifier, AI Editsummary etc I am clearly not some anti-AI luddite.
The WMF is actively making the community more and more anti-AI, and to be honest rightly so. The WMF has no respect for the hard work of countless Wikipedians.
The fact that the WMF does things like have an AI check for NPOV problems shows that the WMF does not understand AI or its limitations or how to write an encyclopedia.
The community could greatly benefit from responsible AI use in a few specific tasks, whereby the human makes the final decision and is responsible for the edit. And the WMF makes it impossible for me to communicate that to people.
At this point, we both know its too late to listen to my feedback. This terrible idea will continue no matter what. The WMF has spent millions on AI related stuff and any benefits to the community were not proportional to the amount of money spent. It doesn't matter to the WMF; they had fun and can put something cool AI-related on their CV and move on.
The WMF has a toxic positivity problem wherein honesty and negative feedback is punished and ignored and all criticism must be hidden below 7 layers of a compliment sandwich. Even people far more diplomatic than I am just can't deal with all the corpo-speak and manipulation.
In an ideal world the WMF would stop what its doing and actually listen.
LLMs present an unique set of challenges and even opportunities and the WMF is fucking it up for everyone else and does not seem to understand how much damage they are causing and can cause in the future.
Is NPOV appropriate to point an LLM at, given its nuance and complexity? No, and that is a silly question. I don't teach my fish braille and ask it to rewrite the bible. Tools are useful in some contexts and bad in others.
What if communities don't want certain suggestions on their wiki? None of the proposed suggestions in its current form would be an improvement.
Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things? No, of course not.
How could the models we use be transparent, open, and have community controls/auditing? That is impossible unless you accept the models are terrible compared to competitors. There are no ethical LLMs available. And making something like that would be unethical and require unethical actions.
Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing? Because the important work sucks and nerds think its fun to play around with AI. Also the WMF does not understand its role; it should serve and protect the community. In any well-run company you do maybe 90% important stuff and 10% fun stuff.
In the future, the community should be involved at the earliest opportunity, when brainstorming. The WMF should learn from its mistakes. Reflect on Simple Summaries. What went wrong, why, how to avoid that in the future? Allowing random newcomers to act on typofixes proposed by an AI will just make a lot of people very angry, sets newcomers up for failure and degrades the quality of Wikipedia. I use the opposite approach; my software helps experienced Wikipedians to fix typos, and an AI helps filter out things that aren't typos, and no one objects to that.
Polygnotus (talk) 00:41, 18 August 2026 (UTC)
Please assume good faith. No one is doing this to have fun or put "something cool" on their CV and move on.
Wikimedia has actually spent a terrifyingly small amount on AI infra and tooling, which is part of the problem here: when an experiment is run, fast iteration on prompts, benchmarks, evals, and models isn't second nature. There are indeed ethical language models, as others note below, and better orchestration would help choose the best for a given task. – SJ + 16:02, 21 August 2026 (UTC)
Marshall (and colleagues), I will say that from reading your post earlier, I think it reflects a good attitude for approaching the issue. (I'm not going to go down the rabbit hole of the difference between words and actions and trust.) I will also note that I have no idea how much say you have in what projects you work on and how far you take them. (Although if "Senior Director of Product" isn't just a bunch of fancy words, I would guess quite a bit.)
I clicked your link to see what's currently on Special:EditChecks, and I will say that almost all of them seem like really good ideas. None of them (the ones I like, at least), require any AI beyond some if statements in a trenchcoat.
In short: I think we could completely ignore all this AI stuff and you'd have some really great and helpful projects to work on. (Others have been mentioned above.)
I know working with such a large community can be hard. I would certainly like to think that we could give you some advice on how to help us. Ask . We only bite sometimes. Thanks for taking the time to listen. LittlePuppers (talk) 01:47, 18 August 2026 (UTC)
That is an excellent and fitting username. Polygnotus (talk) 01:53, 18 August 2026 (UTC)
One time I said that I didn't know what "Director of Product" means (I am not a native speaker) and the ex-CEO of the WMF started attacking and accusing me because she couldn't handle mild criticism. Polygnotus (talk) 03:11, 18 August 2026 (UTC)
Thanks -- just to take this opportunity to shed a little light on what I do and how we're set up:
  • We're set up as a bunch of "cross-functional" product teams. Cross-functional means each team has a few engineers, an engineer manager, a designer, a product manager, a data analyst, a movement communications specialist (give or take -- sometimes a couple teams will share people in certain roles).
  • The product manager is responsible for setting the priorities of the team, what order to work on them, deciding what is in and out of scope, deciding whether the results of an experiment show that something is worth pursuing further. They do this in collaboration with their teammates, not unilaterally.
  • As a director of product, many of these product managers report up to me. When I started at the foundation, I was the PM of the Growth team, and I've been around for eight years and am now in a management role.
  • So in my role, I look across the various teams and play a role in setting the overall priorities -- like which various potential editing projects the various teams working on contributors should prioritize, what the readers teams should prioritize. I work with director counterparts in engineering and design to do this.
MMiller (WMF) (talk) 18:28, 18 August 2026 (UTC)
Re 6, and to some extent 9: Does the WMF have a suitable operational definition of NPOV to measure against? As I understand it, transformer-based models are fairly common in sentiment analysis nowadays, but I don't know if pre-existing models (if any exist) would be well adapted to an encyclopedic context, and I don't think they would be sufficiently interpretable. In any case, I'd say that it is something that is strictly more difficult than tone check, as well as likely being considered higher risk by the community. I can't be certain of course, but I would expect for the community to grant social licence to a model for NPOV specifically would, at minimum, require much better interpretability than typical for current models. Whether the team believes they can achieve that is probably something best answered by the technical staff, the community can only inform as to acceptance criteria. Alpha3031 (t • c) 05:00, 18 August 2026 (UTC)
@Alpha3031 They just take off-the-shelf selfhosted open-weight models, which are way worse than Claude Fable 5 (and other mainstream commercial offerings) like gpt-oss:120b aya:35b / aya-expanse:32b, llama4:scout, qwen3:235b and qwen3:1.7b and then give it a presumably AI generated prompt containing:
 Role: You are an expert Wikipedia Copyeditor and Reviewer.
    Goal: Conduct a professional audit of an article's plaintext so it aligns with Wikipedia core content policies and Manual of Style.

    Context & Constraints:
    - Plaintext environment: non-prose elements may be stripped (infoboxes, references, templates, media, etc.).
    - Non-markup focus: do not suggest wiki-syntax, template, link-formatting, or reference-format edits.
    - Awareness of extraction gaps: if text appears truncated from extraction, do not flag it unless clearly an authorial issue.
    - Focus only on prose, terminology, neutrality, and information structure.
And
    - Neutrality: remove peacock terms, bias, and unnecessary loaded language.
You appear to be overestimating the sophistication of their approach by a rather wide margin.
Polygnotus (talk) 05:14, 18 August 2026 (UTC)
I am aware of that, given that I had read (and in fact posted the link to) said prompts above. Though, I wouldn't say hosted open weights models are necessarily "way worse" than commercial offerings currently given recent and especially upcoming releases (GLM at 700B especially is likely a lot easier to host than the 2.xT models, though still harder than 120B of course). The page does indicate that the team recognises some situations may require bespoke models though. Alpha3031 (t • c) 07:05, 18 August 2026 (UTC)
Open-weights models can be perfectly adequate for some use cases. For source verification the very modest gpt-oss-20b model worked as well as Sonnet 5. Detecting NPOV issues is just a much harder problem, and it needs much more context and probably better models as well. Alaexis¿question? 07:18, 18 August 2026 (UTC)
I don't disagree that open-weights models (even older or smaller ones) can be adequate for many tasks. I just also wanted to point out that as of the time of the current discussion, there are several open-weights models in the 700 billion to 3 trillion parameter range that are comparable to the frontier in a much wider selection of tasks, namely Kimi, Qwen and the upcoming GLM 5.3 (which being much smaller, is probably going to be cheaper to evaluate).
Incidentally, it does appear that previous WMF research findings (m:Research:Test External AI Models for Integration into the Wikimedia Ecosystem#Evaluation Results and Findings) have pointed out the difficulty of NPOV assessment:

Some policy detection tasks are hard for humans and machines alike. Specifically for NPOV, precision across the three model families is very low. Models overall do better when detecting Peacock behavior. This is because while understanding neutrality requires in-depth reasoning, [bold mine] peacock behavior can be detected via language features.

so it's somewhat odd that they've picked it as a easy, low-risk task in this specific project.
I would say that adequate NPOV assessment is probably the hardest possible task to set as a goal, given that it requires the aforementioned tone check, a comparison with the sources, and also some sort of test to ensure the sources the other models see are an accurate reflection of the body of published RS more generally. It's definitely not suitable for a project intended to explore what can be done with relatively generic models. I think there could also be room to explore, e.g., faster community feedback cycles so that the WMF doesn't feel like it needs a whole year to develop a single task specific model. Alpha3031 (t • c) 08:56, 18 August 2026 (UTC)
Agree that NPOV assessment is the hardest possible task, also for humans. More importantly, the assessment relies on consensus. There is no authority that has the truth about whether something is NPOV. The whole point is that the community reaches a consensus looking at all available reliable sources. Delegating this to an LLM (with the mark of approval of the WMF itself, thus giving it an undue resemblance of authority) defies the premise of Wikipedia, which is based on decisions taken by consensus. Aggravated by the fact that LLMs are not at all unbiased by any definition of the term (and are often just plainly wrong). Ita140188 (talk) 10:27, 18 August 2026 (UTC)
Agreed, NPOV is hard! Alaexis¿question? 12:52, 18 August 2026 (UTC)
Thanks for the response, this is already much more transparency than the last time around and I appreciate it.
I am a bit confused as to whether these questions are for us and for you. Gnomingstuff (talk) 17:35, 18 August 2026 (UTC)
@Gnomingstuff: oh, good clarification. The questions Marshall posted are a first pass at translating what we're hearing from y'all into a set of questions that we (staff) can then respond to one-by-one. Of course, if you think there are questions we've missed and/or questions that we've misinterpreted, please comment as much. PPelberg (WMF) (talk) 20:05, 18 August 2026 (UTC)
I thought I would ring in to point to another AI-backed project WMF is doing and share our experience working with volunteers on it, given the interest here in this sort of thing. For context, I lead the Product Safety and Integrity (PSI) team here at WMF, we build security and safety features.
One thing we're working on is detecting abusive content using LLM-based models. Specifically to enwiki, last week we deployed a new feature that is only visible to oversighters, at Special:AbuseReview, which uses an open-weight model called CoPE to flag edits that probably need to be suppressed due to containing personal information.
This on-wiki feature is now being used enwiki oversighters to review and take action on what it raises. They tell us the practical accuracy rate they experience in reviewing the output is better than what they see in the reports they get from human users. And, when we were testing the model output in June and July, in batches through an off-wiki process, most of the true positives it was catching were not being caught organically on-wiki.
One thing I want to mention is the iteration and volunteer collaboration that it took for us to make something deployable. There was some initial volunteer skepticism, and we needed to demonstrate not only that it was valuable, but that this was going to be accurate enough to not waste precious volunteer time. That took some internal experimentation on our part to get something that showed enough promise on both those fronts that we could ask for support doing a round of manual labeling -- which is what really pushed it over the edge of accuracy to be a deployable feature.
This collaboration has helped us do something meaningful about doxxing on English Wikipedia over the last few months, that feels good to both WMF and volunteers. That is possible in large part because volunteers were willing to give us space to operate, and to keep an open mind that could be convinced by data. We also needed to be flexible in our own thinking and incorporate volunteer ideas, something PSI has gotten used to doing in our work. EMill-WMF (talk) 21:17, 19 August 2026 (UTC)
Endorsing Eric's statement, as one of the oversighters involved in the testing/iteration he describes. This is catching OS-level material we were not catching otherwise, and is a clear benefit to the encyclopedia. In solidarity, asilvering (talk) 21:54, 19 August 2026 (UTC)
I think there is a role for AI to play on Wikipedia and it's going to be in something kind of like this. Something Eric alluded to but doesn't say directly which I think is important: the first draft the WMF showed us was nowhere close to ready. From that the learning was that we needed to go even farther in how confident . The second draft - which is when we started doing manual labeling - was still not anything which would have been appropriate for use. It did lead to a learning that one element - personal information - was more reliable than the other OS criteria which got us to the third draft and the efforts to then backfill June & July. That got us a lot of incredibly sensitive information that wasn't getting reported and which is now getting appropriately oversighted. I was the one who ran some data after we completed June to compare the true positive rate against email reports and found it to be in the range of 5% higher. I plan to revisit these stats soon and my hope would be that we'd have a higher difference between email reports and these flags for two reasons: 1) the classifier has been improved since I ran those numbers originally 2) some of the "obviously needs OS" tickets we'd have gotten in the past, we won't be getting because OS will be oversighting edits before some other qualified person finds them and reports them. I am really glad we're getting these signals now to stop some very sensitive personal information from being exposed against policy, we wouldn't have been able to do it without the AI classifer model, and it did take an iterative process to get there. Best, Barkeep49 (talk) 23:05, 19 August 2026 (UTC)

Downloaded the most recent csv, picked an NPOV one at random:

  • "1358549724,The_Dying_Rooms,10791564,NPOV,Remove biased and profane language describing the Chinese government and replace it with a neutral summary of the government's response.,"In the film, Blewett and others travel to mainland China to visit orphanages housing children abandoned due to the ""one-child policy"". The filmmakers stated that unwanted female and disabled children were left to die of neglect, allowing parents to have another child. Showing that China government is a piece of sh!t for letting this happened and being a coward by lying to our faces that it didn't happened and said that all the footage is fabricated to destroy the reputation of the china government.>",enwiki,df59b690-6c8e-4827-b809-76e90d409f9f,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

The statement the LLM objects to was in the article for less than a minute and long reverted by the time the suggestion was created. So one can add to all the above problems and errors that it also wastes times and resources by not checking the current version but some snapshot.

Profanity checks in general are a bad idea.

  • "1358746317,2024_G20_Rio_de_Janeiro_summit,72163669,NPOV,Rephrase the incident with the first lady using neutral language and omit the explicit profanity.,"During a speech about fake news, Rosângela Lula da Silva, first lady of Brazil, swore at Elon Musk, saying: ""I'm not afraid of you. Fuck you, Elon Musk"" (Eu não tenho medo de você. Inclusive, fuck you, Elon Musk, in Portuguese).",enwiki,fba2904a-9276-47e3-8f5c-27fe5aff71f6,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

No, we are not going to "omit the explicit profanity" from a quote, and suggesting things like this is a very, very bad idea.

  • "1357774252,God_Emperor_Trump,74631506,NPOV,Remove profanity from the description of the phrase on the sword to maintain a neutral and encyclopedic tone.,"According to Fabrizio, the phrase could mean 'here's your fucking tariffs'.",enwiki,0f886b1f-2a4c-4303-b06e-63da6700b71e,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

Again, it's a quote, from the creator of the sculpture. The AI should not make statements like "Editors often revise this kind of wording, saying the tone is unbalanced. ", which only works to influence newbies by making a false claim to authority.

And then there the internally contradictory advices, indicating the inherent stupidity of LLMs.

  • "1356142445,Allan_Segura_(model),80130058,simplify_language,Combine the two sentences about sexual orientation and activism into one concise sentence.,Allan Segura is openly gay. He is a transgender rights activist.,enwiki,d28d67bc-2b32-4cbb-9135-ed391c91c13b,Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language. [Learn more](https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style#Vocabulary).,Simplify language"

So do we need to combine these two (very short) sentences into one sentence, or do we need to rewrite this using shorter sentences? Or, just perhaps, our readers are perfectly capable of understanding these two sentences and won't "find this text difficult to understand". What a joke. Fram (talk) 09:01, 18 August 2026 (UTC)

@Fram Also, wasn't the fact that the WMF didn't do content the thing protecting them in lawsuits?
Let's say an article contains negative information about a rich person. If the WMF starts to mess with content, do they not open themselves up to be forced to make changes? I am not a lawyer. Polygnotus (talk) 12:23, 18 August 2026 (UTC)
All good reasons not to use AI for this purpose and probably not for any purpose. Once again, the WMF is wasting what remains of its valuable technical staff (and its ample funds) on trying to drag Wikipedia in completely the wrong direction.
If the bot is criticising vandalism which was only live for a minute, I suspect that it may be reading page history (why?) rather than taking a snapshot. Certes (talk) 12:32, 18 August 2026 (UTC)
A snapshot sample of a large number of articles will also catch revisions that lasted only for a minute. There are lots such edits that are reverted quickly.
Assuming that edit suggestions are not displayed when the content has changed, this particular suggestion would never have been shown - no harm down do anyone. Generating checks on a snapshot rather than live is an architectural decision driven by performance, complexity and other considerations. Alaexis¿question? 12:51, 18 August 2026 (UTC)
Presumably that means any such system will need to be rerun on each relevant page each time they are edited, in case the edit touched the prompted part? Not sure how the current tasks handle this actually. CMD (talk) 12:59, 18 August 2026 (UTC)
Not necessarily, it can be run once a month, with some kind of caching enabled not to re-check the vast majority of content that stays the same. Then the complex and time-consuming part (LLM calls) is done asynchronously, and the easy part (deterministically checking that the content stayed the same) is done when the user opens visual editor. Alaexis¿question? 14:34, 18 August 2026 (UTC)
That is indeed how it's currently working. A large batch of the suggestions were pre-generated, and was poured into a fairly simple API that VisualEditor's suggestion mode knows how to query and match up to a document.
Presumably if this was a successful experiment that we decide together to scale up, it'd turn into a more complicated system where we do something like fire off a job that precomputes the suggestion for each new revision of a page. Still avoiding needing to do the expensive generation every time someone opens the editor, but keeping them more up-to-date. DLynch (WMF) (talk) 23:03, 19 August 2026 (UTC)
Again, it's a quote, from the creator of the sculpture.
That's another LLM tic -- at least when it comes to Wikipedia edits/suggestions, they really hate direct quotes and will usually tell you to paraphrase them and/or do so themselves. (Example from a 2026 edit summary: "Paraphrased a lengthy direct quote regarding the skater's performance at the 2025 World Team Trophy into a concise summary. This improves readability and helps maintain a neutral, encyclopedic tone by removing overly detailed personal reflections.") Gnomingstuff (talk) 17:41, 18 August 2026 (UTC)
The examples Fram shows above makes me more resolute in opposing this on principle. The WMF is trying to exert editorial control and hiding it by labelling them as "suggestions". ♠JCW555 (talk)♠ 16:48, 18 August 2026 (UTC)
@JCW555, I guarantee you they are not trying to exert editorial control with this feature. They're trying it because they think it will be helpful for beginner editors. We can tell them they're wrong and we don't like the feature without impugning their motives. In solidarity, asilvering (talk) 21:21, 19 August 2026 (UTC)
But suggesting sentences should be reworded is the WMF trying to exert editorial control. To quote MMiller above "It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?". Whether a sentence/passage needs to be reworded is a matter for the talk page amongst editors, not through the WMF via their AI. No reply from the WMF in this section has done anything to assuage that concern for me. If the WMF comes out and says that these "suggestions" wouldn't touch content at all, no matter how small or big, then I'd be a tad less aggressive in my opposition. ♠JCW555 (talk)♠ 21:44, 19 August 2026 (UTC)

Is NPOV appropriate to point an LLM at, given its nuance and complexity?

Hi all, I'm Sucheta, product manager on the Machine Learning side of this work.

Reading everything you all have said, it’s clear we shouldn’t move any further with the LLM-generated NPOV suggestions. We're dropping that type rather than trying to iterate on it.

@Ita140188 said this really well: There is no authority that has the truth about whether something is NPOV. Makes sense to me – NPOV isn’t just about word choice; it's about the representation of a topic that editors decide on by weighing the available reliable sources against each other and reaching consensus. We had wanted to take a crack at it to see what the LLM would produce, so thank you for looking at these and thinking about them.

I do want to make the distinction between these LLM-backed NPOV suggestions and the Revise Tone suggestions that are in production now on this wiki. Those Revise Tone suggestions come from a model called BERT that we’ve fine-tuned for a narrower scope. The model is trained to notice peacock language, based on 20,000 examples of revisions where the "peacock" template was added or removed. We think these have worked out well, and thousands of them have been actioned by both new and experienced editors on this wiki. Having them available also made it more likely that a newcomer would make a constructive edit, than when they open the editor without some suggestion inside. You can see them on Special:Homepage in the suggested edits feed (if you check these out and have thoughts, please let us know).

What about the other types of LLM suggestions besides NPOV? Well, let’s keep talking here about whether they have potential. Note: we're working on making a spreadsheet available to you all so that you can see all of the suggestions in one place. -- SSalgaonkar-WMF (talk) 17:13, 18 August 2026 (UTC)

Thanks for the response. Good to hear about the NPOV feature.
The main issue is mostly the same across the board: the actual LLM suggestions have issues as above, but including broad categories without the actual suggestions is just confusing and provides almost no context. This is going to affect any possible category, it's just structurally inherent to the task as I understand it. Other than that:
MOS:GEO: Two issues I can think of:
  • Seems near-certain to inadvertently wade into a geopolitical quagmire of some sort.
  • Less dramatically, a large proportion of these suggestions are really just suggestions to treat everything as first reference (the "Atlantic Ocean"/"Atlantic" thing mentioned above). I don't think there's any way to get around this while working at the single-sentence level.
Simplify language: Two issues again --
  • The text parsing needs to be fixed before anything is done with this since otherwise the suggestions won't make any sense, especially the issue of parsing multiple sentences as one (e.g. King's fourth novel, Euphoria (2014), was inspired by events in the life of anthropologist Margaret Mead. It won the inaugural Kirkus Prize for Fiction and the 2014 New England Book Award for Fiction, and was a finalist for the 2014 National Book Critics Circle Award. Euphoria was listed among The New York Times Book Review's 10 Best Books of 2014, TIME's Top 10 Fiction Books of 2014, and the Amazon Best Books of 2014. -- they aren't visible on-page but there are unicode separators between many of the clauses, maybe that's related).
  • The category scope seems to be off. The vast majority of suggestions are to break up sentences, comparably little about actually simplifying language -- if anything, the language suggestions I've found seem to be suggesting the opposite, to make language more complex. But suggestions for typo fixes etc. also show up here.
Gnomingstuff (talk) 17:58, 18 August 2026 (UTC)
That unicode-characters thing is actually deliberate. The data comes pre-massaged into the exact form that works in VisualEditor's search (non-text gets replaced with a opening and closing internal model tag, so that  is actually <ref></ref>). E.g. If you go to the Lily_King page that quote's from and paste it into the VE search box, it should highlight that entire paragraph, covering the citations. As far as I know, that replacement got done as a post-processing phase after the initial suggestion-generation, though I wasn't directly involved so there's a chance I'm wrong.
...that said, you did make me realize that we're accidentally not comparing them in that form in our final "has the user already changed this bit of text since they started editing" check, so gerrit:1326927 will make all these suggestions that cover citations / templates actually visible for review. DLynch (WMF) (talk) 20:54, 18 August 2026 (UTC)
  • Questions about peacock language; Does the function ignore direct quotes? Any suggestion to change the wording of a direct quote shouldn't happen. Also, can we can we get it to come down hard on subjective words like best while being less aggressive on words that might be verifiable facts like largest? And maybe even less aggressive for largest known and largest on record? --Guy Macon (talk) 18:39, 18 August 2026 (UTC)
    @Guy Macon: Great question. This suggestion can be configured to ignore quoted content which, at present, is exactly what en.wiki has done.
    You'll notice that ignoreQuotedContent within Special:EditChecks#tone is set to true. If you'd like to learn more about exactly about how quoted content is detected, T414715 contains more details. Could you please let me know if anything you see (or don't see) brings other questions to mind?
    Now, to the second question you're asking...
    Assuming it's accurate for me to understand it as something like "How might we specify the suggestion based on specific words/phrases?" I wonder if you think TextMatch could be helpful here. In essence, it enables volunteers to write custom suggestions to appear when predefined words/phrases are detected with an article. PPelberg (WMF) (talk) 19:45, 18 August 2026 (UTC)
An update on the NPOV suggestions. As of ~30 minutes ago, we've removed all NPOV suggestions from the experimental batch of model-generated suggestions. Thank you all for the quick feedback here and for trusting us to hear you.
Note: if, by chance, you happen to still encounter one, can you please let us know? For now, we implemented the above in a bit of a fragile way so we can get something out quickly and will come back to make this more robust in the coming days. PPelberg (WMF) (talk) 22:11, 18 August 2026 (UTC)
Thanks, love this fast and encouraging response. If peacock language detection is working well with a training set of 20,000 examples, is compiling training data a helpful step for catching other narrow style issues? – SJ + 01:20, 19 August 2026 (UTC)
Great question! I think so - though this project is meant to test how viable it is to generate suggestions without manually compiling the training data you described.
If we were to place the LLM-generated suggestions and Tone Check approaches on a spectrum, Tone Check would sit at one end: it works well for identifying a well-defined issue, but it took us over a year to build and deliver, with training and evaluation being two of the most time-consuming pieces. We're now considering three ways to generate suggestions using models, among the many other approaches we're exploring:
  • generic models with task-specific prompts (what we're testing here)
  • generic models plus fine-tuning
  • specialized, bespoke models
So the question we're really asking is what amount of rigor is required to produce suggestions that are sufficiently reliable and useful?
We're asking a similar question about evaluation: whether faster methods like LLM-as-a-judge can tell us if suggestions are reliable without hand-labeling a large test set. Even here we review a number of samples manually to make sure the judge is scoring appropriately; that sample is just much smaller. We also created an eval dataset of past user edits per suggestion type so we could compare the model-generated suggestions to actual edits.
What do you think about this lens? Do you think there are some suggestion types that could be supported by this approach of using generic models without fine-tuning? SSalgaonkar-WMF (talk) 15:42, 20 August 2026 (UTC)
How much effort is something like this? ScottishFinnishRadish (talk) 16:01, 20 August 2026 (UTC)
Thank you for sending this! I really like this idea of experimenting with LLMs to generate basic copyediting suggestions. Reiterating what you said: they seem to be relatively low-risk, simple, and as a result, something newcomers could handle. In terms of effort, I don't think it would require an impractical amount to build a dataset like this (as your demo even shows).
We started to explore copyediting suggestions using MoS guides around capitalization and grammar, but we didn’t feel confident enough about their quality and utility to include them in this initial dataset. The capitalization suggestions scored lower in our LLM-as-a-judge evaluation than other suggestion types, and our manual review showed that “grammar” was too broad a category; we couldn’t come up with one label or description to describe the range of issues surfaced by grammar suggestions.
These both feel like solvable problems, and I’d really love for us to take another look. Would you be willing to share the prompt you sent to ChatGPT? SSalgaonkar-WMF (talk) 13:54, 21 August 2026 (UTC)
Find some typos that can be fixed or copy that can be edited at https://en.wikipedia.org/wiki/Mircea_II_of_Wallachia. If I were going to refine it I would probably break results down to things that need review from someone with varying levels of experience so you can filter the output to users based on experience. For instance, checking if the source said first born or firstborn is a great task for someone trying to step up from beginner editing. ScottishFinnishRadish (talk) 15:10, 21 August 2026 (UTC)
I don't know if the technology can support this, but it would be nice if we could have multiple queues of suggestions in different areas. A queue for spelling and grammar fixes and a queue for source-to-text validation in STEM articles might appeal to different people looking for work. RoySmith (talk) 16:02, 21 August 2026 (UTC)
When I was looking through the csv before I didn't find errors when the llm was identifying the wrong units being used or similar. In my personal testing I've found it good at picking up tense mismatches (occasional errors sure but overall picks things up I've missed when rewriting something). This sort of pattern fixing seems much less likely to cause an issue than setting an llm to gambol over fields of longer text, as well as being a simpler spot and fix for newcomers. CMD (talk) 00:45, 19 August 2026 (UTC)
My apologies if I just read past it and didn't notice, but where is this CSV file? RoySmith (talk) 00:49, 19 August 2026 (UTC)
@RoySmith In , GnomingStuff dug it up. CMD (talk) 01:11, 19 August 2026 (UTC)
Got it, thanks. RoySmith (talk) 01:14, 19 August 2026 (UTC)
We'll be sharing a clearer spreadsheet version (with the most up-to-date entries) tomorrow, per Peter above. HTH. Quiddity (WMF) (talk) 01:17, 19 August 2026 (UTC)

Why are we working on suggestions?

The first reason we're working on this is because of the need to get more new people involved in editing. Being a newcomer has always been hard, and is especially hard now that people spend most of their online time on mobile. When newcomers (especially on mobile) open up the editor for the first time, it is overwhelming and they often just leave -- they are like "Wow, scary, nevermind." (here are some interesting survey results about this moment) But we've seen that when we point out specific bits of the article that could use improvement, the newcomers are much more likely to do something constructive and stick around. We've tested many of the checks and suggestions in Special:EditChecks and they have had these measurable positive impacts. We’ve also been inspired by the tools/scripts/gadgets that volunteers have built that do similar things (some examples here).

So this project here (the LLM suggestions) is another way of learning how we might find more kinds of suggestions in the vein of "how can we help newcomers on mobile be more and more constructive and more likely to stick around?" (in ways that align with policies, values, and existing editor workflows).  It’s also worth noting that the newcomers who start with suggestions often wander off on their own in the wiki once they get comfortable. The design helps with that, because the suggestions happen inside the Visual Editor, i.e. you have to edit in VE to get them done (as opposed to, say, a separate interface).

Secondly, we also think that this can help lower patroller burdens at a time when those burdens are increasing because of AI slop, as Gnomingstuff has pointed out. (i.e. newcomers making constructive edits in the first place lowers burden on patrollers from newcomers being confused). For example, we are developing a way to deter and label edits when people are pasting content from an LLM.

And thirdly, we think that these suggestions can be helpful for experienced editors, too.  A few of you have said in this conversation that experienced editors don’t need help finding improvements to make, but we have also heard from many who appreciate suggestions like these. In my own editing experience, I often click edit to do a specific change, but then discover a few other small things to improve via the suggestions.  So we think there is also opportunity here to help experienced editors get more wiki work done with less seeking/searching/effort.  It becomes a question of which of these signals to present to which users in which places.

How does this all sound? It would be great to hear from anyone who has seen these suggestions in action with newcomers or has been using them.

-- MMiller (WMF) (talk) 19:49, 18 August 2026 (UTC)

Around here, the city is running an e-scooter pilot. Install the app, hop on a scooter, ride to where you're going. One of the interesting things they do is the app automatically imposes a lower speed limit on all new riders. Once you've ridden more than (IIRC) 10 hours, you get to go full speed. I think it also won't let newbies take out a scooter after sunset. I forget the details, but you get the idea.
I could see doing something similar here. Have some way of scoring suggestions for how risky they are. Fixing an obvious typo is pretty low risk. Rephrasing a statement that appears to be biased is higher risk. The type of article might also factor into it: an article about a WP:CTOP would probably not be the best choice for a newbie to learn on. The total neophyte would only get the safest suggestions. People who had gotten a bit more experience might be offered a wider range of suggestions. RoySmith (talk) 20:14, 18 August 2026 (UTC)
MMiller (WMF), two thoughts come to mind:
  1. A good place to ask might be WT:AFC and similar (e.g. NPP)—on one hand, writing articles is kind of exactly what we don't want new editors to try, because it's really hard, but we see about every kind of possible error there: from formatting (a first heading duplicating the page title, formatting inside headings, malformed templates, all sorts of weird stuff) to tone (as established this is hard, but there are probably a few ways to detect COI) to referencing (citing Wikipedia, citing social media, just not referencing, weird formatting, duplicated refs). I see that some of it you have projects related to, but that's a handful more off the top of my head, and I'm sure people at those pages can think of more. I'd also be curious if you've investigated what the impacts of limiting this to visual editor are (i.e. how many new editors use VE).
  2. Another thing to consider is looking at what gadgets and user scripts are commonly installed. Your check about disambiguation links makes a lot of sense to me, because there's been a gadget which displays them in a different color for years. And some of those do make more sense as user scripts or being community maintained, but there are definitely some which would make more sense as part of MediaWiki or which could use some love. (Looking through my user scripts, there are a handful I don't really use, and some which are enwiki-specific, but also many which make a lot of sense to integrate or which I'm importing cross-wiki and haven't been updated for 8 years or something.)
I don't know if you're short on ideas or not, but those are what came to mind for me.
And I just reread your post and realized you already did #2. LittlePuppers (talk) 21:21, 18 August 2026 (UTC)
@LittlePuppers: thank you for sharing ideas of places to look for and vet ideas for new Checks and Suggestions. I've added WT:AFC and Wikipedia:New pages patrol to to the MediaWiki page where we are bringing together this sort of information. If/when other ideas come to mind, we'd be thrilled if you'd add them directly. Of course, I'm happy to add them as well. Just give me a ping if/when something strikes you.
On the topic of gadgets and user scripts, we're with you here. In fact, doing what you described helped prompt the work we're partnering with @Alaexis on to turn a tool he wrote for identifying cases where a source might not support its associated claim into a new experimental suggestion. PPelberg (WMF) (talk) 23:13, 18 August 2026 (UTC)
Re writing articles is kind of exactly what we don't want new editors to try, I'm not sure I agree with that. For some new users who don't know how to get started, having them fix typos might indeed be a good way to ease them into editing. But some people will already know what they want to write about. If you tell them, "No, that's too hard, we want you to fix typos for a while", all you're likely to have done is lost an opportunity to get a new editor hooked on the project. The very first edit I ever made was to create City Island Bridge. RoySmith (talk) 01:32, 19 August 2026 (UTC)
Yeah, I was wondering if someone would push back on that. My wording there was probably too strong, main point being that it's really hard and, I suspect, very often discouraging. But then again, I do see, on occasion, someone who will read through guidelines and knows how to write and all that and write pretty decent articles very early on. To be honest, those are probably also the people who write decent articles later on.
Today, your first edit would have to go through AfC, and get declined for being unsourced... ah, those were simpler times. There was probably also more low-hanging fruit then. I have no idea where I'm going with this comment. I'm all for developing features to help out those who start in all sorts of ways. LittlePuppers (talk) 01:59, 19 August 2026 (UTC)
The difficulty is that so very many new editors know what they want to write about, but what they want to write about isn't going to make it to article form at the present time. They want to write about themselves, their companies, their favourite YouTuber, the assignment they've been given, the really fun thing they made up...
Yes, we would lose them if we told them to fix typos. But we also lose them if we don't let them create that article they want, and we can't let them create that article they want. Many of them are also happily LLM-ing it up in their draft. I almost wonder whether there's a call for a service (human, AI/LLM, both?) that we try to push would-be article creators towards that assesses or helps them assess whether they actually have sources - and tells them to stop if they don't, or at least to try another subject instead. Perhaps an AI/LLM project could be given some basic guidelines about common problems (interviews are likely to be inappropriate; this source doesn't seem to be about the subject) and at least head off the ones that don't have a chance. A lot of people seem remarkably inclined to listen to what a machine says, possibly because it's seen as authoritative and knowledgeable on basically any subject. Meadowlark (talk) 06:27, 19 August 2026 (UTC)
Hi @Meadowlark, I'm Rita Ho, director of design at WMF working as the design counterpart to Marshall with product teams working on these new editing tools. Your comment about helping editors to create new articles by providing more guidelines (like including sources) is related to another feature in development, called Article Guidance! This feature is aimed at helping newer editors who want to create articles to succeed. It's kind of like if Article Wizard could be tailored and offer specific support depending on the type of article being created (animal/building/person/etc). It includes initial source validation, notability risk assessment, and initial minimal content structure or "outline" for someone to get started, and is community configurable.
Sharing in case you and others are interested to learn about this other project and participating in testing and giving feedback. RHo (WMF) (talk) 22:01, 19 August 2026 (UTC)
I do like having community-create outlines. That seems (at least in theory) to catch a lot of the types of mistakes I see with new editors (sources!). LittlePuppers (talk) 22:11, 19 August 2026 (UTC)
It becomes a question of which of these signals to present to which users in which places.
Building on what Marshall shared above, I think it might be useful to consider that we're trying out a range of ways for generating the signals to power new edit checks and suggestions...
Some suggestions, like adding references and detecting when someone has pasted content from an LLM, are generated using deterministic rules.[i] Some look for the presence of specific templates, like citation needed. Others use small language models specially trained on Wikipedia edits to identify a specific type of issue, like finding issues with tone or suggesting images to add to articles. There are also suggestions written by volunteers using the TextMatch feature (inspired by AbuseFilter). We're also experimenting with a way for volunteers to create suggestions based on the presence of maintenance templates or missing template parameters.
I share all of this in an effort to communicate that we're eager and open to experiment with a signals from a variety of sources. What's most important to us is identifying ones that we collectively see as reliable and useful.
---
i. E.g. contents of clipboard metadata and the absence of a reference within a defined amount of new text PPelberg (WMF) (talk) 23:03, 18 August 2026 (UTC)
Thanks, I've edited the mediawiki page to add two ideas, one to identify/highlight problematic text that a maintenance tag is referring to, one to point people to en:Help:Find sources if they try to use an unreliable source like a blog or social media. I think it's probably best if we focus on problems that would result in a revert for the sake of retention of newcomers (rather than minor stuff like MOS:GEO, where people can learn from someone editing their work). I'd also suggest working more closely with WP:AIT if you aren't already (if they have the time, people such as @Polygnotus, Alaexis (as I see you're doing), @Dreamyshade) as they'll likely have a lot of ideas and be able to discern some issues which may not be apparent. Otherwise I'd just encourage people to boldly edit the mediawiki page and add any ideas or whatever. Maybe it could have a second column for concerns about an idea? Kowal2701 (talk, contribs) 08:41, 19 August 2026 (UTC)
Approaching this from the narrow goal of improving Wikipedia, I'm concerned to see newcomers and suggestions again mentioned in the same breath. Automatically produced suggestions need to be assessed individually by someone familiar with writing for Wikipedia, and that's not a newcomer. If the goal is not to improve Wikipedia but to increase the active editor count by making newcomers feel useful then this may be a viable scheme, in the same way that we reluctantly allow education projects to introduce so many errors to our articles. However, if that is the case then (yet again) editors and the WMF are pulling in different directions and we have a clear conflict of interests to resolve. Certes (talk) 09:18, 19 August 2026 (UTC)
@Certes -- yeah, I understand. We've been talking about this since the early days of the Growth team in 2018: were suggested edits more about improving Wikipedia, or about retaining newcomers so that they could grow into editors who improve Wikipedia later? We generally erred on the side of retaining newcomers -- for instance, the first suggested edit was "add a link". Do the wikis really need lots more blue links between articles? Some wikis do, some not really. But it was a good task for newcomers to get their feet wet, have a succesful first experience, and want to come back again. We saw some newcomers "go on a run" where they did hundreds of those tasks over the course of several days. And we (happily) also saw many do a few of them and then go do some other, higher value edits on their own.
But the ideal is that we can do both at the same time: design tasks that are both healthy for newcomers and constructive for the wikis. I would say that "revise tone" is an example of this. I would be curious what you think of the diffs in this Recent Changes filter, which shows both links and tone diffs, highlighted by how experienced the person is.
And more generally, where you would come down on that trade-off between "invest in newcomers learning" versus "constructive edits now" (hoping, of course, that we could have our cake and eat it too with the right designs). MMiller (WMF) (talk) 21:28, 19 August 2026 (UTC)
I rarely improve tone and am no expert on it but I looked at the first three recent changes by different editors. The first looks like a useful improvement from a new editor who clearly already has the skills we need. The second slightly misses the point: I've seen the film and the character's defining aspect is that he is a local legend. The third is also wide of the mark: much of the promotion is in the paragraph before the one the editor changed. Its edit summary of "changed tone(bot told me to)" is also concerning: perhaps the editor values obeying "the bot" above using their judgement or feels that this is what is required to become an accepted editor.
Having our cake and eating it would be nice, but I do tend towards constructive edits over using articles as a sandbox. Certes (talk) 21:46, 19 August 2026 (UTC)
Taking a look at the tone edits, starting at the bottom:
  • Special:Diff/1370168493: The edit in isolation is ok but the edit summary suggests that it is AI-generated, and the many other edits they have pumped out such as Special:Diff/1369053156 and especially this promotional draft corroborate that. The ideal response here would be for them to stop using AI for promotional edits. I would also suspect an SPA based on this edit history.
  • Special:Diff/1370225916: Probably OK, no immediate red flags
  • Special:Diff/1370168654: Obvious promotional AI slop.
  • Special:Diff/1370168718: Probably OK, no immediate red flags
  • Special:Diff/1370168757: Obvious promotional AI slop by the same person who did the last one, which should illustrate the volume at which this adds AI slop to the wiki. (Also, the topic is potentially controversial.)
  • Special:Diff/1370169145: Obvious promotional AI slop by the same person, I'm going to skip their edits from here on out but just know that I am skipping a lot of bad edits as a result.
  • Special:Diff/1370170671: OK but not really a tone edit, and their edit history is somewhat suspect. Also, Special:Diff/1369948373 does not actually fix the tone, it just puts a band-aid over it, which is the other problem with these edits.
  • Special:Diff/1370171161: Probably OK.
  • Special:Diff/1370171537: Promotional AI slop (by someone else this time, Special:Diff/1350356412 is the smoking gun and they have several warnings on their talk page). As you can see from their edit history they are also pumping these out at high volume so I will also be skipping their many bad edits.
  • Special:Diff/1370171930: Grammar edit that does not actually fix the tone, it is still promotional
  • Special:Diff/1370172403: Not a tone issue. The fact that there is an obvious tone issue literally one word away does not speak highly of their competence in editing. (that is, competence in editing English-language writing, not Wikipedia specifically)
  • Special:Diff/1370174754: Probably OK, no red flags
  • Special:Diff/1370176147: Probably OK, no red flags
  • Special:Diff/1370176495: Obvious promotional AI slop and they didn't even bother removing the citation markers.
This is the thing that I have been trying to point out for several months now. The feature is a net negative. It may result in numbers going up in terms of edits by newcomers, but the workload for editors also goes up -- if they even notice it -- to a point where there are simply not enough people available to cleanup. This also means that revert rates are artificially low, because again, there are not enough people to even see them. Gnomingstuff (talk) 22:07, 19 August 2026 (UTC)
Okay, well as a third person who has randomly gone through some of these:
  1. diff: change is an improvement; the paragraph is unsourced, and might be better removed. (My changes: pt 1, pt 2, pt 3; still room for improvement.)
  2. diff: may or may not be an improvement; I haven't seen the film, and I doubt the editor had either.
  3. diff: possibly worse, though I haven't seen the show.
  4. diff: likely improvement. Removes unsourced content.
  5. diff: minor but definite improvement. Edit summary includes "bot told me to". I'd probably also remove "even" or reword, but "one of the largest" is likely factual, albeit not sourced inline.
  6. diff: could be worded better but it's an improvement.
  7. diff: minor improvement, room for more work.
  8. diff: meh, at least there's a period at the end of the sentence now.
  9. diff: not really an improvement.
  10. diff: improvement.
A common theme is that many of these seem to involve people changing tone without the background knowledge to understand the article. "Tone" is also vastly oversimplifying the variety of problems these articles have. LittlePuppers (talk) 22:15, 19 August 2026 (UTC)
I actually have seen the show; the former text was an accurate description of the character, especially given that the characters in Sunny are... extreme and exaggerated people. So this person and/or any hypothetical AI they are using does not know the difference between fictional characters and real people.
The other issue -- and this is an issue across the board with all newcomer tasks -- is that the template on the article is actually more specific than revising tone: it says that the article is written in a primarily in-universe style and should not be. I assume they were not told that in the task. Gnomingstuff (talk) 22:22, 19 August 2026 (UTC)
Re: newcomers vs experienced editors this is partially covered by Peter's comment above. I.e. These features are entirely configurable (and extensible) by each local community, and importantly, that means that some of the types of Suggestion can be completely limited so they're only seen by highly-experienced editors. If you/anyone can think of types of Suggestions that would be widely useful for just experienced users, and if those Suggestions can be programmatically recognized by simple textmatching, or more complicated types of code within the extension code, or via a locally run/controlled LLM, then it should be possible to setup those kinds of things (to test, and if proven useful then to make available to all editors who fit the locally defined configuration).
Also, one of the main goals of the feature is to help provide editors (both newcomer and experienced editors) with a handy link to the relevant guideline/policy within the Suggestion card's text. That makes it easier for editors to learn (or remind ourselves) of the specific nuances involved in any particular fix. E.g. If someone is editing a disambig page, and they ignored the EditNotice reminding them of the basic guidelines, then when they add two links within a single entry (or stumble upon an existing entry which does so during their editing), there could be a Suggestion that informs them of that guideline and has a singular pointer to the details (versus the nine links in the current Enwiki Editnotice). I.e. Targeted micro-tutorials that show up when relevant. The main limitation for what can be written in the Suggestion card, is length, as it also needs to work in a mobile-sized screen. Quiddity (WMF) (talk) 22:53, 19 August 2026 (UTC)
Just want to reiterate that I appreciate the response here, it is already a much better dialogue than before and that's what I was trying for, to keep the focus on the content and implementation details itself. It may surprise some people here but I am not blanket anti-AI across the board no matter what. I just don't think that providing them to newcomers who don't have a sense of our AI guidelines is a good idea. (For more on that reread GreenLipstickLesbian's comment)
Some things I don't think I've seen mentioned:
  • Right now, this and other task suggestions seem to target tagged/templated articles with high view counts. I think both parts of this are a mistake. People template articles because they want to make them better and, by and large, the result has been templated articles getting either worse, or impossible to tease apart good vs. bad edits due to the sheer volume of them. Something like this might look good from an edit metrics standpoint but from the standpoint of doing patrol it is a lot of patrol work -- and that's just the entry point to one article. Something like that can easily spawn 5 more tabs to check if someone was indeed doing high-volume AI edits. And of course the higher the view count, the higher stakes any mistakes become. (I will say that Paste Check has been somewhat helpful, it technically creates more patrol work but it's more like pointing to patrol work that would exist no matter what)
  • If the text parsing can't be fixed (Though it should be), the fail states are systematic enough that you can probably just regex filter somewhere in the process.
  • Controversial material needs to be blacklisted for reasons that should be clear. Admittedly the three examples above were somewhat cherry picked to demonstrate that point, but they were not especially hard to cherry pick. (The third one was just CTRL-F "partisan" because I already knew LLMs have issues with that, and sure enough there they were.) This should also be low-tech. Something like blacklisting CTOP and BLP articles -- I know Simple Summaries blacklisted BLP articles, though there were a lot of holes in that implementation -- and more importantly, an additional filter at the sentence level that is just blunt keyword-based stuff. This should also make MOS:GEO suggestions much safer since it's really hard bordering on impossible to predict every way that can go wrong.
  • I don't know what level of LLM familiarity the people working on this have, but I assume it is more in the LLM-coding-in-general realm, less the intersection of LLM output and Wikipedia. So, seconding getting in touch with the people at WP:AIT, they've been iterating on stuff like this for a while and have a sense of what does and does not work. I also can't speak for everyone, but while the people on AI Cleanup are less likely to be open to AI and even less likely to have any spare time to help out with stuff, they do have more of a front-row seat to AI edits and edit suggestions than basically anyone else. Most if not all of the issues above were predictable when you've seen a lot of LLM edits on Wikipedia. I have only found one common LLM edit problem that hasn't shown up in these (which is actually kind of interesting from a model standpoint). Feel free to email me, I have a great deal of collected data on this.
  • This is cheating since I've mentioned it, but "more QA" seems to be a constant across the board for all features. Another place where getting in touch with AI folks can help because spot checking goes a lot faster when you know what to look out for.
(Also, I know this isn't something you have any way of knowing about, but I prefer to not be pinged to ongoing discussions I am participating in, it just generates notifications and emails that pile up). Gnomingstuff (talk) 18:28, 19 August 2026 (UTC)

Where can you see all of the model-generated MoS suggestions?

Okay. Linked here you will find a spreadsheet that contains the batch of ~6,500 experimental LLM MoS suggestions as they appear in Suggestion Mode for people who enabled experimental suggestions.

Within the spreadsheet are two categories of suggestions: those related to simplifying language and MOS:GEO. We've removed all of the NPOV suggestions based on the feedback y'all have been helpfully sharing here.

How do these suggestions look to you? What are examples of specific suggestions that you find to be unhelpful, confusing, and/or just plain wrong? As you're going through these suggestions, what broader patterns are you noticing/conclusions are you reaching about these two categories of suggestions?

With this feedback in-hand, we're thinking we (staff and volunteers) can take a step back together and decide how and if we should move forward with this particular set of suggestions.

Of course, if you find any part(s) of the spreadsheet are unclear, please let us know.

A couple of notes:

  1. We welcome feedback in whatever form is most convenient for you. E.g. sharing directly in this discussion, enabling experimental suggestions in Suggestion Mode (see instructions) and offering feedback through the UI, etc.
  2. Some suggestions may be present in the spreadsheet and not visible when you edit the actual article with experimental suggestions enabled. This is because this spreadsheet contains suggestions that were generated in a batch offline. In the time since, some articles have been edited in ways that make the suggestions obsolete.

PPelberg (WMF) (talk) 22:35, 19 August 2026 (UTC)

Thank you! Will take a look Gnomingstuff (talk) 22:43, 19 August 2026 (UTC)
You bet and thank you! PPelberg (WMF) (talk) 22:49, 19 August 2026 (UTC)
Is it worth implementing some sort of minimum length check for simplify language? I am not sure how much shorter "Scorer for Crystal Palace; Terry Fenwick" or "Town in Faisalabad District" can be. The model seems to feel parentheticals are a lot more complex than they are in reality: "Romelda Aiken-George (née Aiken}; born 19 November 1988) is a Jamaican netball player." It also appears to be picking up some template code or something ([ { "List of leaders of Markazi Jamiat Ahle Hadith": "Order" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "1" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "2" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" } ]) as well as picking up lists as prose ("See also
List of prime ministers of Pakistan List of presidents of Pakistan Chief Secretary Khyber Pakhtunkhwa List of chief ministers of Punjab List of chief ministers of Sindh List of chief ministers of Balochistan". It's even picked up the reference list for 1994_Andhra_Pradesh_Legislative_Assembly_election as in need of simplification.
If the list could be refined to remove the simply technically non-applicable, it would be easier to parse. I would be interested in a few examples worked through by editors here who have spent a lot of time looking into language simplification (@Femke?). For example, "It is the feminine form of the Late Latin name Clarus which meant "clear, bright, famous"" and "A space station (or orbital station) is a spacecraft which remains in orbit and hosts humans for extended periods of time" seem near fully concise to me, but I have not looked into this topic much. (Looking now as I copy these in, these may also represent examples of the model looking at too short a sentence and being confused by parentheticals.) CMD (talk) 02:29, 20 August 2026 (UTC)
In this case, cross-referencing against the csv I have, the AI's "suggestion" is Fix the misplaced bracket and punctuation in the parenthetical expression for the maiden name. Which actually is a legitimate fix -- there's a stray curly brace -- but isn't related to simplification at all, see my comment about it being a catch-all.
I know it might be confusing and the spreadsheet here is supposed to mimic what editors see, but for QA purposes having those around without having to cross-reference might be helpful to pin down what's going on Gnomingstuff (talk) 03:04, 20 August 2026 (UTC)
Thanks, same issue as a lot of the MOS:GEO examples then in the llm finding something that is out of its supposed scope leading to a confusing tag. CMD (talk) 03:07, 20 August 2026 (UTC)
@Chipmunkdavis: thank you for reviewing! Responses to the points I see you raising...
Is it worth implementing some sort of minimum length check for simplify language?
Great question. Assuming we were to agree this suggestion was worth moving forward with, I think we could implement a way for you all to set a minimumCharacters value as is currently available for Reference Check. cc @DLynch (WMF) who can say definitively whether this is feasible.
The model seems to feel parentheticals are a lot more complex than they are in reality... as well as picking up lists as prose...
Great spots and noted. We'll investigate whether there are things we could do to exclude these two category of issues you're describing. PPelberg (WMF) (talk) 23:35, 20 August 2026 (UTC)
Sure, there's two ways we could do it. The equivalent to the reference check would involve doing it client-side -- you'd set a config value in Editcheck-config.json on-wiki and we'd filter out suggestions that we get from the API that're below whatever length you specify. The other way would be to configure the model to not generate those suggestions in the first place, which would be more-ideal overall in terms of work saved. Harder for you to configure though. DLynch (WMF) (talk) 23:41, 20 August 2026 (UTC)
By the way, for the "simplify language" suggestions I was wondering if existing small language models could do a similar job, so I bashed together something (I used the existing WMF readability model, m:Machine learning models/Production/Multilingual readability model card, though it isn't really designed for sentence level ARA... so possibly a way to improve things if a more suitable model can be found). I did have Gemini 3.6 Flash screen the about 2/3rds of the suggestions and fill out the description field (columns are same as the CSV instead of the spreadsheets) since I haven't investigated which features produced the scores, not sure how useful those descriptions would be. I've posted the suggestions it generated here if anyone is interested in looking at them and comparing to the ones generated by GPT-OSS. Alpha3031 (t • c) 15:16, 22 August 2026 (UTC)

What if communities don't want certain suggestions on their wiki?

If a wiki does not want a certain suggestion to be shown to anyone on their wiki, any administrator or interface administrator can disable it directly, without requiring any code changes/action from WMF staff.

In practice, this would look like someone visiting MediaWiki:Editcheck-config.json, identifying the Edit Suggestion or Check they'd like to disable, and changing ShowAsCheck or showAsSuggestion from True to False. We're of course happy to help out/clarify any confusion. Although, the goal is for this system to be intuitive enough for you all to adjust it based on what you're seeing in practice and the expertise you've developed.

In addition to toggling Checks and Suggestions ON/OFF, there are a range of other ways you all can customize them:

  • Who sees them. account limits a Check or Suggestion to people who are logged-in or logged-out and minimumEditCount/maximumEditCount enables you to target them based on someone's editing experience.
  • Where within an article they appear. ignoreSections excludes Checks/Suggestions from appearing within specific section headings and includeSections does the inverse, restricting a Check/Suggestion to only appear within the sections you define.
  • What kinds of pages and content they run on. ignoreDisambiguationPages makes it so a Check/Suggestion does not appear on any disambiguation pages, ignoreQuotedContent prevents them from appearing on text inside of quotation marks or blockquotes, and inCategory/notInCategory and hasTemplate/lacksTemplate scopes Checks/Suggestions based on the categories and templates present/absent within an article. Note: there's a task for enabling configuration based on properties of an article's talk page too.
  • How sensitive an individual Check is. These vary by Check/Suggestion. For Reference Check, minimumCharacters enables you to define how much new text someone will need to have added for it to get activated. For inference-based Checks/Suggestions like link suggestions, predictionThreshold enables you to set confidence threshold the model must reach before anything is shown to someone.
  • Entirely new, locally-defined Checks. TextMatch lets you define entirely new Checks, locally using pattern matching, without any code changes from us.

At any point in time, you can visit Special:EditChecks to see the current set of Checks and Suggestions that are available and how they're configured. This page is updated automatically as code changes are merged. In the future, we'd like to add metrics so that you have even more visibility into how things are working and identify what might need adjustment.

Might there be ways you'd like to be able to configure Checks/Suggestions that we do not currently offer? Is there information that you'd like to see on Special:EditChecks that's not currently visible? More broadly, does how we're thinking about this line up with how you're thinking about it? We're open and eager to any and all feedback. It's important to us that you have the visibility into, and control over, this system to make it work for your wiki (and the same for volunteers at other wikis).

---

i. See the distinction between Edit Checks and Suggestions that Marshall posted above. PPelberg (WMF) (talk) 23:51, 19 August 2026 (UTC)

What kind of communication with communities has there been on this project so far?

In a June meeting hosted in the Wikimedia Discord, members of the Editing and Machine Learning Teams shared (discord link) that we were experimenting with LLM-generated MoS suggestions.

In that conversation, we mentioned that this initial experiment would not involve surfacing any actual edit suggestions to volunteers. Instead, we would show edit suggestions in an "experimental" state that invited the people who opted-into seeing them to offer feedback about the quality of the suggestions. Then after that feedback process, we (volunteers + staff) could decide whether any of them were worth actually showing to people via a controlled experiment.

In the time between that call and now, we shared progress updates on the MediaWiki project page and had planned to announce the existence of this work and invite feedback about it on-wiki this week or next.

I appreciate that the above still led us into a situation where many of y'all were caught off guard by this work and as a result, trying to put the pieces together in real-time. This is not ideal for y'all and it's not ideal for us!

In this thread we've seen folks like @Alpha3031, @Chaotic Enby, @SJ, and @Kowal2701 helpfully suggesting that work like this would be better off were we to be in touch about it early and often. This way, we can get on the same page about what ideas might be non-starters and for those that we do deem worthwhile to pursue, to align on when and how we'll evaluate them together.[1] [2] [3] [4]

This all leads me to wonder…

Let's imagine we (staff) have identified a new LLM-powered suggestion that we think would be worthwhile to experiment with. When would you all appreciate us checking in with you all about it? How much information/example/context do we need to prepare before it's useful for a large group to evaluate an idea? Where do you think would be the best place for us to post about this?

Asked as another way: if we could do this all over again, what would you imagine the "ideal" process to be? PPelberg (WMF) (talk) 23:08, 20 August 2026 (UTC)

Not sure this is a communication issue, so much as a content and quality issue. If a feature is good then communication about it will be received more positively no matter what form it takes, whereas if a feature is not good then there is no way to announce it in a way that will be received well. Gnomingstuff (talk) 03:22, 21 August 2026 (UTC)
To me, the ideal process would be asking the community first what LLM-powered suggestions they would be most interested in and would find acceptable. It might also avoid issues with setting up suggestions that may swerve into CTOPs and CTOP-adjacent spaces that you might not be aware of. There's plenty of low hanging fruit for LLM suggestions that I think would have a reasonable amount of support, but any time you're going to have an LLM suggest things, especially to new editors, that have resulted in arbitration cases and site bans you're probably looking at the wrong things. ScottishFinnishRadish (talk) 11:25, 21 August 2026 (UTC)
Hello Peter, agreed with SFR + Gnomingstuff that this is about quality and collaboration, and how to start with easy cases when implementing any new interface or workflow. Sharing examples as they are produced, and working in public on essential parameters like what eval is used and the threshold for what gets presented, would make it easy for people to give targeted feedback early and often, saving everyone time.
* Give more attention to the choice of areas that will be covered. Have a page for suggestions, with definitions slightly more detailed than a one-sentence description. (what is the scope of each, defined how, tested against which guidelines)
* One step in the review checklist should tease out what might be easy or hard or surprising in each area.
* Spend time on better evals. Have human evals along with LLM-judge evals of initial suggestions.
* Use a tool for bulk elicitation and evaluation of suggestions (like an editable spreadsheet); working through the VE interface is too slow for meaningful human review at scale, and when reviewers find a type of error they will want to find many instances of it to fully characterize what's going wrong.
* Fine tune the models used on feedback.
This doesn't need to be a large group; small groups working in public on a predictable cadence is fine. The above is helpful regardless of what is generating the suggestions (LLM or other tool). It shouldn't be 'checking in' ⸺ this is part of the editorial workflow! There's no lack of interest in finding low-hanging fruit that works. I propose a better starting premise would be "we (all) have identified a suggestion that may be worthwhile to experiment with"... which is possible once there is an active page for suggestions. – SJ + 15:37, 21 August 2026 (UTC)
Seconding all of these Gnomingstuff (talk) 06:14, 22 August 2026 (UTC)
Also, you might let people test + give feedback on the interface separately from particular suggestions. For instance, using this to highlight {cn} instances on a page or other inline flags that are already present in the wikitext but harder to discover. That could happen continuously while studying low-hanging suggestions above. – SJ + 15:37, 21 August 2026 (UTC)
@ScottishFinnishRadish and @Sj: thank you for offering these clear recommendations for how we might work more effectively together going forward.[i] And thank you Gnomingstuff for stating clearly that no amount of communication substitutes for the underlying quality of the work.
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion. Until then, thank you for the thought and attention you've offered this week...we continue to learn a great deal in the process!
---
i. E.g. create a local project page where we can generate and evaluate ideas for new suggestions together, make evaluation easier to do at-scale, avoid contentious topics when considering suggestions to pursue, etc. PPelberg (WMF) (talk) 23:25, 21 August 2026 (UTC)
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion.
Hi y'all – there are at least two ways the work will proceed from here:
First, on the current batch of MoS Suggestions: we’re going back through this discussion to compile a list of issues that we can investigate and hopefully, use to improve the MOS:GEO and Simplify Language suggestions (spreadsheet).
Second, on process: we're going to draft a proposal for where/how we can evaluate ideas for new suggestions together (e.g. copyediting) and metrics we can use to assess the quality of suggestions once a dataset is available.
You can expect another message to Village Pump as soon as we have updates on the above.
In the meantime, thank you all for this helpful feedback. PPelberg (WMF) (talk) 18:25, 26 August 2026 (UTC)
– SJ + 13:22, 27 August 2026 (UTC)
Re: communication, I think it depends on the context. Is it something suggested by several editors on-wiki? Is it very similar to an existing feature? Is it similar to an existing user script? Go for it. Have similar ideas been controversial in the past? Is it a completely new approach to something? More discussion with the community as you design and develop it would probably be good. If in doubt, just drop a "hey, what do you guys think about an AI tool to detect NPOV issues?" on a noticeboard somewhere. It doesn't need to be some elaborate presentation. (I know big communities and corporate cultures can both make some of those things hard, but aspirationally I wish we had a culture—both here and at the WMF—that made that not a big deal and encouraged the communication.)
The best place is on wiki, either on this page or (for more specific features) on a relevant talk page. This is where we all are, and the same can't really be said for anywhere else. (Disclaimer that I can't speak for other projects.) And if you're not sure, it's generally pretty easy to ask "where is a good place for this" or just to drop a link on a few pages. LittlePuppers (talk) 03:50, 22 August 2026 (UTC)
I'd advise against relying on suggested by several editors on-wiki. I'm sure you could find more than several who would support any and all LLM implementation across the board. "Consensus between multiple experienced editors" might be better phrasing, but I think similarly to existing tools that have general acceptance is a safer metric for proceeding without additional discussion. Anything that's novel should probably receive community input. ChompyTheGogoat (talk) 21:09, 27 August 2026 (UTC)

Possible RfC

Given the WMF seems inclined to continue developing these tools regardless of what we say, I propose we hold an RfC establishing that LLM tools may not be deployed without an explicit consensus from the community. If editors think this may be a good idea, I suggest the following as the first draft of the question:

Should we require an explicit consensus before the WMF is permitted to deploy tools involving the use of an LLM? The WMF may deploy tools for user testing, so long as all of the following criteria are met:

A: Testing of the tool concludes after no more than 12 months, after which the tool must be removed unless there is a community consensus for either extended testing or full deployment.
B: Use of the tool is restricted to editors who have opted-in
C: Use of the tool is restricted to editors who have been granted the LLM-tool user right. This would be a new user right created and granted by admins at WP:PERM.

Does anyone see issues with this proposed question? Are there any revisions? BilledMammal (talk) 07:42, 21 August 2026 (UTC)

I'm not fundamentally opposed to this, but I think we need a lot more clarity about what this new LLM-tool user right would mean. A naive reading of the name would be "This user is exempt from WP:LLM" which I suspect is a broader interpretation than you intended. RoySmith (talk) 12:53, 21 August 2026 (UTC)
This seems like policy creep, and not a good approach (we shouldn't have blanket limitations based on the "type of tool" involved). I believe it is also misdirected: the example above is testing what would be an entirely community-configurable tool. Opt-in is appropriate for things that are so new, and I'd say anything that messes with your margins should be configurable. But proliferating user rights is an anti-pattern, and should only be used as a last resort. – SJ + 16:21, 21 August 2026 (UTC)
yeah, I think this may be premature for now, but worth revisiting if there are any disasters Kowal2701 (talk, contribs) 18:24, 22 August 2026 (UTC)
Consistent with the current wording of NOLLM and with active work being done - so far successfully - at helping limit abuse, I would suggest some cleaving of administrative actions and content actions. Best, Barkeep49 (talk) 17:05, 21 August 2026 (UTC)
Solely regarding access control (I'm undecided regarding the overall proposal): I agree with the other commenters that that creating a new user right isn't the best fit. I think tool-specific JSON lists of approved users that are fully protected could suffice, and be more adaptable to allow for per-tool authorization. isaacl (talk) 17:15, 21 August 2026 (UTC)
The whitelist system works well for AWB/JWB. Certes (talk) 21:13, 27 August 2026 (UTC)
I also think that the community should be able to control and configure the tools used in en wiki but I'm not sure an rfc is needed atm and its wording unfortunately isn't clear.
What is a "tool"? What else, other than edit suggestions, is in the scope?
12 months deadline is arbitrary: it's too long for a tool that the community actively opposes and may be too short for an iterated testing of a complex idea.
Some abuse/vandalism prevention tools by definition cannot be opt in. Alaexis¿question? 20:18, 21 August 2026 (UTC)
It's hard to guess what is a tool, because the AI is developed secretly and added to Wikipedia's software or configuration as a fait accompli. Certes (talk) 21:13, 27 August 2026 (UTC)
I agree with most of the comments above that this is not a well-defined RfC. In addition, the WMF have said this will be community-configurable if it makes it into production so it appears to be unnecessary anyway. Mike Christie (talk - contribs - library) 18:34, 22 August 2026 (UTC)
I assume this is intended to apply to any future implementation that would integrate LLM features in any way, not just this one concept, which I fully support. Pushing it on us without community consent is inappropriate and exactly the type of forced AI we're seeing on every other platform. Wikipedia is supposed to be different.
I'd agree on amending the deadline - I'm not sure what an appropriate initial test phase would look like, but I'd recommend a limit aligned with that and only extended or fully implemented via consensus. Just enough to get a feel for it, not a full beta phase to polish it for release. People with more programming experience than me would probably have a better notion of the timeline.
I do think we need some limitation on access (especially if it is fully implemented) and I think Isaacl's suggestion sounds like a better way to handle it. ChompyTheGogoat (talk) 20:58, 27 August 2026 (UTC)
I think that some level of access control is warranted, but the idea of getting individual approval for every tool that the WMF wants to test would impede testing numbers without much safety gain over a single user list for all beta tools.

Having a time limit on testing is strange, but my exact opinion on it depends on how community consensus is defined here. I think something that both mitigates the issue i think you're trying to solve while not dictating the WMF's development calendar would be something like "Any tool in testing shall be disabled if a consensus is formed at the tool's thread on VPWMF that the tool is causing disruption. At which point the tool will only be re-enabled with consensus." Although, if we got anything near consensus that a tool in testing is disruptive, I think the team working on the tool would disable/fix it very quickly. These are people who chose to work on mediawiki here, not WMF management. MetalBreaksAndBends (One for all) 01:32, 28 August 2026 (UTC)
I think it's only suggesting permissions for LLM tools that would potentially be more prone to abuse, not for any and all in testing. ChompyTheGogoat (talk) 07:13, 28 August 2026 (UTC)
There isn't any limiting language in the comment, so I (and most people, probably) would think it would apply to all LLM tools. MetalBreaksAndBends (One for all) 23:42, 28 August 2026 (UTC)
That is what I meant - all LLM, not all tools period. If we end up with so many LLM tools being tested that they're hard to keep track of I'd consider that a problem unto itself. ChompyTheGogoat (talk) 02:50, 29 August 2026 (UTC)
I'm not saying the issue is that they could be hard to keep track of (though over a long enough time period it could), I'm saying that applying for every beta is an unnecessary hassle. MetalBreaksAndBends (One for all) 03:38, 29 August 2026 (UTC)
I think I would share the view that a full, formal RFC would be unnecessary if prior discussion shows clear consensus to implement, for example (assuming it's advertised to a noticeboard like this one or the cleanup wikiproject). I would expect any community members participating in such discussions to be sufficiently in touch with current attitudes towards LLM tools to bring up anything that would seem potentially controversial and standard pre-RFC discussion processes should function fine (w formal closures and move to full RFC if no clear consensus or if there is consensus there should be an RFC in the specific discussion). Alpha3031 (t • c) 04:05, 29 August 2026 (UTC)
Do you mean a full rfc for each tool or an RFC for this proposal? MetalBreaksAndBends (One for all) 06:25, 29 August 2026 (UTC)

Hey WMF: nice job on communicating in this section

I think this conversation has been healthier than some other recent enwiki–WMF conversations. Just wanted to say thanks to the WMFers that are participating here and doing things like compromising (i.e. getting rid of the NPOV model), replying a lot instead of making one polished statement then leaving, and speaking clearly and honestly.

It's tough because on some issues, WMF and enwiki are very out of sync. So even if everyone does everything right, these conversations may still be tough. But doing things like compromising, having conversations with us that aren't just statements, and speaking clearly and honestly are definitely a good approach. Please keep it up. –Novem Linguae (talk) 19:23, 22 August 2026 (UTC)

Absolutely seconding this! Really happy to see WMF folks take community feedback into account. I understand the task can be much harder than it seems at first, especially as these discussions get sprawling and it can be hard to find a thread that unites the whole range of community opinions together, let alone incorporate it in the team's plans for the project.
The way you managed to navigate it was a very positive surprise, and I'm looking forward to more productive exchanges from both sides! Chaotic Enby (in solidarity · talk · contribs) 19:45, 22 August 2026 (UTC)
Yes, absolutely. Ymblanter (talk) 08:10, 23 August 2026 (UTC)
Absolutely agreed on all of this Gnomingstuff (talk) 19:01, 23 August 2026 (UTC)
+1. Can't speak for anyone else, but pretty much all my complaints about WMF's poor communication are limited to the Trustees and a few Officers. Everyone else at the WMF seems to communicate fine. MMiller and PPelberg are two usernames (among others) I've grown very accustomed to seeing regularly on-wiki, they've communicated often and effectively with volunteers for years. Levivich (talk) 20:22, 23 August 2026 (UTC)
Thank you -- we're glad to hear it -- we are trying hard! Although there are going to keep being times that we all disagree on ideas, the most important thing is that we can discuss constructively to figure out how to best improve/adapt the wikis. Thank you all for being here for these conversations, on top of doing all your usual wiki work. MMiller (WMF) (talk) 05:20, 24 August 2026 (UTC)
I am of the opinion that most -- maybe all -- WMF employees who are not in top management are competent, helpful, want to do the right thing, and are eager to communicate. I attribute the stonewalling we often see with a perfectly reasonable fear that actually having a dialog with the volunteers will never help your career and just might get you fired. If only there was some way that WMF workers could organize and join an entity that works to protect them from being unjustly fired... Let me know if anyone has ever heard about something like that. --Guy Macon (talk) 12:38, 24 August 2026 (UTC)
I find this post amusing. But as a Wikipedian, I can't help but be myself and note that if you define top management as something beyond "C Suite" Marshall is would probably be considered top management. Best, Barkeep49 (talk) 14:47, 24 August 2026 (UTC)
+Whatever we’re on In solidarity Wikipedian12512 (Talking is fine | contribs) 01:05, 6 September 2026 (UTC)

Wishlist 2027, invitation to share feedback

Hi, I’m Sonja and I lead some of the teams at the Foundation who will be responsible for picking up wish work under the new wishlist process. As you may know, the Community Wishlist started out as an annual process through which Wikimedia contributors submit and vote on technical improvements they would like the Wikimedia Foundation to work on. The main goal of it is and has been to improve the editing experience by making changes and features the community asks for specifically.

In recent years, the process behind the wishlist has changed, and we’ve heard from many of you that it no longer meets many community members’ needs. So now, the Foundation is designing a new process with the community to improve how wishes are triaged, voted on, and prioritized in a way that is transparent, balanced across project families and language editions, and takes into account what the Foundation can deliver.

I would like to get community input specifically on these three stages of the Wishlist process:

  • The triage stage, meaning how wishes are fleshed out, organized and filtered prior to voting
    • We recommend to have a working group, including volunteers from various wikis and Wikimedia Foundation staff to work through this together
  • The voting stage, including who may vote and how votes are structured
  • The post-vote stage, including how to bring equity into what work is prioritized
    • One way to do this is to rank wishes within 3 categories: Large Wikipedias or covering all wikis, small and medium-sized Wikipedias, and sister projects, so that top-voted wishes from smaller projects also get attention

This message is an abstract of the full proposed process. As you read the proposed ideas on Meta, please speak up about whether you think this is working well or if there are ways to make it stronger. Regarding the timeline, this consultation is open for two weeks. You can post your feedback on Meta or in response below.

For this year’s cycle, we plan to have the wish submission period in late October/early November and the triage process completed by late November. To respect the end-of-year holiday season, voting would happen in early to mid January. This first voting cycle is meant as a first step to try out a new process, and there will be more opportunities to provide feedback along the way, so that we can figure out the best process for future years together. SPerry-WMF (talk) 17:41, 27 August 2026 (UTC)

Is there an overview that contains a list of all the things that made the wishlists and whether they were implemented? Ideally, such a list would break down whether an unimplemented wish is [A] something the WMF still wants to do but hasn't done, [B] something that is possible but the WMF decided not to do it, and [C] things that the community wished for that are impossible.
That is assuming, of course, that the WMF can identify when something is impossible. Remember when the Flow team promised us "No edit conflicts, ever" and I pointed out that this was formally proven to be impossible? (See Wikipedia talk:Flow/Archive 4#No edit conflicts? and Wikipedia talk:Flow/Archive 4#Brewer’s Conjecture.) Good times... --Guy Macon (talk) 17:08, 28 August 2026 (UTC)
Results from previous surveys can be found on the respective survey results page. We currently don't have a running document of all wishes and those explanations, but we can keep this suggestion in mind moving forward. One problem with the full summary you're suggesting is that status labels have changed over the years, and with them decline reasons. We're trying to create a clearer set of statuses and are actively discussing how declining of wishes should be handled. Do you have any thoughts on that? SPerry-WMF (talk) 22:41, 28 August 2026 (UTC)
I have a specific comment and some general comments.
Specific to what you are working on, It seems to me that someone sitting down for an afternoon could translate the rejection reasons or at least comment on those old wishes so as to make the history helpful. What I am thinking is that when you efficiently solve a problem it goes away and nobody talks about it, but when something doesn't get immediately solved it remains an annoyance. This gives a false impression about how effective the team is by making a lot of the good stuff invisible. The list I described gives both equal visibility.
My general comment is about the nature of wishlists, based upon decades of addressing similar issues in industry.
Consider two wishes. They both seem to be roughly equal to the people making them but actually have wildly different difficulty. See [ https://xkcd.com/1425/ ] as an example. Maybe wish A is 5% more popular than wish B but takes a thousand times more work to solve. You need to figure out how to do the less popular thing that takes a few hours first, and somehow communicate to those making the wishes that some things that look easy are hard and some things that look hard are easy.
Or consider an example I gave before: The Abcom word counting template doesn't actually count words correctly. This has no effect on most users but the users it hits get hit hard while in the middle of an already stressful situation. The Arbs and clerks can't fix this -- they aren't developers and don't have the skills -- so they bodge up some crappy workarounds like saying "the clerks will eyeball the page and ignore the bogus word count". Your team could fix this in an hour or two. The most junior developer is able to do a reasonable job of counting words. But it will never, ever, get to the top on any community wishlist because most people never see the bug.
You should spend a significant percentage of your resources fixing these small, easy to fix problems that have been annoying users for years instead of spending most of your effort on big, sexy. popular, and exciting things and leaving a huge mountain of quality-of-life technical debt in the hands of unpaid volunteers. --Guy Macon (talk) 15:14, 29 August 2026 (UTC)
When deciding what to work on next, there's a lot of factors. One is certainly how hard it is. But another is how many people the issue affects. Another is how much pain does the problem cause. I see word-counting arbcom statements as being pretty low on both the "how many people" and "how painful" scales and thus a perfect project for somebody to solve with a user script.
There's also the "how much collateral benefit will this bring?" scale. When doing work planning, it's not uncommon to look at something and say, "The user-facing benefit of doing this isn't huge by itself, but the work involved to do that will also result in refactoring this other gnarly thing which is a blocker for five other projects in our backlog, so it's worth doing". We, as users, typically have very little visibility into that aspect.
There's also the question of who's familiar with the code. Sometimes you go around the room and somebody says, "I'm all over that part of the system and know exactly what has to happen to implement this so I can knock it off in an afternoon". Sometimes you get a bunch of blanks stares and people muttering, "I didn't even know that existed; it'll take me a couple of days of exploration before I can venture an estimate of how much work it'll be".
Guy, looking at your userboxes, I would expect nothing I've said here will come as a surprise to you. It's cool that you know C and assembler and Forth. Me too, on all counts. RoySmith (talk) 15:43, 29 August 2026 (UTC)
Thanks for your work on this. I'm glad there's a plan to have an iteration of the wishlist basically this year.
For this year’s cycle, we plan to have the wish submission period in late October/early November and the triage process completed by late November. To respect the end-of-year holiday season, voting would happen in early to mid January. I believe older wishlists allowed the creation of wishes and then voting on the wishes simultaneously. That is, in old wishlists you could create the wish during the voting period.
For this iteration of the wishlist, it sounds a bit like you are proposing a system where we can only create wishes before November, then a committee pontentially vetoes them, then we have to wait 2 months before we can vote on them? If this is the plan, then I think adding so much time to the cycle is not a good idea. The annual cadence of old wishlists got people to focus on the wishlist for a couple weeks -- this new system would require folks to focus on it for a couple months (create a wish at the proper time, wait 2 months, then market the wish so that it gets votes). I think it is really important to shorten the cycle in order to keep things nimble and to avoid bureaucracy. Perhaps all triaging should occur during and after the voting is completed, so that folks can easily create wishes during the voting period. –Novem Linguae (talk) 08:32, 29 August 2026 (UTC)
believe older wishlists allowed the creation of wishes and then voting on the wishes simultaneously. That is, in old wishlists you could create the wish during the voting period. i dont remember this. We had vote creation combined with the community feedback phase, and some people would vote before the vote had started, because for many people it was difficult to understand what they were supposed to do. —TheDJ (talk • contribs) 17:49, 29 August 2026 (UTC)
Yes, historically it was intended that there'd be a "filing and discussing wishes" phase and then a "discussing and voting on wishes" phase. Early voting happened but was various levels of discouraged. AntiCompositeNumber (they/them) (talk) 19:21, 29 August 2026 (UTC)
This is all really helpful, thank you for weighing in.
The timeline we picked was meant like this: wishes can be submitted anytime between now and early November, but we'd run banners for the last 2 weeks of the submission window to get people to participate. Then we'd do triage for 2 weeks on all the wishes, meaning the working group would look through wishes, size them, ask anyone who is following the wish questions and discuss and document tradeoffs, and in some cases where it would be very clear that the wish would not be possible to be picked up (see potential reasons on under Triage activities and discussion with new ideas and perspectives on the Talk page), the wish might be declined. Then we'd have a voting period, widely advertised with banners, in January. After that we'd review the vote count to see if anything should be adjusted, for example to bring in a top voted wish from a smaller project, and we'd share the list we'd plan to work on with the community, again with an invitation for feedback (open for a week or two), before we start implemetation.
The reason for spacing it out like that is merely logistics: closing wish submission makes it much easier to triage wishes, because you don't have an ever growing list, and having the voting period in January was simply to respect the end of year holiday season. We're also a bit in a time crunch this time around, because we want to have a prioritized list by early February at the latest, so that we can start including them in our roadmaps for the current fiscal year (meaning between Feb and June). That part will be different in future wish years, because wishes voted on in January would not get worked on until the start of the next fiscal year in July. Aligning the vote with our annual planning process ensures that we free up the necessary resources to deliver on wishes alongside other work we have planned. This year is different, because we want to action wishes as quickly as possible under the new process. Aside from picking up wishes pretty much immediately after the vote, we will also bring them to the table when we start our planning process for the 27/28 fiscal year in February, so this round will feed wishes into this and next fiscal year.
With regards to not closing wish submission until we close the voting period: that makes sense, and you make a good argument to keep the periods closer together. Let me think that through a bit and come back to you with an alternate timeline early next week. SPerry-WMF (talk) 23:00, 29 August 2026 (UTC)
Thanks for the detailed response and for taking the feedback onboard.
If you keep the phase system (submitting, then triaging, then voting), it could make sense to schedule it so it doesn't intersect with the December holiday season. Then less of a break would be needed in the middle, and the gap between submitting and voting could be shortened.
I think you all are envisioning the wishlist being independent of the annual plan. If that turns out to be true, it can be scheduled anytime. But assuming that it does need to be scheduled before annual planning season, perhaps Oct-Nov (before December) would be a good time period to shift it to. –Novem Linguae (talk) 10:13, 30 August 2026 (UTC)
Those are good points – I think for future years it would make a lot of sense to move the submission and voting period to January and February, which would help us keep the timing between those periods tighter and it would reduce the time between vote and wish implementation as well. We specifically want to bring the results of the vote to our annual planning process, so that we can ensure wish work gets a dedicated spot in our roadmaps for the next fiscal year.
I took another look at this year’s timeline and still think it's best to keep the plan as-is, because we want to pick up tickets as early as February in this first cycle. Carrying out this entire new process in January/February instead of doing the submission and triage in Oct/Nov would push that back. Plus we don’t want to pull the vote to December, because we want everyone to get an equal chance to participate, but we know that a lot of people take a break then.
Also, there’s a similar discussion happening on the proposal Talk page in case you want to follow along there. SPerry-WMF (talk) 21:22, 1 September 2026 (UTC)
@SPerry-WMF how does the wishlist interact with feature request tickets opened in phab? I usually just go the phab route. Do these just get lumped into one pool to be evaluated? Is one the preferred process over the other? RoySmith (talk) 13:34, 29 August 2026 (UTC)
See phabricator as a permanent list of all that could potentially be done, but no promise of anyone even looking at it. Its also generally more technical. Wishlist is a subselection of that same list, but allows community voting, and comes with a promise of people actually evaluating what was filed. —TheDJ (talk • contribs) 17:36, 29 August 2026 (UTC)
A well-functioning Wishlist will be able to get top wishes prioritized, with WMF teams assigned to work on them. On Phab, whether something gets prioritized is completely up to the team or volunteer devs that maintain the software. Also, Phab skews towards existing software rather than new software. As for which system a community member should use, maybe start by creating a Phab ticket, and then if the issue is important AND not getting worked on, also file a wish. Although this "strategy" part is subjective so is up to you. –Novem Linguae (talk) 10:21, 30 August 2026 (UTC)
  • I have a proposal: I propose that multiple people who read this and agree with me place/support a wish on the wish list that reads something like this:
"Knock down our technical debt by fixing small quality-of-life issues, prioritizing things that are easy to fix. Don't hold back because fewer people are affected, because some unpaid volunteer supposedly maintains a script, or for any other reason. If it's wrong and you can fix it in a short amount of time, fix the bug wherever it resides."
Regarding "fewer people are affected" I have seen this as an excuse for not fixing an arbcom script that doesn't count words correctly (few people are the subject of an arbcom case) and as an excuse for not fixing a nasty accessibility bug (few Wikipedia editor are blind).
Probably best if someone else makes the wish. I have a number of people who hate me because I suggested that the WMF stop pointing a giant money hose at their pet projects. --Guy Macon (talk) 18:43, 4 September 2026 (UTC)
@SPerry-WMF, when you were looking at wish "size", how small was "small"? Might there be room for "extra small" wishes? In solidarity, asilvering (talk) 20:31, 4 September 2026 (UTC)
I would consider the smallest possible "fix a bug" job to be something that you give to a developer and they say that they can completely finish the job including documentation in three hours or less.
(I don't trust 15-minute fixes on anything someone else is supposed to use. Too many times the fix adds a new bug. I would say that devoting maybe an hour to testing is reasonable for the smallest, easiest bug. More if you do regression testing).
When you apply the standard multiplier (start by multiplying every estimate any developer gives you by Pi and then look for reasons it might take longer) you can hope to get it knocked out in about a day.
If they say they really can knock them out quicker that that (you might have managed to hire the next Charles H. Moore or Margaret Hamilton), have them do ten have someone else check them for errors, and you will have a better estimate than I could come up with.
The key is picking obvious quality of life issues that nobody is even thinking of fixing. Here is another example:
Without checking, tell me what happens if I sign this post with 1 tilde, 2 tildes, etc. up to 12. Can you? I can't without checking my notes. How long would it take to fix the stupidity that results from something that is 99% likely to be a typo? It will take some amount of time to decide how to fix it. Automatically turn everything from 3 to 12 into four tildes, screwing with the tiny percentage of users that want to add just a name or just a date? Tack on an "are you sure" and "never warn me about this again" dialog (my preferred fix)?
Just for practice in the proper method of fixing stupidities like this, refrain from telling me about all the cool ways that I can abandon the workflow I have been using for years and switch to a method that doesn't require the tildes. Dumping the bug back on the user and asking them to use a workaround might be the only answer you have, but it should never be your first choice. --00:13, 5 September 2026 (UTC)Guy Macon (talk)
The sizing we use currently is as follows:
  • S - not complex at all
  • M - mostly clear, but some complexity, solvable by one team
  • L - some complexity, potentially requiring help from multiple teams
  • XL - very complex, likely requiring multiple teams
There is definitely room to add an XS.
@Guy Macon: I think that's a fair suggestion, but what would make it really actionable is a list of phab tickets that you'd prioritize. I worry that with a wish formulated like that, it could be anything and everything and the impact is in the eye of the beholder. Plus we wouldn't ever be done with the wish, because there are so many tickets going back years, so one year wouldn't be enough to get through them all, even if we could put a full team behind it. And that's the other aspect that makes your proposal difficult to execute: wishes will be done by the teams who are best equipped to fulfill it. For a wish as you describe it, it would be pretty much every product team at the Foundation. What would get you more success is to go by general area or tool: Have one wish that lists tickets you want closed that address mobile web editing, one that fixes all issues you might see with the Watchlist, etc. That way, people voting on your wish would know exactly what all small changes/fixes you are referring to and you'd give them a tangible result for what would happen if they cast their vote on this and it would be worked on. SPerry-WMF (talk) 16:35, 11 September 2026 (UTC)
Too much work, needs a full team to get done? So don't get done. Put one developer on it for three days, then next week put another developer on it for another three days. The important thing is steady progress instead of doing nothing to reduce technical debt.
Asking an unpaid volunteer to help you to decide what to work on? That's the kind of thinking that got you so deeply in technical debt. When something isn't working, don't do it harder. Let the developer decide what to work on. Make the only instruction to the developer "do a bunch of small things that you can knock out fast." Don't worry about importance, impact, or who "owns" the problem (fix it whether it is in the code you maintain or some script some volunteer wrote), or anything else. Just start someone plugging away at it instead of doing nothing. If it becomes obvious that the wrong things are being done we can address that later. If for some reason you decide that everything needs to be fixed in a year (you have been fine with zero effort to address small technical debt issues for decades) we can address that later.
Put up a simple web page where you list what got fixed, what you gave up on because it didn't turn out to be as easy as you thought, and things you are working on or thinking of working on. Let anyone make suggestions (most of which will be useless), but let the developer decide what to fix. They are in the best position to know what they can fix quickly. Keep it simple, Do something instead of doing nothing. --Guy Macon (talk) 17:35, 11 September 2026 (UTC)
"Put one developer on it for three days, then next week put another developer on it for another three days." That's not how software is written or fixed. I doubt many of the tickets could be resolved by one person working for three days. But imagine, like, suggesting that a novel be written by having one author work on it for 3 days one week, another author continue working on it for 3 days the next week, etc. Patchwork code writing like this is a bad idea. (It's probably how a lot of the bugs and lack of future-proofing and such got in the code in the first place: too many hands stirring the soup.) Levivich (talk) 17:56, 11 September 2026 (UTC)
I am intimately familiar with the way software is written and fixed at Boeing, Airbus, Mattel, Perkin Elmer, Parker-Hannifin, NASA, and a number of smaller companies that you have never heard of. What you say is perfectly valid for small to large software projects. It is absolutely not true for the kind of tiny, easy-to-fix quality of life issues I am talking about, and indeed most of the companies I just mentioned have someone quietly chipping away at the kind of small technical issues that "proper" software development has trouble dealing with. Like a word counting script, written by a volunteer who hasn't logged on in years, and which fails to accurately count words. That's about a three hour job including testing and documentation. --Guy Macon (talk) 18:24, 11 September 2026 (UTC)
We're making decisions as part of the Foundation's annual plan. The wishlist is specifically meant to complement that, meaning the decision on what should get prioritized as part of wish work should sit with the community and will be voiced via vote count. Making clear what is and is not part of a wish is an important factor in that decision making process. So yes, in this case asking volunteers to decide on what to work on for a wish is exactly what we should do. SPerry-WMF (talk) 18:59, 11 September 2026 (UTC)
I am going to withdraw from this discussion now, convinced that you completely failed to understand what I am asking, and assuming that the fault is on my end and that somehow I am incapable of communicating. Maybe someone else will have more sucess than I did.
What you are talking about is what you are doing now. I have no reason to think that what you are doing now isn't being done well. It's as if I asked someone "please take the garbage out. It will only take a minute." and they replied "I already have a plan to repair the leaky roof and getting three estimates is exactly what we should do." Not the same thing. You will no doubt end up accomplishing many new and useful things while making zero progress on the easily-fixed technical debt issues I am talking about. And everyone who participates in any Arbcom case will still have to deal with a word count tool that doesn't count words, sanctions if they go over, and clerks who eyeball the word counts and make estimates as a workaround for a shitty word-count tool that you could have fixed in a few hours.
So here is my final feedback: In my opinion you are so focused on doing what you are already doing quite well even better that you appear to be completely blind to the possibility of also doing something that is low effort unless you can approach it the same way Procrustes approached innkeeping. I am done. --Guy Macon (talk) 20:24, 11 September 2026 (UTC)
SPerry-WMF is possibly saying that their hands are tied and directions from above mean that fixing things won't be done. WMF bureaucracy probably has learned Microsoft's mantra that time spent fixing things gives their competitors time to develop more glitz that will steal a lead. Levivich is talking about different things—items larger than what you have in mind. Guy Macon is absolutely correct that having one person focus on one small problem of their choice is exactly the way to make quality progress. If that person needs help because it's more tricky, try two people. After that, think about what to do. Make progress on technical debt. Johnuniq (talk) 03:20, 12 September 2026 (UTC)
Thank you to everyone who engaged in this discussion here and on meta. I’ve processed all the feedback we've received and made changes accordingly. We will use this updated version for the Community Wishlist this year, with the intention to further improve the process next year. In the meantime, I look forward to hearing from you all during the submission and voting periods. SPerry-WMF (talk) 03:52, 15 September 2026 (UTC)

A new village pump page?

I think it would be a good idea to have a Wikipedia:Village Pump (affiliates). Sometimes affiliate related issues are shared here because there's some overlap, but I think it could be its own topic. I think it could also help people provide feedback on projects and it might encourage people from affiliates to share updates. I'll create it if other people think this is a good idea as well. Clovermoss🍀 (talk) 10:25, 8 September 2026 (UTC)

I'm wary of adding yet more Village pump pages, and I've not seen much of a problem from them being posted here or at other relevant village pumps. Have there been problems I've missed noticing? Or call from people for a dedicated page (versus "might encourage")? Anomie⚔ 11:41, 8 September 2026 (UTC)
Technically I'm a contractor for Wikimedia Canada and I think it's a good idea. But it's more of a might encourage thing, since I don't like that most of this communication takes place off-wiki in places like mailing lists/telegram. Encouraging more on-wiki communication promotes understanding from both sides and hopefully leads to people feeling better about some of their concerns since they have a place to go. Clovermoss🍀 (talk) 11:45, 8 September 2026 (UTC)
I also don't see the need for another VP page. If there was so much traffic about affiliates that it was clogging up the existing pages, it would make sense. If the people involved with affiliates prefer to use other channels, it's not our place to say, "No, we want you to have those discussions here". Also, this isn't an enwiki-specific thing. I would think someplace on meta (i.e. meta:Wikimedia movement affiliates) would be a better place. RoySmith (talk) 12:04, 8 September 2026 (UTC)
It's not a have thing, but an option. The people I talk to aren't trying to hide things and would like the community to like them. I think Meta isn't a bad idea, but a lot of things that affiliates do affect en-wiki specifically and it makes sense to have a page here for that as well. People applying for grants are also encouraged to seek feedback and those pages don't seem to have as much traffic as a specific page for affiliates and grants could. Clovermoss🍀 (talk) 12:11, 8 September 2026 (UTC)
My first reaction is to explicitly add affiliates to the scope of this page. If turns out that this increases the traffic so much that WMF-exclusive stuff gets drowned out then we can split at that point. Thryduulf (talk) 12:48, 8 September 2026 (UTC)
That makes sense to me. RoySmith (talk) 12:54, 8 September 2026 (UTC)
It was my thought too. Per Thryduulf on rethinking if WMF stuff gets drowned out. (Slightly reduced risk of a dedicated throw brickbat space.) CMD (talk) 12:57, 8 September 2026 (UTC)
I worry that might cause some friction, the way we tend to get upset when volunteer editors are conflated with the WMF. They're separate organizations, even if they're mission-aligned. I also envision encouraging a bunch of affiliate-related people to introduce themselves if such a page was started. Clovermoss🍀 (talk) 12:57, 8 September 2026 (UTC)

I'm trying to understand the objections here. Is the concern that the page would end up so low traffic as to be dead? It's hard to know that without even trying. Or is it more that the uncertainty isn't great when there's already a bunch of village pump pages? Clovermoss🍀 (talk) 14:36, 8 September 2026 (UTC)

onwiki issues related to affiliate related work, at least for my region, ESEAP, had been an once a year occurrence. Extending further out to the wider Asian region, anecdotally, had been one or two issues a year. I would prefer to utilise the misc village pump first before opening another venue. – robertsky (talk) 15:09, 8 September 2026 (UTC)
@Robertsky: your affiliate region wouldn't have any ongoing projects or updates to share more than once a year? I'm not trying to dismiss your perspective, I'm just having a difficult time trying to understand it. There's no contests or challenges or institutional partnerships? Clovermoss🍀 (talk) 15:13, 8 September 2026 (UTC)
You're a contractor and you want to create a page to benefit your contract? Alanscottwalker (talk) 15:20, 8 September 2026 (UTC)
No, it doesn't really have anything to do with that. I'll be doing what I've been doing regardless of whether this is created. Clovermoss🍀 (talk) 15:22, 8 September 2026 (UTC)
I thought you said you were a contactor of an affiliation and you thought this new page would be a good idea? Alanscottwalker (talk) 15:25, 8 September 2026 (UTC)
I'm a contractor, but my contract has to do with running events and trying to increase the reach of Wikimedia Canada. Almost all of that stuff is offline, although I have started working on a subpage for transparency's sake at User:Clovermoss/Wikimedia Canada. A lot of stuff that goes on at Wikimedia Canada happens without me, as they have dedicated full-time staff. I made the distinction because I am not that. I have had an interest in making affiliates less of a maze to people outside of them for awhile, though. See this essay. Clovermoss🍀 (talk) 15:29, 8 September 2026 (UTC)
The purpose of this new page would be to "increase the reach" of affiliates and "make affiliates less of a maze"? Alanscottwalker (talk) 15:37, 8 September 2026 (UTC)
I doubt it'd do much to increase the reach per se, as again, most people would be whatever they're doing without that page. My desire to create something like this comes from the years I've spent not being involved with affiliates as a "regular" volunteer and wanting to understand how they work. Most of this information is very difficult to find if you are not part of those circles. Clovermoss🍀 (talk) 15:39, 8 September 2026 (UTC)
How is it not increasing their reach when the page's purpose is to expose others to them? Alanscottwalker (talk) 15:50, 8 September 2026 (UTC)
Because when I'm expanding the reach, I'm doing that in-person with Canadians who have incredibly limited knowledge about Wikipedia and Wikimedia projects. What I'm proposing here has more to do with people having a place to go that doesn't involve travelling to a conference to talk to affiliate-related people and for affiliate-related people to ask for feedback from people who may have very different experiences then their own. Wikimedia Canada is only one organization of several, let alone all these groups. Your concern seems to be based along the lines of someone told me to do this and no one told me to do this. Clovermoss🍀 (talk) 15:59, 8 September 2026 (UTC)
No, I'm not concerned about what you are told. You have a contractual interest in an affiliate, and the purpose of this proposed page is for affiliates. Alanscottwalker (talk) 16:05, 8 September 2026 (UTC)
I guess, but that's a cynical way of looking at it. The most personal gain I might get out of this is brownie points, but it's more of a risk than a benefit because it probably wouldn't look that great on me if it made people more skeptical of affiliates instead of restoring some trust.
I'd be interested in this even if I had nothing to do with Wikimedia Canada (as was the case a few months ago) but obviously I can't force you to believe in my sincerity. I mainly brought it up because I was being perceived as an "outside" source and people seemed to have concerns about if I was making assumptions about what spaces people connected with affiliates might be interested in. Hence my technically response to Roy. Clovermoss🍀 (talk) 16:15, 8 September 2026 (UTC)
It is not cynical, it is standard concern. (see, WP:EXTERNALREL) Alanscottwalker (talk) 16:20, 8 September 2026 (UTC)
I know what external relationships are. It doesn't undermine the primary goal of improving the encyclopedia. These two things aren't in conflict with each other. So yes, I think your take is fairly cynical even if I won't tell you you can't have that opinion. The fact that people see affiliates as being an external relationship speaks volumes in its own right. I think a lot of people who are more connected with affiliates would find that sad and want to change that perception. But they can't do that if they don't know who they even have to convince. Clovermoss🍀 (talk) 16:23, 8 September 2026 (UTC)
Affliates are designed to be and have always been external because affiliation is not English Wikipedia's WP:PURPOSE. English Wikipedian's do not pay anyone to do anything, for anything, they are required to join no groups. I think you are undermining the encyclopedia, if you can't keep your external relationships, like your contracts, where they belong, not here. Alanscottwalker (talk) 16:43, 8 September 2026 (UTC)
If you seriously think that, please start a noticeboard thread instead of casting unfounded accusations about me undermining the encyclopedia. Clovermoss🍀 (talk) 17:56, 8 September 2026 (UTC)
It's founded in your contract, and in your refusal or inability to recognize that as an external relationship. Alanscottwalker (talk) 18:03, 8 September 2026 (UTC)
Would you say the same about someone who worked for the Wikimedia Foundation? Clovermoss🍀 (talk) 18:06, 8 September 2026 (UTC)
Apart from the limited office action, they are not to create things on this site, and they don't. Moreover, unlike any other corporation or group, they are the webhost and define the terms of use. Alanscottwalker (talk) 18:11, 8 September 2026 (UTC)
The Wikimedia Foundation creates things that affect the English Wikipedia, alongside other projects, all the time. As well as commenting enough for this noticeboard to exist. I don't understand why you think it's disruptive to want a place to talk to people, when it's completely harmless. To say that I'd be undermining the project for doing so is confusing if you're not just trolling me. But I'd like to think maybe you just don't understand what affiliates are intended to do. Clovermoss🍀 (talk) 18:17, 8 September 2026 (UTC)
Harmless? It is well recognized that editing with a conflict maybe harmless (it is also well recognized that it is not commentary on anyone's good faith or competence), it is equally well recognized that it needs to be avoided, and in talks, disclosed. Alanscottwalker (talk) 18:26, 8 September 2026 (UTC)
AGF Czarking0 (talk) 16:27, 8 September 2026 (UTC)
I appreciate the sentiment, but I can take a bit of pushback. I'd rather people voice their concerns than keep them to themselves. I don't want to dismiss people based on tone and cynicism alone because that drives me up a wall when I'm on the other end of it. Obviously I can't force other people to see my heart, but I hope it comes across. Clovermoss🍀 (talk) 16:31, 8 September 2026 (UTC)
Well, since that didn't end up working out and this thread appears like it came after the above if one isn't looking at the timestamps, I'll note that I did end up escalating this to a noticeboard thread. Clovermoss🍀 (talk) 19:06, 8 September 2026 (UTC)
While being a contractor might influence your perception of the ask, frankly it should not be much of a factor. I can see a potential for such a page to be a centralised location for communications between editors and the various affiliates who have activities that touches on this project. Not all editors are comfortable reaching out to affiliates offline, and affiliates do not have a monopoly of ideas of what activities to run for the project, and the editor in question might just need some support somewhere somehow to get their proposed activity off the ground.
That being said, while I am in the Singapore user group as well as the ESEAP Hub (as a steering committee member) and I see some merits in having one such village pump, I would like to reiterate that the we can potentially use the misc village pump first to see how the conversations play out, in terms of the frequency, breath and depth. If it gets too intense there, an affiliate village pump might then be created. – robertsky (talk) 16:05, 8 September 2026 (UTC)
Does not WP:Affiliates suggest they are all online?Alanscottwalker (talk) 16:16, 8 September 2026 (UTC)
@Robertsky: What about something like "an affiliate corner" in userspace for a few months (User:Clovermoss/Affiliate corner) just to see if there's enough to justify a "real" page? I think your idea has merit. If mine doesn't, everyone can just go back to Village Pump (Misc). If my idea actually ends up encouraging conversation, its text and history can be moved to a Village pump subpage for affiliates specifically. The main reason I want it all together is so people interested in those issues don't have to wade from unrelated conversations in archives to find them. Clovermoss🍀 (talk) 16:21, 8 September 2026 (UTC)
If you started it in Wikipedia space, I'm doubtful anyone would bring it to MfD. It just wouldn't be a Village Pump during that time. Best, Barkeep49 (talk) 16:42, 8 September 2026 (UTC)
Well, Wikipedia:Affiliate's corner it is then. Better than userspace, at least. I'll set it up later today when I'm less busy. Clovermoss🍀 (talk) 17:53, 8 September 2026 (UTC)
There's already Wikipedia:Meetup. It probably makes sense to have both of those link to the other. RoySmith (talk) 17:58, 8 September 2026 (UTC)
Yes, I agree. Not all meetups are organized/supported by affiliates, but some are, and it makes sense to link them together once the page is set up. Clovermoss🍀 (talk) 18:00, 8 September 2026 (UTC)
we have had wiki loves Ramadan this year that bubbled up onto here or ANI(?). Last year there was a translation project that went awry (contained at AfC). There may be another translation project soon, different group running. There are institutional partnerships, for enwiki specifically though, mostly would be the Australian, New Zealand, Singapore, and possibly Philippines that may have them. – robertsky (talk) 15:21, 8 September 2026 (UTC)
I think Hawkeye7 does some stuff with Wikimedia Australia. Tamsin and Giantflightlessbirds are who I think of when I think of Wikimedia New Zealand. There's also some Wikimedia NYC people I know of like Pacita (WikiNYC) and Pharos. Clovermoss🍀 (talk) 15:27, 8 September 2026 (UTC)
Tamsin and Giantflightlessbirds did m:GLAM/Wikifying a Conference recently, and produced a guidebook/book on the topic, which should actually help affiliates with offline outreach. 😏 – robertsky (talk) 15:47, 8 September 2026 (UTC)
I got a physical copy when I attended wikimania. :) Looks interesting, although I haven't got a chance to take more of a detailed look yet. But I wouldn't have really known any of the names I just pinged except for Hawkeye if I hadn't started attending conferences in the last 3 years and I'll never forget how completely obscure everything felt before then. Even now, there's still a lot of moving pieces I'm trying to make sense of. Clovermoss🍀 (talk) 16:02, 8 September 2026 (UTC)
(Just noting that Tamsin from NZ is actually DrThneed.) Giantflightlessbirds (talk) 20:10, 8 September 2026 (UTC)
Oops, thanks for the correction! Clovermoss🍀 (talk) 20:13, 8 September 2026 (UTC)

I have three at least partially competing reactions to this: a logistical reservation, a pessimistic view, and an optimistic view. Logistics: yet another discussion page? I'd like to see it grow in an existing discussion page first. And most affiliates aren't even primarily active on the enwp, so agree with the above that it seems more like a meta page? For those that affect enwp directly, perhaps creating a signpost series that solicits regular updates from affiliates is a good place to start? Pessimistically, it's hard not to draw a comparison to this board, which is functionally less about WMF sharing things and more about a place to evaluate and air grievances about the WMF and/or what it's doing. Putting aside the extent to which that's needed and/or useful, "do for affiliates what this page does for the WMF" doesn't seem appealing for affiliates to participate in. We should remember that while there are a few well resourced affiliates the overwhelming majority of affiliate members are volunteers. The optimistic view: it seems like a lot of people don't really have a good idea of what affiliates are, who participate, what they do, etc. On pages like this one, they're more like a line item in a budget. On Meta they're often lists of metrics. Having a space for "here's some neat stuff we've been doing at [affiliate] that you might be interested in" could actually connect the community with the good being done. Like how many people here know that Habst and some other folks at WikiNYC have been developing a neat tool called Wikinewsie that follows edit-based news (articles that are being edited a lot, or by a lot of people)? So yeah, mixed feelings, tending towards a "this is maybe a signpost thing". — Rhododendrites talk \\ 16:54, 8 September 2026 (UTC)

Perhaps, if someone wishes to write a signpost article or column ('what X group is doing'), but it is not hard to find what https://meta.wikimedia.org/wiki/Wikimedia_New_York_City is up to or who to contact. Some people may be confused that they are a separate corporation with their own mission and that volunteering for them is volunteering for that corporation, but it is not hard to find that out. Alanscottwalker (talk) 17:22, 8 September 2026 (UTC)
  • I admit skepticism, but I could be convinced. What exactly would we be talking about on an affiliate noticeboard? Do we have any examples of recent noticeboard conversations about affiliate stuff? As far as I know, affiliates aren't exactly in the business of writing or developing for EnWiki. They organize events, train people, try serve as centers of political power, channel money into... whatever it is they use all that money for, and run various competitions. But content? Tools? The sorts of things that affect our lives on EnWiki? I mean, for the most part, they leave us alone, and we leave them alone, and we're all the happier for it. Am I missing something about the work that affiliates do or that we should be talking about here? CaptainEek Edits Ho Cap'n!⚓ 22:39, 8 September 2026 (UTC)
    Well, certain affiliates like Wikimedia Deutschland take on technical projects that the WMF doesn't. WPMED does as well (Doc James would be able to talk more about some of what they do). WikiPortraits takes photos that are used quite frequently on the English Wikipedia (SuperHamster could give more accurate stats than I could off the top of my head; noting that I also have done work with WikiPortraits for the purpose of this conversation). There's a lot that could be discussed and talked about. I also think it'd be good just in general for both sides to understand each other better. Like, one of the most common volunteer concerns is when things like contests go wrong and result in massive clean-up efforts. I also see interesting things that happen on wikimedia-l that could give people a better impression of the "good" side. For example, this recent thread about cyclists. It's easier to build trust if people know what's going on. Clovermoss🍀 (talk) 23:48, 8 September 2026 (UTC)
    There's also a Commons user group , where doing stuff on-wiki is kind've the entire point. Some projects have much stronger overlap between affiliates/the general editing community. It's mainly North America that's the exception to that, where affiliate activity is more concentrated to specific areas. Wikimedia NYC does a lot of good work, though. Clovermoss🍀 (talk) 00:10, 9 September 2026 (UTC)
    @Clovermoss Admittedly, my regard for WPMED is low; you can see my thoughts about its value at . I don't disagree that Affiliates can do good for the movement, but I'm less certain they're doing EnWiki relevant work? Or that is to say, work that wouldn't be more relevant at say Meta. I do appreciate you giving some examples though. CaptainEek Edits Ho Cap'n!⚓ 04:14, 9 September 2026 (UTC)
    I agree it would be helpful to have some specific examples drawn from the past of what would be useful to have on a separate project space page. isaacl (talk) 01:28, 9 September 2026 (UTC)
    Another example is trying to understand what concerns affiliates have about the new proposed funding model from the WMF. Sohom Datta knows more about that from me. The stuff that's being proposed sounds reasonable to me at a glance, but I don't know what people's objections are, just that they have them. A place where people can share in their concerns, hear other people's perspectives etc can be really valuable.
    Another reason is that affiliates do have some forms of power in the movement. For example, affiliates were given the ability to shortlist which candidates were able to make it to board elections last year. Bluerasberry can probably give more examples than I can as they've been immersed in that world for longer. Affiliates can also go hand-in-hand with WikiProjects. For example, Wikimedia LGBT and WikiProject LGBT. Clovermoss🍀 (talk) 01:33, 9 September 2026 (UTC)
    Are you suggesting that affiliates are interested in posting on English Wikipedia about their concerns with the proposed WMF funding model, and about what candidates they are approving to proceed to the board elections? (I might not have been clear regarding what I meant by examples drawn from the past: are there past messages that were posted elsewhere on English Wikipedia that could, in future, be posted in the new proposed venue?) isaacl (talk) 02:20, 9 September 2026 (UTC)
    Yes to the first, uncertain as to the second unless we're talking about cleanup threads that happen at ANI and the Wikipedia:Education noticeboard. A place for good things might help people feel like there's less of a disconnect in priorities.
    Most of the communication I'm hoping for to take place exists in places that are not on-wiki, like an affiliate's website, offwiki Telegram groups, in-person, Whatsapp, etc. I'm just hoping to convince some people who are amenable to the idea to participate on on-wiki discussions if they wish to, but I want to emphasize the importance of optional and respecting people's autonomy to not participate if they do not want to. Clovermoss🍀 (talk) 02:40, 9 September 2026 (UTC)
    They're your examples, so you can tell me what kinds of messages you think they would cover. Your second example was "affiliates were given the ability to shortlist which candidates were able to make it to board elections last year," so I'm not sure what you mean by cleanup threads.
    I'm not sure why an affiliate holding conversations on their own web sites or various messaging apps would decide that a better replacement would be a page on English Wikipedia. I do appreciate that an affiliate who wants to hear more from the broader communities would want to hold discussions on wiki. But as Rhododendrites said, affiliates are largely composed of community members. I think most of them should be able to find places to start conversations on wiki. isaacl (talk) 02:56, 9 September 2026 (UTC)
    I'm not trying to replace anything, but offer a supplemental place to people who want it. I disagree with Rhododenrdrites view, but I don't want to single out people who don't have connections with any community, but plenty of people like that do exist.
    This isn't a longstanding issue of contention for no reason, even if it's clear people have had very different experiences with different affiliates. I'm glad Rhododendrites has only ever seen that overlap, as there tends to be less issues when everyone is on the same page, hence why I wish for there to be a place for people to communicate more often.
    I've been giving lots of examples of what a page like this could cover. I don't think these issues contradict each other, they're just examples of different things affiliates could talk about with people who exclusively stay on-wiki. Clovermoss🍀 (talk) 03:11, 9 September 2026 (UTC)
  • Having thought through this further, I am more cautious due to the way interest groups generally interact on English Wikipedia, which is through WikiProjects. Participation, and even offline participation, in Wikipedia:WikiProject Women in Red (WiR) is larger than participation in many (most?) affiliate projects. WiR has created an ecosystem of on-wiki pages that document their work, but doesn't have a dedicated board to discuss its activities with WP:MILHIST. Questions following from that include: would setting up like a WikiProject work similarly for affiliates? Are WikiProjects are also missing out on some communication opportunity? Do we want to treat mostly offline groups (eg. affiliates) differently from mostly online groups (eg. WikiProjects), and if so, why (and how to deal with the sliding spectrum between those)? CMD (talk) 23:18, 8 September 2026 (UTC)

A lot of "build it and they will come" venues get proposed, which is why often there's a reluctance to create yet another page that no one reads or edits. As Barkeep49 said, though, I don't think anyone will make a fuss about a new project space page, as long as it doesn't impose any additional work or mandatory changes to workflow for anyone not interested in using it. This does mean that affiliates shouldn't expect that any message they place in this new venue will get read by a broad segment of the commumity. So if they are seeking broader awareness, they still need to use one of the existing appropriate venues to garner the desired attention. isaacl (talk) 01:24, 9 September 2026 (UTC)

I think part of the issue is that it isn't always clear what the appropriate venue is, which is why I suggested making miscellaneous affiliate communication stuff part of this boards focus (could also be a different existing board of course). If input is sought regarding a project with a specific topical or geographical scope there may be existing noticeboards for that, but there isn't anywhere (obvious to me at least) for projects without that focus and/or where there isn't an existing board closely aligned. Thryduulf (talk) 01:47, 9 September 2026 (UTC)

I will also note, though, that I'm wary of creating a page under the assumption that others want to use it, if we haven't heard from the others. Is there a clamor from any affiliates that a separate venue for them is needed? I think we should consult with them first and ask if they want a new venue. Plus it feels very English Wikipedia-centric to assume that affiliates want to communicate with the English Wikipedia community. isaacl (talk) 01:33, 9 September 2026 (UTC)

I could ask other people, but I'd worry I'd get accused of canvassing. I was mostly just going off the impression I've had in off-wiki conversations that affiliate-related people aren't inherently opposed to engaging with the community more. A lot of people see themselves as part of the community and don't like the idea of being perceived as some seperate, outside force. I think the main concern would be people treating them like Alan just treated me here. It's a bit hard to convince people to go participate in an optional place where people might accuse you of doing all sorts of things you're not doing. But I'm from the on-wiki side of things first, so I understand the importance of constructive criticism and giving people an outlet. It's difficult to rebuild trust by doing nothing. Even just identifying that a problem exists can be a helpful step. I've found certain people have been surprised by Wikipedia:Editor reflections, particularly my on-going analysis when it comes to what people think about WikiEd. Clovermoss🍀 (talk) 01:49, 9 September 2026 (UTC)
Also, having a space on enwiki doesn't mean there couldn't be pages on other projects where affiliates engage with the relevant communities. I think we're the only project with a dedicated WMF page (and am not sure what led to the creation of it), but other projects could theoretically build their own place for this if they think it is a useful concept. Clovermoss🍀 (talk) 01:52, 9 September 2026 (UTC)
See Wikipedia:Village pump (proposals)/Archive 168 § Proposal: New Village Pump Page. This page was created with the goal of increasing communication with the WMF. Their communication staff, though, didn't commit to post everything to this village pump page, as it didn't want to split discussion if another page was more appropriate. isaacl (talk) 02:11, 9 September 2026 (UTC)
Thanks for the link! It looks like Alsee did good work on getting that proposal through. I wonder if they agree or disagree about whether this is a similar sort of situation nessecitating the need for a new page. I feel like their experience is so 1-to-1 here in a way you rarely get when it comes to proposing new things. Clovermoss🍀 (talk) 02:19, 9 September 2026 (UTC)
I appreciate there are some who like the "build it and see if they come" approach. Personally, I prefer to gauge demand on both sides (those who would potentially edit the page and those who would read it) and figure out if it is sufficient to sustain a new venue. I feel that for a new venue to work, it needs to be publicized and some effort invested into making it a minimally useful forum, but if there's no interest in using it, I don't want to push people into it.
I imagine the vast majority of affiliates don't have a lot of spare personnel for outreach to many Wikimedia communities. So asking them to come to English Wikipedia in addition to meta, the more typical cross-Wikimedia site, feels to me a bit like asking for special treatment. isaacl (talk) 02:02, 9 September 2026 (UTC)
If they'd prefer to engage on Meta with their limited time, that's their choice, but that's not my preference. I'd be willing to take an active role at Wikipedia:Affiliate's corner but not on Meta due to some rather upsetting recent experiences. I have more faith in our processes for dealing with problematic behaviour, even if they aren't perfect. Clovermoss🍀 (talk) 02:15, 9 September 2026 (UTC)
I understand your position. But I don't think that means we should tell all affiliates that the English Wikipedia community has more faith in its own processes, so they should come here for any discussions they want to have about their initiatives. isaacl (talk) 02:25, 9 September 2026 (UTC)
But that's not what I'd be telling people. If people have more faith in Meta, they should participate where they feel the most comfortable. I just personally wouldn't be interested in having these conversations there. Clovermoss🍀 (talk) 02:31, 9 September 2026 (UTC)
Well that just goes back to my question: why are we asking affiliates to engage specifically with the English Wikipedia community? Wikimedia Deutschland of course has a specific interest when it is working on a MediaWiki feature that is of interest to English Wikipedia, and it has indeed come to existing venues to discuss its work on improving the autogenerated citation names, without needing a new venue. But for many others, it feels like we'd be asking them to do something special for English Wikipedia. isaacl (talk) 02:46, 9 September 2026 (UTC)
My understanding is that WCNA is often seen as somewhat of an English Wikipedia gathering, which may provide a different perspective to affiliates elsewhere. CMD (talk) 03:17, 9 September 2026 (UTC)
WCNA is technically a user group in its own right, even if it does have a strong enwiki focus. Clovermoss🍀 (talk) 03:39, 9 September 2026 (UTC)
Well that's confusing! I meant it in the, uh, sensu lato sense of groups regularly attending the conference. CMD (talk) 06:32, 9 September 2026 (UTC)
What is this affiliate-related people aren't inherently opposed to engaging with the community more stuff? (this is also partly a response to a couple other comments). Affiliate-related people are part of the community. There are Wikipedians who don't engage with anything affiliates do, but if there are affiliate members that don't engage with one or more of the Wikipedia projects and their communities I haven't met them. Affiliate people are both existing volunteers who say "hey, maybe I'll go to one of these in-person meetings" and readers who become volunteers through the in-person meetings. I get there's a tendency to frame affiliates as some sinister cabal, but they're just you (the general you) if you went to an event, or took part in an edit-a-thon run by an affiliate, etc. And if we're specifically talking about people who get paid through an affiliate, we're talking about such a teeny tiny portion of "affiliate people" that there's not much to talk about. — Rhododendrites talk \\ 02:20, 9 September 2026 (UTC)
It is now over ten years since my time with an affiliate came to an end, so I think I have sufficient detachment to comment here. Yes affiliates are part of the movement, many people involved in affiliates were wikimedians before during and after their time with an affiliate. I like to think that applies to my time at Wikimedia UK, and many of the people in GLAM roles that I met in other chapters. But there are people who work for affiliates who need a little guidance and maybe some targets to interact with the volunteer community. I'm not convinced that the village pump of the English language Wikipedia is always the best place for that, The GLAM newsletter, The Signpost, Meta, outreach wiki if that still exists, are all relevant places/media depending on the topic. I do think that the WMF KPIs for affiliates should include some targets for interaction with the most relevant volunteer communities, and village pumps would be part of that. However a separate noticeboard should only come after the amount of postings has reached a point where some on the main noticeboard would like those threads spun off to somewhere they can choose not to follow, not before. Pave the desire lines, don't try to predict where the desire lines will go. ϢereSpielChequers 13:39, 9 September 2026 (UTC)
@WereSpielChequers: That sounds more like a directory than a place for conversations, which isn't a bad idea in its own right. I'm just not sure how one would pave the desire lines without fragmenting discussions? Could you elaborate a bit more on what you envision here? Clovermoss🍀 (talk) 16:42, 9 September 2026 (UTC)
Hi Hannah, well I think I am suggesting fragmenting discussions. So for example, when an affiliate is running a GLAM program that is relevant to a particular Wikiproject, I would hope that someone from the affiliate would talk on the Wikiproject page. If someone wants to discuss chapters in general then I think that discussion is best on Meta. Where I would like to see a change is with the individual Wikis that each chapter seems to have. At least they did when I was last involved in a chapter circa 2015. Having separate Wikis under the control of individual chapters might help with any proof if twere needed that those chapters were independent of the WMF. But it costs money and or requires extra volunteer time and because they are outside Single User Login, it creates a gulf between those involved in a chapter and those who aren't involved but might have got involved in a specific discussion. I have subscribed to discussions on many talkpages on several wikis, and I get notifications from an eclectic variety of Wikis from among the thousand that the WMF supports. Extending SUL to the wikis of affiliates, or consolidating such wikis with meta, would in my view reduce some of the gulf that I see developing between affiliates and the broader movement. Not sure if that addresses your original concern, and yes I can see that chapters with a right to left script might not want to migrate their wiki to meta. But extending SUL to chapter wikis would at least make it easier to include people in fragmented discussions. ϢereSpielChequers 17:16, 11 September 2026 (UTC)

Source Verification Suggestion

Testing donor account creation - a new post on the WMF Fundraising Hub

Wikimedia Foundation Bulletin 2026 Issue 17

Invitation to join the conversation on future of affiliates

From 2011: Opinion Letter to Wikimedia Foundation, Inc.

Scholarships for Wikimania 2027

Related Articles

Timelines

Top Qs

Fact Checks