OMG THANK YOU SO MUCH I CANT Believe i got noticed,if you can (and mind) can you put it on the front page like some of the other article audios (you dont need to credit me,idrk) and also can you tell me how the people of the project came up with the idea to make the ai voice thing
What are the steps to contribute to the AI voice project
Hello,
I'm interested to contribute creating AI voices audios but I'm not sure about the process. Do I wait for someone to request for a specific article? Or is it up to me to decide which articles to AI voice? InfiniteMonkeyMachine (talk) 15:30, 3 December 2024 (UTC)
Hello, thanks for asking. It's up to you. However, please do not upload many audio files if the quality is bad. Moreover, so far only open source software has been used (it would be good to keep it that way) and if you're using something other than SoniTranslate it would be helpful if you added information about that to the guide which also has info about which types of commons minor issues to look out for (e.g. adjusting the source text to spell out abbreviations). Prototyperspective (talk) 16:08, 3 December 2024 (UTC)
Audible links
Screenreaders often read out linked URLs, which you can imagine makes Wikipedia pretty unusable for a lot of blind people. Could we read the blue text, but have some sort of chime or sound in the background to indicate it's a link to another page? HLHJ (talk) 01:57, 13 December 2024 (UTC)
They also read a lot of things one doesn't want them to which makes them quite suboptimal even if the voice sounded more natural (like [1][2] if there's refs etc). Also one can't replace abbreviations. The linked Commons help page gives some insight about how much is done to get the final source text.
Good idea, however when it comes to wikilinks I don't think this would be useful to a listener and would only disrupt the listening experience which should be as pleasant as listening to a podcast and there would be lots of sounds due to the many wikilinks. I had a similar idea for section headers as well as tags like [citation needed], there I think would be very useful: It would also be nice if it played a distinctive sound for every section header and subsection header and/or said "Section:" before the header (see below).
There could also be two audios per article where one is intended for visually impaired people and one for everybody (and I think nonvisually impaired people would make up a way larger fraction of users of these which listening data about podcast-listeners supports). I think most of the former would actually also prefer a listening experience that is like listening to a podcast or audiobook so I think it would need feedback from some actually visually impaired people about what they need or prefer, e.g. how they usually navigate Wikipedia articles. A lot of unlinked terms also have Wikipedia articles and each wikilink usually is wikilinked only at first occurrence and not also e.g. where the user may be interested in the other article. Thus, mainly due to the former and because it disturbs the listening experience without being interactive I don't think it would be good to add this functionality as described.
Another thing I thought of was if a podcast player could have a button that you can press and it goes through the wikilinked articles of the currently read sentence so you can interactively go there or let it read its lead. However, I don't know of a way to make it play a sound for wikilinks anyway and it's not readily possible (wikilinks are text like any other in the current way the article text is fetched and in the tool there is no feature to play some distinctive sound or similar for specified keywords). Thanks for feedback. Prototyperspective (talk) 10:45, 13 December 2024 (UTC)
Yeah this seems correct. It would be so beneficial if someone would utilise the newer ai voices for wikipedia. Should be like a podcast, that’s the thing about screenreaders - they just don’t know understand how to read wikipedia articles. Coneill1774 (talk) 21:22, 28 December 2024 (UTC)
As well, I think that the audio for wikpedia articles definitely should not be in file form, but should still be a screenreader, that simply understands how to interpret wikipedia articles. Coneill1774 (talk) 21:24, 28 December 2024 (UTC)
That's simply speaking practically impossible and also has no benefit while having it as file has lots of benefits such as less server-load and ability to easily download the file to your device and use podcast-player features like skip x seconds back, offline listening, and so on. Prototyperspective (talk) 21:56, 28 December 2024 (UTC)
TTS engine used?
Am I correct that the sample audio files were generated using OpenAI's TTS engine (or similar commercial option)? I tried using piper-tts, which seems to be the most-recommended open source engine, and it sounds nowhere close. Unless I missed an open-source alternative?
For context, I would love to build an on-demand TTS feature into the Wikipedia Android app, where the user could randomly listen to an article as if it were a podcast. But as a user of this feature, the only way I could see myself listening to it for an extended time is if it sounds sufficiently human, and unfortunately open-source options don't seem to be at that level yet.
[edit]: It looks like the "SoniTranslate" tool is just a frontend for orchestrating the transcription/generation/etc, but it doesn't generate the speech on its own. It seems to default to using Microsoft's Edge TTS, which is technically a free API (for now) but definitely not open-source, unless I'm mistaken. If you do a web search for the names of the voices, such as "en-US-BrianNeural-Male", they are associated with Microsoft Azure AI Speech. Does this need to be taken into account when attributing the files on Commons? Dmitry Brant (talk) 04:41, 7 January 2025 (UTC)
Hi everyone, I'm new on this field.
It seems we need a map of existing high quality TTS engines, something lile:
If you bump into such things, we could be interested.
[EDIT] I started to list notorious text-to-speech models in a new row in {{Generative AI}}. The page Text-to-speech model is missing. There surely are reference in Tacotron and other academic articles to create such an article some days where we can then list major models properly. Yug(talk) 🐲 11:49, 18 February 2025 (UTC)
The Text-to-speech article is very outdated and a new article may be good to have. In the Generative AI template, the row for "Speech" shouldn't just have these 3 items, this is misleading. I think most major freely-usable ones including bark have been integrated into SoniTranslate. Here is an issue about integrating Kokoro-TTS which seems to also be performant, including a link to a benchmarking comparison. If anybody knows of an AI voice they think would be better than the en-US-Brian en-US-Ava voices I usually use or the best option, I recommend creating a WP-article audio as a sample. Prototyperspective (talk) 15:19, 20 May 2025 (UTC)
Am I correct that the sample audio files were generated using OpenAI's TTS engine No, they weren't. They were created as described in the page linked in the file description with the AI voice named there. (That page is also linked from the WikiProject page.)
"en-US-BrianNeural-Male" Yes, that's the voice I most commonly used (the multilingual variant of it). As for attribution, see c:Template:PD-algorithm. If you know of an AI voice not associated with any particular company that sounds at least as good, please name it here or create an issue at the SoniTranslate repo to have it integrated there. Maybe there already is one of that kind since I haven't tried all the many options integrated into SoniTranslate.
where the user could randomly listen to an article as if it were a podcast Please see the page that describes the process how these audios created and maybe try creating an audio for an article for why this is unlikely to work well. Further issues include that it would lead to lots of server strain from all the users generating the files anew, that one would have to wait minutes for it to start, etc. Prototyperspective (talk) 14:30, 20 May 2025 (UTC)
Hello Prototyperspective, Recent progresses on the side of ML TTS, Lingua Libre (support for phrases and texts recording planed for 2025 Q3), your well written vision on meta, and possible grants avenues make things more possible. This vision next step would need dedicated work, therefore money. There are the avenues which I would like to explore in the first half of 2025:
Best regards, Yug(talk) 🐲 11:51, 25 February 2025 (UTC)
Can i use Capcuts text to speech feature to contribute,i plan on contributing to Wikipedia policy pages
I want to highly contribute but have little to none sources to,However i have Capcut but i think using Capcut has restrictions and is limited for Wikipedia.
Don't know what you're talking about. Can capcut be used to narrate text? If so what's the quality and I don't think this would be fine, for example because capcut is not open source software and not needed when there is open source software that can do this quite well. Prototyperspective (talk) 11:54, 20 May 2025 (UTC)
Script to read articles aloud
Hello everyone, some of you might be interested in WikiNarrator, a new user script to read selected passages or whole articles aloud. More information on installation and usage can be found at User:Phlsph7/WikiNarrator. Please let me know if you have new feature ideas for the script or encounter problems. Phlsph7 (talk) 10:06, 20 May 2025 (UTC)
Looks interesting! I wanted to try it out to see how the voice sounds like, whether it can read the full article at once, and how it deals with various difficulties such as abbreviations but I can't make it show. Does it not work in Firefox? I use Firefox mainly because it is open source. I clicked Tools->Show/hide WikiNarrator and tried selecting some text but can't see any button. It could be useful especially for articles that do not have an audio of the full article available and one could consider adding it to the project page if it works for Firefox too. Prototyperspective (talk) 11:53, 20 May 2025 (UTC)
Hello Prototyperspective and thanks for taking a look at the script! Firefox is not that common anymore so I wrote the script primarily for Chrome and Edge. I just made some adjustments and the script is working for me now on Firefox as well. Could you give it a try to see if it works for you too? Phlsph7 (talk) 14:40, 20 May 2025 (UTC)
Thanks, I still don't see anything but I'm not sure how the tool is supposed to show up. Maybe it's blocked by uBlock or something? I check whether the isFirefox is true in the script on my side and it is. Prototyperspective (talk) 14:47, 20 May 2025 (UTC)
Now it does show up (in the bottom right) after clicking after clicking "Show/hide WikiNarrator" twice but the voice dropdown is empty and it does not play any audio when selecting text and clicking play. Prototyperspective (talk) 15:23, 20 May 2025 (UTC)
The browser needs to load the voices before the script can be used. It's probably the case that they are not yet loaded when the script opens. Maybe waiting some time before interacting with the script could solve the problem. If you have some familiarity with the firefox console, you could try opening it (F12) and enter speechSynthesis.getVoices(); . If this shows an empty array ([]) then this means the voices haven't been loaded yet. Meanwhile, I'll see if there is a better solution for this. Phlsph7 (talk) 15:31, 20 May 2025 (UTC)
Okay, FYI it never loads those voices. Waited for minutes before clicking the show button as well as after making the tool show up but speechSynthesis.getVoices(); keeps returning an empty array. Prototyperspective (talk) 15:46, 20 May 2025 (UTC)
Hm, I'm not sure if that's a problem with the script or with the browser. It could also be that no voices are installed in the browser for some reason. For me, they are installed by default but different browsers have different voices. I ran into various problems with speech synthesis in Chrome as well but there were usually some workarounds. Could you try whether the example at https://eeejay.github.io/webspeechdemos/ works for you? In the meantime, I modified the script: it should show a warning message now if the voices haven't been loaded. Phlsph7 (talk) 15:57, 20 May 2025 (UTC)
Much better if there's some error message. No, there are no voices in that dropdown either. Prototyperspective (talk) 17:26, 20 May 2025 (UTC)
In that case, it's probably a problem with the browser, see also https://stackoverflow.com/questions/46617366/speechsynthesis-getvoices-not-listing-voices-in-firefox . When I installed firefox today, the voices worked right away. Maybe updating to the newest version could solve the problem. There are also firefox extensions that provide reading functionality, like https://addons.mozilla.org/en-US/firefox/addon/read-aloud/ . I'm not sure where the browser gets the voices from, possibily the operating system (like windows, linux, etc). This would mean that if no voices are installed with the operating system, the browser wouldn't have any access to voices either. Phlsph7 (talk) 17:45, 20 May 2025 (UTC)
Experimentation for a screen reader approach
Screen readers I found earlier either cost money or have low voice quality. Now I tried something new, the following seems to work on Firefox for a natural-sounding audio playback:
The default voices sound really bad there too (you could try it first to check if it's different for you) – right click on the Read Aloud addon icon and select "Manage Extension", then go to preferences and select "Google Translate English". I don't know which voices work well but that one is quite good.
Go to the Stylus preferences, create a new style, give it a name, select URLs starting with, enter e.g. https://en.wikipedia.org/wiki/, give the profile a name, then paste this CSS and save:
5. On a Wikipedia page, click the stylus button in your browser toolbar and enable the style you created
6. Select all the text of the article
7. Click the Read Aloud addon and click play (note that it also shows the text currently being read in the popup and one click it to jump to another part, which is an advantage)
Problems:
One may need to click play several times for it to work and I haven't yet tested if it reads an entire article nonstop or if it break in between.
This is for desktop but one would generally use this on mobile – the Stylus and Read Aloud addons are also available for Firefox on mobile but it would be very cumbersome to set it up and to select all the article text for example.
It's also cumbersome to have to launch the Firefox app, select the article and scroll to the place where one last left off to then select the text and click play. With audio files one can have them in an audio player – launching quickly with a tap where one just needs to tap play – and quickly change between which audio/article to listen to and continue where one has left off etc.
One can't download the audio to listen to it offline and I don't know if it would cause issues if e.g. many thousands of people request audios from GT, often redundantly for the same article.
The voice is good but not as good as the ones used in the demos.
One can't correct mispronunciations like who instead of W.H.O.
People realistically don't know about these addons and how they can be used – it needs a built-in way, in part also because it's important that one can just click play at the article, as one can with the human-read spoken Wikipedia audios. Wikispeech seems far from there when it comes to the voice quality if that's what would be used. Furthermore, as a first step probably the audio player needs to be modernized – see W317: A proper audio player (voting now open).