[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/mlp/ - Pony


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: altOP.jpg (1.26 MB, 2119x1500)
1.26 MB JPG
Welcome to the Pony Voice Preservation Project!
http://youtu.be/730zGRwbQuE

The Pony Preservation Project is a collaborative effort by /mlp/ to build and curate pony datasets for as many applications in AI as possible.

Technology has progressed such that a trained neural network can generate convincing voice clips, drawings and text for any person or character using existing audio recordings, artwork and fanfics as a reference. As you can surely imagine, AI pony voices, drawings and text have endless applications for pony content creation.

AI is incredibly versatile, basically anything that can be boiled down to a simple dataset can be used for training to create more of it. AI-generated images, fanfics, wAIfu chatbots and even animation are possible, and are being worked on here.

Any anon is free to join, and there are many active tasks that would suit any level of technical expertise. If you’re interested in helping out, take a look at the quick start guide linked below and ask in the thread for any further detail you need.

EQG and G5 are not welcome.

>Quick start guide:
docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Introduction to the PPP, links to text-to-speech tools, and how (You) can help with active tasks.

>The main Doc:
docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit
An in-depth repository of tutorials, resources and archives.

>Online speech generation
haysay.ai

>MOSS-TTS-1.7B-PNY v0.1: a finetune of MOSS-TTS + custom vocoder for 48KHz audio
HuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNY
Colab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBB
Public demo: http://198.53.64.194:35029/

>Active tasks:
Research into animation AI
Research into pony image generation

>Latest developments:
pastebin.com/4p00iUZM (embed)

>The PoneAI drive, an archive for AI pony voice content:
drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp

>Clipper’s Master Files, the central location for MLP voice data:
mega.nz/folder/jkwimSTa#_xk0VnR30C8Ljsy4RCGSig
mega.nz/folder/gVYUEZrI#6dQHH3P2cFYWm3UkQveHxQ
drive.google.com/drive/folders/1MuM9Nb_LwnVxInIPFNvzD_hv3zOZhpwx

>Cool, where is the discord/forum/whatever unifying place for this project?
You're looking at it.

Last Thread: >>43289079
>>
FAQs:
If your question isn’t listed here, take a look in the quick start guide and main doc to see if it’s already answered there. Use the tabs on the left for easy navigation.
Quick: docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Main: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit

>Where can I find the AI text-to-speech tools and how do I use them?
A list of TTS tools: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.yuhl8zjiwmwq
How to get the best out of them: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.mnnpknmj1hcy

>Where can I find content made with the voice AI?
In the PoneAI drive: drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp
And the PPP Mega Compilation: docs.google.com/spreadsheets/d/1T2TE3OBs681Vphfas7Jgi5rvugdH6wnXVtUVYiZyJF8/edit

>I want to know more about the PPP, but I can’t be arsed to read the doc.
See the live PPP panel shows presented on /mlp/con for a more condensed overview.
2020 pony.tube/w/5fUkuT3245pL8ZoWXUnXJ4
2021 pony.tube/w/a5yfTV4Ynq7tRveZH7AA8f
2022 pony.tube/w/mV3xgbdtrXqjoPAwEXZCw5
2023 pony.tube/w/fVZShksjBbu6uT51DtvWWz

>How can I help with the PPP?
Build datasets, train AIs, and use the AI to make more pony content. Take a look at the quick start guide for current active tasks, or start your own in the thread if you have an idea. There’s always more data to collect and more AIs to train.

>Did you know that such and such voiced this other thing that could be used for voice data?
It is best to keep to official audio only unless there is very little of it available. If you know of a good source of audio for characters with few (or just fewer) lines, please post it in the thread. 5.1 is generally required unless you have a source already clean of background noise. Preferably post a sample or link. The easier you make it, the more likely it will be done.

>What about fan-imitations of official voices?
No.

>Will you guys be doing a [insert language here] version of the AI?
Probably not, but you're welcome to. You can however get most of the way there by using phonetic transcriptions of other languages as input for the AI.

>What about [insert OC here]'s voice?
It is often quite difficult to find good quality audio data for OCs. If you happen to know any, post them in the thread and we’ll take a look.

>I have an idea!
Great. Post it in the thread and we'll discuss it.

>Do you have a Code of Conduct?
Of course: 15.ai/code

>Is this project open source? Who is in charge of this?
pony.tube/w/mqJyvdgrpbWgZduz2cs1Cm

PPP Redubs:
pony.tube/w/p/aR2dpAFn5KhnqPYiRxFQ97

Stream Premieres:
pony.tube/w/6cKnjJEZSCi3gsvrbATXnC
pony.tube/w/oNeBFMPiQKh93ePqTz1ns8
>>
>>43451491
I added MOSS-TTS by Delta to the OP.
Since we had some developments and reached 500, I figured it would be good to make another thread.
If we die too soon or become too much of a bump thread, I'll wait a bit before making the next one.
>>
File: 1628825457742.jpg (191 KB, 1600x1538)
191 KB JPG
I want to say, Delta, thank you for getting us back ai voices with emotion control, its very cool.
>>
>>43448113
https://www.youtube.com/watch?v=IOtG5f0-YcM
reposting Vul new cool song from last thread
>>
I wish there was more development in the VR side of things (other than one Anon doing his own thing)
>>
Next MOSS-TTS-PNY release (sometime) will split emotion into emotion and delivery classes.
Also, last night someone slid into my Xitter DMs and pointed out there was enough data, so I'll most likely add the voices of the Mare Fair mascots.
>>
>>43456294
Nice! I still haven't tried MOSS yet, but I may need pony voices soon for a project, so it's good to see it being updated.
>>
>>43456294
>split emotion into emotion and delivery classes
whats 'delivery classes' ?
>>
>>43456623
Linux user talk for 'more control'
>>
>>43456294
excited!!!!!!!!!!!!!!!!!! please do this soon!!!!!!!!!!!!! it will make my pony audiobooks so much better!!!!!!!!!
>>
>>43456623
Emotion is well, emotion. Delivery is how you say something, like normal, shouting, whispering. You can be shouting something angrily or happily, or whispering happily or sadly, for example. So, it makes sense to decouple them.
>>
Any good TTS for Derpy? and for Sombra?
>>
File: lyra jumping.gif (751 KB, 393x323)
751 KB GIF
>>43456959
>You can be shouting something angrily or happily, or whispering happily or sadly
oh fuck yeah, thats something I wanted for years, since trying to get a "scared whisper" vs "smug whisper" was a sisyphean task of constantly getting wrong tones and vibes from the generated clips.
>>
>>43456793
share them?
>>
File: IMG_2607.png (265 KB, 1147x1049)
265 KB PNG
RVC v3 coming soon.
>>
File: 1664300976353425.jpg (6 KB, 203x250)
6 KB JPG
>>43457420
Oh, nevermind, disregard. This announcement is apparently 3 years old.
>>
>pone
>>
>>43456294
We will be able to mix and match emotions to create new ones? (Similarly to IndexTTS2

(e.g 20% sad 80% angry)

Also we will get new emotion IDs?
>>
>>43458312
+1 to this question
>>
>>43409987 >>43410118 < - audiobook anon here. I replaced the subpar narrator TTS with the more quality-sounding Gemini 3.1 TTS.

before: https://files.catbox.moe/o3wv7b.wav

after: https://files.catbox.moe/pnmty9.wav

hopefully, I should have the entire shebang finished in the coming week.

does anybody have a good sound effects library I can source from? doubly grateful if its a torrent for a pirated premium library
>>
>>43459095
>after: https://files.catbox.moe/pnmty9.wav
files broken, cant play it.
As for the sound effects, that bit difficult as everything I have is a large mix bag of whatever I managed to google and download at the moment I needed it. I guess I can give you some starting points that USUALLY worked out fine for me (but I would recommend trying to find and build your own audio packs, since you will know best what you are going to use in your projects):
https://sounds.spriters-resource.com
https://sound-effects.bbcrewind.co.uk/search
https://www.bfxr.net this one is decent at getting random scifi sounds but bit pain in ass to use
https://opengameart.org/art-search-advanced?keys=&field_art_type_tid%5B%5D=13&sort_by=count&sort_order=DESC
>>
>>43459238
>files broken, cant play it.
probably on your end. here's a vocaroo instead.

before: https://voca.ro/1oN59Zz09OK5

after: https://voca.ro/1il1QV91jwoA

>As for the sound effects, that bit difficult as everything I have is a large mix bag of whatever I managed to google and download at the moment I needed it. I guess I can give you some starting points that USUALLY worked out fine for me (but I would recommend trying to find and build your own audio packs, since you will know best what you are going to use in your projects):
thanks, anon! these are great.
>>
>>43459291
>after: https://voca.ro/1il1QV91jwoA
Man, the Gemini is a really massive upgrade from the previous narrator, with the exception of very few odd bits it sounds like a real person reading from a script.
Just don't break a bank while paying for google services. Maybe it would be worth it to gen 5 minutes of audio with this tts voice and use rvc and so-vits to create a local ai voice version of it?
>>
>>43410118
what voice model architecture for the ponies are you using here? sovits, rvc, moss? it's really good
>>
>mares
>>
>>43457409
different audiobook anon hi, posting new project, ponybook! uses moss-tts server to generate audiobooks in real time
you plug an llm in for the first time to annotate who's speaking and when, add emotion tags to every line, and assign voices to characters. you can change the voice whenever you want. comes with an anon voice that i think fits! also, the included model is a finetuned voicemodel with a phenome hybrid that ive been working on.
as part of this release im dropping a demo, and the first full 8 hour long audiobook i generated,
excuses by mobius_
https://u.pone.rs/obiitvsw.m4b
generated by ponybook
>narrated by twilight sparkle
https://github.com/maresmaremares/ponybook
>>
File: pp2.png (419 KB, 1600x1612)
419 KB PNG
>>43460757
oh snap, you're the bonzipony dev, right? great seeing you here, dude. were you also the anon last thread who used fable to improve the tts' pronunciation? if so, i might actually use your finetune.

i only listened to the first minute or so of the audiobook, but from what I've heard it seems it suffers from an issue that I would like resolved in the next update from Delta, if possible: The character voices can fluctuate between the different performances given to them in each season. This is most apparent with Applejack: S1 AJ and post-S1 AJ are two vastly different takes on her voice - the other M6 to a lesser extent - to the point that I feel like they should constitute their own model. Also, I'm not sure if this is a memory-saving measure you put in place because the github says that it can regenerate entire paragraphs, but it almost seems like the tool is generating new lines for the narrator per-sentence rather than per-paragraph. Because of this, it is quite apparent within a paragraph where one voice clip starts, and where one ends, botching the natural rhythm and cadence of the narration in the process. This could remedied by making the narrator generate per-paragraph, though if it's already an option in the tool, then ignore me.

all in all, exciting times ahead for ai content! really hoping everybody starts producing some bangers. we have more than enough tools already - no excuses
>>
>>43461125
>>43460757
ah snap, i forgot to give you props. genuinely, nice work anon. you made a neat tool with bonzi, and this is looking like another banger, too. keep up the good work!
>>
File: 1660070598192988.gif (541 KB, 401x340)
541 KB GIF
>>43460757
Darn, If I wasn't prepping for convention I would had check it out. Anyhow, thanks for all your hard
work Anon!
>>
>>43458312
>>43458721
In some future version if it's possible with the amount of data I have, but not anytime soon.
>>43460757
Very cool.
>>
>>43451491
Bump.
>>
pump
>>
>>43463891
>>
File: vaporpopbubblegum.gif (127 KB, 125x125)
127 KB GIF
>>
.
>>
>>43463891
>>
>>43463891
rump
>>
Wake up, thread.
>>
File: 1465240649798.gif (208 KB, 1458x950)
208 KB GIF
>>43468700
>>
>>43468337
Made with Moss-TTS
>>
>>43468195
This.
>>
>>43469920
>>
Has anyone tried to go through the catalogue of early fan music written / sung from the perspective of an pony, but used AI to dub over the voice to actually sound like said pony's voice?
>>
>>43468973
Great start. Keep on improving.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.