Welcome to the Pony Voice Preservation Project!http://youtu.be/730zGRwbQuEThe Pony Preservation Project is a collaborative effort by /mlp/ to build and curate pony datasets for as many applications in AI as possible.Technology has progressed such that a trained neural network can generate convincing voice clips, drawings and text for any person or character using existing audio recordings, artwork and fanfics as a reference. As you can surely imagine, AI pony voices, drawings and text have endless applications for pony content creation.AI is incredibly versatile, basically anything that can be boiled down to a simple dataset can be used for training to create more of it. AI-generated images, fanfics, wAIfu chatbots and even animation are possible, and are being worked on here.Any anon is free to join, and there are many active tasks that would suit any level of technical expertise. If you’re interested in helping out, take a look at the quick start guide linked below and ask in the thread for any further detail you need.EQG and G5 are not welcome.>Quick start guide:docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/editIntroduction to the PPP, links to text-to-speech tools, and how (You) can help with active tasks.>The main Doc:docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/editAn in-depth repository of tutorials, resources and archives.>Online speech generationhaysay.ai>MOSS-TTS-1.7B-PNY v0.1: a finetune of MOSS-TTS + custom vocoder for 48KHz audioHuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNYColab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBBPublic demo: http://198.53.64.194:35029/>Active tasks:Research into animation AIResearch into pony image generation>Latest developments:pastebin.com/4p00iUZM (embed)>The PoneAI drive, an archive for AI pony voice content:drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp>Clipper’s Master Files, the central location for MLP voice data:mega.nz/folder/jkwimSTa#_xk0VnR30C8Ljsy4RCGSigmega.nz/folder/gVYUEZrI#6dQHH3P2cFYWm3UkQveHxQdrive.google.com/drive/folders/1MuM9Nb_LwnVxInIPFNvzD_hv3zOZhpwx>Cool, where is the discord/forum/whatever unifying place for this project?You're looking at it.Last Thread: >>43451491
FAQs:If your question isn’t listed here, take a look in the quick start guide and main doc to see if it’s already answered there. Use the tabs on the left for easy navigation.Quick: docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/editMain: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit>Where can I find the AI text-to-speech tools and how do I use them?A list of TTS tools: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.yuhl8zjiwmwqHow to get the best out of them: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.mnnpknmj1hcy>Where can I find content made with the voice AI?In the PoneAI drive: drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCpAnd the PPP Mega Compilation: docs.google.com/spreadsheets/d/1T2TE3OBs681Vphfas7Jgi5rvugdH6wnXVtUVYiZyJF8/edit>I want to know more about the PPP, but I can’t be arsed to read the doc.See the live PPP panel shows presented on /mlp/con for a more condensed overview.2020 pony.tube/w/5fUkuT3245pL8ZoWXUnXJ42021 pony.tube/w/a5yfTV4Ynq7tRveZH7AA8f2022 pony.tube/w/mV3xgbdtrXqjoPAwEXZCw52023 pony.tube/w/fVZShksjBbu6uT51DtvWWz>How can I help with the PPP?Build datasets, train AIs, and use the AI to make more pony content. Take a look at the quick start guide for current active tasks, or start your own in the thread if you have an idea. There’s always more data to collect and more AIs to train.>Did you know that such and such voiced this other thing that could be used for voice data?It is best to keep to official audio only unless there is very little of it available. If you know of a good source of audio for characters with few (or just fewer) lines, please post it in the thread. 5.1 is generally required unless you have a source already clean of background noise. Preferably post a sample or link. The easier you make it, the more likely it will be done.>What about fan-imitations of official voices?No.>Will you guys be doing a [insert language here] version of the AI?Probably not, but you're welcome to. You can however get most of the way there by using phonetic transcriptions of other languages as input for the AI.>What about [insert OC here]'s voice?It is often quite difficult to find good quality audio data for OCs. If you happen to know any, post them in the thread and we’ll take a look.>I have an idea!Great. Post it in the thread and we'll discuss it.>Do you have a Code of Conduct?Of course: 15.ai/code>Is this project open source? Who is in charge of this?pony.tube/w/mqJyvdgrpbWgZduz2cs1CmPPP Redubs:pony.tube/w/p/aR2dpAFn5KhnqPYiRxFQ97Stream Premieres:pony.tube/w/6cKnjJEZSCi3gsvrbATXnCpony.tube/w/oNeBFMPiQKh93ePqTz1ns8
Any updates on MOSS?
>9
>>43515526https://www.youtube.com/watch?v=FziZYANzClshttps://u.pone.rs/rfxsptss.zipI made an album recently that uses Delta's new RVC NG for the between-song dialogue. It's actually pretty nice, especially since it's actually possible to get more varied energy out of it for yelling and whispering and such, whereas the old RVC struggled with that. Old RVC is still on the singing though.
>>43516297Nice!
>10
Can AI operate with flash/toonboom puppets yet? Because now it's clear that generating sloppy shit from scratch will never git gud.
>>43517328
Last bump.
>>43517853ToonBoom has a product called Ember which integrates with Harmony and does some very limited AI stuff, like extending background images and generating scratch tracks (roughouts of character dialogue that you use for timing your animation until the professional voicework is done). As for 3rd-party AI integration with Toon Boom Harmony, I know nothing.
>>43518906is it good
>>43515531Nothing on MOSS for now, I've been noticing a lot of people could use something more lightweight for gamedev or other purposes, so last night I began developing a small non-autoregressive model (well, kind of semi-autoregressive because it uses GRUs -- you need some form of autoregression to model pony speech well). This one is 32M parameters.GIF is on LJSpeech, very first test.
>>43519884Really good to have news and Godspeed, Anon.
>>43519884>you need some form of autoregression to model pony speech wellTo expand on this point (history time!), I spent the better part of ~late 2023 trying to make a pony VITS, which always failed despite all the upgrades I stacked on it because it was entirely non-autoregressive and thus had no way of modeling how the speech changed over time. Oct 2023 I got a decent Tacotron 2 working with the predecessor of the alignment framework I'm using now: https://zdisket.github.io/conv-att-c/
>>43517853>Can AI operate with flash/toonboom puppets yet?You can spend the time training an AI to animate flash puppets (AFAIK it hasn't been done, yet.) Or, you could spend that time making animations yourself.
10
>>43519884Very, VERY excited to use this for H:E. I think it'll be the default model shipped in the package if it's hyper-optimized for weaker graphics cards.
>didnt even noticed the thread was back 2 days ago in the catalog fuck me, I guess?>>43517853You know what I would love to see? some way for the specialized ai to operate characters inside Blender, since this one already has some python code integrated into it.
>https://u.pone.rs/cltgsbgu.mp4On kind-but-not-exactly related note, the ai video stuff has really progressed since last two years, now instead of looking whenever somebody has an extra arm or morphs into part of the furniture in the background after few seconds, instead the telltale sign is just audio being weird and not matching what is happening in the video and contextually people just behaving weirdly when interacting with anything for longer than two seconds.It's nice seeing the more tech literate normies using it to create cool shit, like the guy who is making a audiobook/drama continuation of the ST Voyager.
>https://www.youtube.com/watch?v=730zGRwbQuE>imagine
>>43517853How exactly has the last 4-5 years of AI innovation convinced you that it will never improve past what it can do today?
>>43521207Unf
>>43519884>>43520316Can you tell us more about the architecture of the TTS model itself? Is it based on anything, or is it something custom made?> Also, what upgrades does the new attention framework have?
Has anyone fine tuned an LLM on character personality yet? I'm interested in setting up a roleplay model that uses MOSS TTS for dialogue output
making love with ai mares
>>43519884Is there a repo for your version of MOSS with all the voice models?
>>43523262hot
Mares
>>43525311>11
I really wish the whatever open sourced song models were small enough to run on a 10gb gpus and not requiring 48gb for like 45 seconds of audio like they do right now
>>43524703snowpity
>>43525749What models are you referring to Anon?
>>43526330I think last open source music model was YaE (or some othe rshit with just three letters) from china. It has been a year so maybe there was so e improvements on makir it more useful but I do remember seeing early stara with "use 32GB to create 30s of music in just under 15m" .Doesn't help that suno and other ai music services really suck balls now when trying to create a proper good sounding music with them. They always are pushing for pop like sounds for some reason.
>https://github.com/Speedstu/CUDA-for-AMD-WindowsSomehow related news, for people with newer AMD cards can use and and run the actual cuda scripts without messing with the Linux processes.
>https://www.youtube.com/watch?v=U7YU0w9Mo_I>vr game project with ai thinking and ai talking maresoh shit son, its happening!
>>43527122mares are for booping
>>43527605>https://u.pone.rs/irpozdou.mp3Oh hey, the thread is back. Let me share some simple rvc cover from some time ago (song names is MartinLechowicz - ogrod eden).
>>43527730comfy
>https://u.pone.rs/cwxoechz.mp3
>>43521207I won't have to imagine, someday
>>435286503D printed mares?
>>43528647kino
>>43527605boopa the snoofa
>>43528650we will be looking forward to your project Anon. I find it weird that there isnt really anybody trying to make LLM that can both talk and interact with the outside world using robot arms/bodies.
>Breeze-TTS-2 >minimum gpu required is 12GB>can create audio at 3.1 real-time speed, with fast enough LLM is may allowed for streaming audio in such speed you can talk back and forth just like in real space>can clone audioman, all the tech changed between now and the past six years feel like a magic. we will have ai mares for rest of the time!
>>43530442:excite:
>>43527122whatever happened to blob anyway?
>>43526535>They always are pushing for pop like sounds for some reason.Probably overfitting.
amres
>>43532366yeah, I guess whomever is doing the training is grabbing whatever dataset is cheapest to make/buy, and with how much pop trash is overflowing all the radios and playlist, this would reflect the poisoning of the training process.
>https://docs.comfy.org/tutorials/audio/minimax/minimax-music-3>https://huggingface.co/Comfy-Org/MiniMax-Music-3/tree/main>ai song generating model that can supposedly work on 8gb vram using comfy uiwelp, I guess I know what I will obsessively try to mend over this weekend, sadly it doesn't look like it can generate the vocals and music as two separate files, otherwise hooking it up to rvc would have made an almost instant pony sang music.
>>43532894i love them
>>43533644could be promising
snoofa
>>43527122promising stuff anon!
>https://u.pone.rs/gwsvgfvy.mp3damn, look what I just found out, all the way from 2019
>>43536295vintage
>>43536295was this uhhh tacotron?
>>43536801I'm 100% possitive it was the tts they found before tacotron. The one that needed at least 5 hours of audio to even begin the training process .
>>43536974We've come really far.
>>43537545in mares
>>43537954amres
>>43538340
ilovethem
https://u.pone.rs/gzapkdtb.wavPlaying with new Suno. I put Pony Zone through it.
>>43539562thats pretty nice, but both the instrumental and vocals feel like they are missing 10% of extra "umpf" to make it really good.
>>43539562not bad
>>43539562Can't wait to test out the "suno at home" model to see how it really compares to the online models.
>>43539562unf rara unf
snowpity
Post from the sister thread /chag/, that Anon really cooked something really fucking good here>>43541285>I asked Opus 5.5 to create an MLP RPG video with code>https://u.pone.rs/rlvvvpsw.mp4>>43541680>It's just:>>Using only code, create a very cool video of an MLP video game RPG, SNES style, with HP bars and stuff.>It's minimalist by design, since I wanted to test some of the claims that were made about its ability to do that. I just let it run in the background with Claude Code until it finished.>>AI video generation>Keep in mind there is no AI image generation or video generation. Anthropic doesn't have those.>It wrote a SNES-style pixel renderer in code and constructed everything with that. Same for the music.>>43541456>Asked it to isolate everything. So here are the sprites, backgrounds, musics and sfx.>https://u.pone.rs/oqtrcmto.zip
>>43542043oh wow
What's the best way to get a moan or groan from SoVitS 4.0 settings?
>>43542493Holy moley.
>>43542493>sex hackerWhat is that?
Anyone want to help me with a half lewd SFW sex hacker game for possibly this upcoming gamejam? SFW meaning no genitals showing...https://pastes.io/KJVSMmia
>>43542043that sequence at the end where all the M6 came together was friggin kino. AI is slowly getting less sloppy with each new release. we will unironically start genning our own MLP episodes in 2027.
>https://www.youtube.com/watch?v=aB8prowklygponi ai stuff
>>43544142noice
>>43533644>The comfy version that has all the cuda working is not able to run the workflow due to outdated notes>The new comfy version with updated nodes is not able to run it because of incompatibility with old cuda drivers>BUT the cpu version can open the workflow just fine>after 15 minutes the model generation didn't even moved to 1%Fuck, the python module compatibility are one again kicking my ass. I will need to sit and think about this because having offline music generator sounds rad as fuck.
>>43545000Are you using the python version or the standalone one?
>>43545000>https://u.pone.rs/iqoitdyt.mp3Gentlemen, we got them, the suno at home
>>43545928>>43533644This was the initial test, using the lowest quality models from minimax music3, the first clips I made was just 5 seconds and it took 2m50s to generate some instrumental (despite adding some text for it to sing), then I generated some 10s clip and that took about 12s to makes so clearly the early time was just from model taking its time to load up. For the 30s clip it took 39s to generate however I did noticed that the node set up will crop a second or two from the end of the generated clip, so now everything has additional 2 seconds added .To load the int8 diffusion, text and vae models it took just about 20GB ram (how much vram? no fucking idea, my installation of linux decided to be a piece of shit and break whatever code that was able to monitor the vram use on my end).Im now testing the mid tier models fp16 and I can see its already too 27.6 ram, its been 20 minutes and it looks like its only just about half way done.Also one note, for whatever reason the comfyui MM3 nodes will not save the audio automatically, and when connecting to specialised mp3 nodes it will throw a pissy error about arrays not matching, so sadly it looks like I wouldn’t be able to just load up batch of 20 songs and let it run overnight (however, this could be just me problem, because I had yet to see anybody else complain about it)>>43545343on the windows I first tried doing some specific version (cant tell you which ones, it was ages ago) , but every time there was some update, modules just got over riden and things just randomly stop working. Above I was trying to use the premade zip environments that come from official website but as you seen that was unsuccessful.As for my linux set up, I just downloaded the official appimage file and let it install whatever it needed to install to make it work. Other than audio not saving automatically it looks like it all works just fine.
>https://u.pone.rs/eauqqatx.mp3...and it crashed, I guess using mid and high tier models for music generations will be limited to people with battlestation loaded with 16~32 Vram and 64Gb ram.Im not planing on dumping all the music slop I may end up making in the thread, however now that I don't have the "you can only make X gened clips per day" hard limitation hanging over my head, I can finally look into some old musical projects.Also I discovered that the audio files indeed are being saved, however insead of using the "output" file inside of its folder, like a normal program would have it set it, it created a "comfyui-shared/output" in the home directory, to dump all the files as flac. So at least that mystery was solved.
>>43546118comfy maresic anon
>>43546118Have you worked with a node based workflow before anon?
So, ElevenLabs V4 came out and I think the instant cloning might be the new SOTA compared to MOSS-TTS-PNY. The voice is insanely accurate (about 95% accuracy compared to MOSS' 100%), with the remaining inaccuracies being slight enough for RVC to be more than capable of resolving. The emotional delivery and control is a substantial step change, too. The only issue rn is that Elevenlabs, even on the >$20 creator plan, still jews you out of lossless wav output format and limits you to 192kbps mp3.Here's how it handles Pinkie, one of the hardest voices for TTS software to accurately depict: https://files.catbox.moe/fn48ru.mp3
>>43548829and here's twilight: https://files.catbox.moe/bmtibk.wav and a very brief audiobook i slopped up with both of them: https://files.catbox.moe/v9wckn.wavi was able to get wav outputs too, for some reason
>>43548829okay, listening to it again, i think SOTA is overdoing it. MOSS still mogs for now, but I think a combination of MOSS for more prosaic lines and 11labs for the more adventurous and complex lines is the ideal.
>>43548742No, this is my first time messing around with confyui (with exception of trying to run some image stuff like a year ago). The only changes to the original that I made was editing the MM3 main note to output the seed, and download some GitHub that converts the INT to text so I could use the seed numtas part of the file name .
>>43548829>>43549146one last post. I decided to crank up the similarity slider all the way to 100 (it was at 75% before), and i think it sounds even better than before. also cloned dashie and rarity's voice, and they sound pretty swell, too.https://files.catbox.moe/85peyw.mp3https://files.catbox.moe/icynpq.mp3https://files.catbox.moe/9lajv1.mp3https://files.catbox.moe/8ddwul.mp3
>>43549657waow accurate
>>43549657not gonna lie, this sounds almost perfect if not for some odd focusing on some parts of the words, pauses (or lack of them) and the flow of speech, but still its like 95% sounds like it could be some kind of unused clips from show archives.
mares
>>43551422This
>>43549657I don't often check this thread (only when I see it on the front page), but if this is the level what can be achieved then damn. Pretty good. If you ask me to differentiate between this and the actual VAs I would be in trouble.
>>43552205yeah, now all we need is that anon that was working on a program that converts normal stories to script-like-format to create a version of this app that could work without all the python jank and we are almost near the steps of having an automated pipeline of fics being converted to professionally sounding audiobooks.
>>43548829>>43549657These are amazing, way beyond anything I ever imagined would be possible way back when PPP started. Even if it's in the hands of the Jews for now it's super encouraging to hear this level of quality, especially if it doesn't require a reference voice recording as input.How much effort did it take to put them together? All generated from one take or did you need to edit together several individual gens?
>>43549657>7 years later the only way AI voices sound good is via a closed paywalled platformFuck this shit takes too long, I thought we'd fully own mare voices by now
>>43552909Just one single take, crazily enough. I did have gemini place inline emotion tags intelligently in the generated story, though, which is partly responsible for the expressiveness. i'll drop the dataset i used for instant cloning tomorrow once i finish the rest of the M6, if anyone wants to jew out the bux for 11labs
>>43515527>>Do you have a Code of Conduct?>Of course: 15.ai/codelol xd lmao
>>43553823We should REALLY update the thread OP
>>43552949>I thought we'd fully own mare voices by nowwe did
more conversation tests. Twilight and Pinkie streaming Minecrafthttps://files.catbox.moe/g9y34c.mp3https://files.catbox.moe/ekzsb8.mp3
>>43552949Delta stuff is like 94% there
>>43554149neat!
>>43554360yup
>>43552949*Closed paywall platforms that use our incredible amount of fan labor as a basis
>>43519884Kickass man! Thanks!
>>43519884+1 for interest in a smaller tts model
>>43519884BLESSED
>>43520720>specialized ai to operate characters inside BlenderLike create animations or use as an avatar?
>>43516297kino
>>43548829>1:30>2:10holy fuckPinkie's laughter is unironically one of the most precious sounds to me and it's been so hard to clone historicallyif you're telling me an AI can finally just let me listen to her being all giggly and sweet I'm gonna fucking cry
>https://u.pone.rs/booklrom.mp3BEHOLD, first /mlp/ song done with minimax 3 (and rvc/sovit voice conversion). Lyrics written by >>43544247 Anon
>>43558702shit, I fucked up, that was mppp OP, lyrics are here its the >>43552239 poster
>>43558702>>43558715Overall it is good! A few words got pronounced pretty badly (bureaucratic, requirement), and I expected the last sentence to close the song completely (instrumentals already closing at "Raise your glass to Pissfilly").Maybe some anon could rewrite the lyrics a bit so it flows better.
>>43559281Yeah, there are some problems with the two generated test pieces were it swapped few sentences around. I feel like its good chance its caused by the use of the low tier int8 model, sadly I don't have a 24gb vram to test out the large minimax 3 music models to compare the quality (some Anons on /g/ stated that it has clearly reach the same quality as suno outputs from last year).
>>43554149These were super cute. I listened last night in bed with the other ones. Thank you so much for making them, Anon! All of the audio files got put into today's rewatch Cytube by the way, people were overall very impressed with the high quality. Super exciting times.
>>43542043thanks for the zip anon!
>>43558702Great stuff Anon!
>>43528650If I don't get to fuck a real 3d printed llm powered and ai voiced pony by this time next year I will be incredibly blueballed.