[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/mlp/ - Pony

Name
Spoiler?[]
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
Flag
File[]
  • Please read the Rules and FAQ before posting.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


File: altOP.jpg (6 KB, 250x176)
6 KB JPG
Welcome to the Pony Voice Preservation Project!
http://youtu.be/730zGRwbQuE

The Pony Preservation Project is a collaborative effort by /mlp/ to build and curate pony datasets for as many applications in AI as possible.

Technology has progressed such that a trained neural network can generate convincing voice clips, drawings and text for any person or character using existing audio recordings, artwork and fanfics as a reference. As you can surely imagine, AI pony voices, drawings and text have endless applications for pony content creation.

AI is incredibly versatile, basically anything that can be boiled down to a simple dataset can be used for training to create more of it. AI-generated images, fanfics, wAIfu chatbots and even animation are possible, and are being worked on here.

Any anon is free to join, and there are many active tasks that would suit any level of technical expertise. If you’re interested in helping out, take a look at the quick start guide linked below and ask in the thread for any further detail you need.

EQG and G5 are not welcome.

>Quick start guide:
docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Introduction to the PPP, links to text-to-speech tools, and how (You) can help with active tasks.

>The main Doc:
docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit
An in-depth repository of tutorials, resources and archives.

>Online speech generation
haysay.ai

>MOSS-TTS-1.7B-PNY v0.1: a finetune of MOSS-TTS + custom vocoder for 48KHz audio
HuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNY
Colab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBB
Public demo: http://198.53.64.194:35029/

>Active tasks:
Research into animation AI
Research into pony image generation

>Latest developments:
pastebin.com/4p00iUZM (embed)

>The PoneAI drive, an archive for AI pony voice content:
drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp

>Clipper’s Master Files, the central location for MLP voice data:
mega.nz/folder/jkwimSTa#_xk0VnR30C8Ljsy4RCGSig
mega.nz/folder/gVYUEZrI#6dQHH3P2cFYWm3UkQveHxQ
drive.google.com/drive/folders/1MuM9Nb_LwnVxInIPFNvzD_hv3zOZhpwx

>Cool, where is the discord/forum/whatever unifying place for this project?
You're looking at it.

Last Thread: >>43451491
>>
FAQs:
If your question isn’t listed here, take a look in the quick start guide and main doc to see if it’s already answered there. Use the tabs on the left for easy navigation.
Quick: docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Main: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit

>Where can I find the AI text-to-speech tools and how do I use them?
A list of TTS tools: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.yuhl8zjiwmwq
How to get the best out of them: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.mnnpknmj1hcy

>Where can I find content made with the voice AI?
In the PoneAI drive: drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp
And the PPP Mega Compilation: docs.google.com/spreadsheets/d/1T2TE3OBs681Vphfas7Jgi5rvugdH6wnXVtUVYiZyJF8/edit

>I want to know more about the PPP, but I can’t be arsed to read the doc.
See the live PPP panel shows presented on /mlp/con for a more condensed overview.
2020 pony.tube/w/5fUkuT3245pL8ZoWXUnXJ4
2021 pony.tube/w/a5yfTV4Ynq7tRveZH7AA8f
2022 pony.tube/w/mV3xgbdtrXqjoPAwEXZCw5
2023 pony.tube/w/fVZShksjBbu6uT51DtvWWz

>How can I help with the PPP?
Build datasets, train AIs, and use the AI to make more pony content. Take a look at the quick start guide for current active tasks, or start your own in the thread if you have an idea. There’s always more data to collect and more AIs to train.

>Did you know that such and such voiced this other thing that could be used for voice data?
It is best to keep to official audio only unless there is very little of it available. If you know of a good source of audio for characters with few (or just fewer) lines, please post it in the thread. 5.1 is generally required unless you have a source already clean of background noise. Preferably post a sample or link. The easier you make it, the more likely it will be done.

>What about fan-imitations of official voices?
No.

>Will you guys be doing a [insert language here] version of the AI?
Probably not, but you're welcome to. You can however get most of the way there by using phonetic transcriptions of other languages as input for the AI.

>What about [insert OC here]'s voice?
It is often quite difficult to find good quality audio data for OCs. If you happen to know any, post them in the thread and we’ll take a look.

>I have an idea!
Great. Post it in the thread and we'll discuss it.

>Do you have a Code of Conduct?
Of course: 15.ai/code

>Is this project open source? Who is in charge of this?
pony.tube/w/mqJyvdgrpbWgZduz2cs1Cm

PPP Redubs:
pony.tube/w/p/aR2dpAFn5KhnqPYiRxFQ97

Stream Premieres:
pony.tube/w/6cKnjJEZSCi3gsvrbATXnC
pony.tube/w/oNeBFMPiQKh93ePqTz1ns8
>>
Any updates on MOSS?
>>
>9
>>
File: a1183363622_2.jpg (25 KB, 350x350)
25 KB JPG
>>43515526
https://www.youtube.com/watch?v=FziZYANzCls
https://u.pone.rs/rfxsptss.zip
I made an album recently that uses Delta's new RVC NG for the between-song dialogue. It's actually pretty nice, especially since it's actually possible to get more varied energy out of it for yelling and whispering and such, whereas the old RVC struggled with that. Old RVC is still on the singing though.
>>
>>43516297
Nice!
>>
>10
>>
Can AI operate with flash/toonboom puppets yet? Because now it's clear that generating sloppy shit from scratch will never git gud.
>>
>>43517328
>>
Last bump.
>>
>>43517853
ToonBoom has a product called Ember which integrates with Harmony and does some very limited AI stuff, like extending background images and generating scratch tracks (roughouts of character dialogue that you use for timing your animation until the professional voicework is done). As for 3rd-party AI integration with Toon Boom Harmony, I know nothing.
>>
>>43518906
is it good
>>
File: alignment_progress_fast.gif (631 KB, 1200x400)
631 KB GIF
>>43515531
Nothing on MOSS for now, I've been noticing a lot of people could use something more lightweight for gamedev or other purposes, so last night I began developing a small non-autoregressive model (well, kind of semi-autoregressive because it uses GRUs -- you need some form of autoregression to model pony speech well). This one is 32M parameters.
GIF is on LJSpeech, very first test.
>>
>>43519884
Really good to have news and Godspeed, Anon.
>>
>>43519884
>you need some form of autoregression to model pony speech well
To expand on this point (history time!), I spent the better part of ~late 2023 trying to make a pony VITS, which always failed despite all the upgrades I stacked on it because it was entirely non-autoregressive and thus had no way of modeling how the speech changed over time. Oct 2023 I got a decent Tacotron 2 working with the predecessor of the alignment framework I'm using now: https://zdisket.github.io/conv-att-c/
>>
>>43517853
>Can AI operate with flash/toonboom puppets yet?
You can spend the time training an AI to animate flash puppets (AFAIK it hasn't been done, yet.) Or, you could spend that time making animations yourself.
>>
10
>>
>>43519884
Very, VERY excited to use this for H:E. I think it'll be the default model shipped in the package if it's hyper-optimized for weaker graphics cards.
>>
File: pip 1719060717748766.png (1.74 MB, 2000x2000)
1.74 MB PNG
>didnt even noticed the thread was back 2 days ago in the catalog
fuck me, I guess?
>>43517853
You know what I would love to see? some way for the specialized ai to operate characters inside Blender, since this one already has some python code integrated into it.
>>
>https://u.pone.rs/cltgsbgu.mp4
On kind-but-not-exactly related note, the ai video stuff has really progressed since last two years, now instead of looking whenever somebody has an extra arm or morphs into part of the furniture in the background after few seconds, instead the telltale sign is just audio being weird and not matching what is happening in the video and contextually people just behaving weirdly when interacting with anything for longer than two seconds.
It's nice seeing the more tech literate normies using it to create cool shit, like the guy who is making a audiobook/drama continuation of the ST Voyager.
>>
File: other 1360359.png (1020 KB, 3000x4815)
1020 KB PNG
>https://www.youtube.com/watch?v=730zGRwbQuE
>imagine
>>
>>43517853
How exactly has the last 4-5 years of AI innovation convinced you that it will never improve past what it can do today?
>>
>>43521207
Unf
>>
>>43519884
>>43520316

Can you tell us more about the architecture of the TTS model itself? Is it based on anything, or is it something custom made?

> Also, what upgrades does the new attention framework have?
>>
Has anyone fine tuned an LLM on character personality yet? I'm interested in setting up a roleplay model that uses MOSS TTS for dialogue output
>>
making love with ai mares
>>
>>43519884
Is there a repo for your version of MOSS with all the voice models?
>>
>>43523262
hot
>>
Mares
>>
>10
>>
>>43525311
>11
>>
I really wish the whatever open sourced song models were small enough to run on a 10gb gpus and not requiring 48gb for like 45 seconds of audio like they do right now
>>
>>43524703
snowpity
>>
>>43525749
What models are you referring to Anon?
>>
>>43526330
I think last open source music model was YaE (or some othe rshit with just three letters) from china. It has been a year so maybe there was so e improvements on makir it more useful but I do remember seeing early stara with "use 32GB to create 30s of music in just under 15m" .
Doesn't help that suno and other ai music services really suck balls now when trying to create a proper good sounding music with them. They always are pushing for pop like sounds for some reason.
>>
>https://github.com/Speedstu/CUDA-for-AMD-Windows
Somehow related news, for people with newer AMD cards can use and and run the actual cuda scripts without messing with the Linux processes.
>>
>https://www.youtube.com/watch?v=U7YU0w9Mo_I
>vr game project with ai thinking and ai talking mares
oh shit son, its happening!
>>
File: boop 1780899346428650.png (1.01 MB, 2400x2000)
1.01 MB PNG
>>43527122
mares are for booping
>>
>>43527605
>https://u.pone.rs/irpozdou.mp3
Oh hey, the thread is back. Let me share some simple rvc cover from some time ago (song names is MartinLechowicz - ogrod eden).
>>
>>43527730
comfy
>>
>https://u.pone.rs/cwxoechz.mp3
>>
File: 20260208_215353922.jpg (965 KB, 4000x3000)
965 KB JPG
>>43521207
I won't have to imagine, someday
>>
>>43528650
3D printed mares?
>>
>>43528647
kino
>>
>>43527605
boopa the snoofa
>>
>>43528650
we will be looking forward to your project Anon. I find it weird that there isnt really anybody trying to make LLM that can both talk and interact with the outside world using robot arms/bodies.
>>
>Breeze-TTS-2
>minimum gpu required is 12GB
>can create audio at 3.1 real-time speed, with fast enough LLM is may allowed for streaming audio in such speed you can talk back and forth just like in real space
>can clone audio
man, all the tech changed between now and the past six years feel like a magic. we will have ai mares for rest of the time!
>>
>>43530442
:excite:
>>
>>43527122
whatever happened to blob anyway?
>>
>>43526535
>They always are pushing for pop like sounds for some reason.
Probably overfitting.
>>
amres
>>
>>43532366
yeah, I guess whomever is doing the training is grabbing whatever dataset is cheapest to make/buy, and with how much pop trash is overflowing all the radios and playlist, this would reflect the poisoning of the training process.
>>
>https://docs.comfy.org/tutorials/audio/minimax/minimax-music-3
>https://huggingface.co/Comfy-Org/MiniMax-Music-3/tree/main
>ai song generating model that can supposedly work on 8gb vram using comfy ui
welp, I guess I know what I will obsessively try to mend over this weekend, sadly it doesn't look like it can generate the vocals and music as two separate files, otherwise hooking it up to rvc would have made an almost instant pony sang music.
>>
>>43532894
i love them
>>
>>43533644
could be promising
>>
snoofa
>>
>>43527122
promising stuff anon!
>>
>https://u.pone.rs/gwsvgfvy.mp3
damn, look what I just found out, all the way from 2019
>>
>>43536295
vintage
>>
>>43536295
was this uhhh tacotron?
>>
>>43536801
I'm 100% possitive it was the tts they found before tacotron. The one that needed at least 5 hours of audio to even begin the training process .
>>
>>43536974
We've come really far.
>>
>>43537545
in mares
>>
>>43537954
amres
>>
>>43538340
>>
ilovethem
>>
File: RaraEyes2.png (479 KB, 1800x1800)
479 KB PNG
https://u.pone.rs/gzapkdtb.wav
Playing with new Suno. I put Pony Zone through it.
>>
>>43539562
thats pretty nice, but both the instrumental and vocals feel like they are missing 10% of extra "umpf" to make it really good.
>>
>>43539562
not bad
>>
>>43539562
Can't wait to test out the "suno at home" model to see how it really compares to the online models.
>>
>>43539562
unf rara unf
>>
snowpity
>>
Post from the sister thread /chag/, that Anon really cooked something really fucking good here
>>43541285
>I asked Opus 5.5 to create an MLP RPG video with code
>https://u.pone.rs/rlvvvpsw.mp4
>>43541680
>It's just:
>>Using only code, create a very cool video of an MLP video game RPG, SNES style, with HP bars and stuff.
>It's minimalist by design, since I wanted to test some of the claims that were made about its ability to do that. I just let it run in the background with Claude Code until it finished.
>>AI video generation
>Keep in mind there is no AI image generation or video generation. Anthropic doesn't have those.
>It wrote a SNES-style pixel renderer in code and constructed everything with that. Same for the music.
>>43541456
>Asked it to isolate everything. So here are the sprites, backgrounds, musics and sfx.
>https://u.pone.rs/oqtrcmto.zip
>>
>>43542043
oh wow
>>
What's the best way to get a moan or groan from SoVitS 4.0 settings?
>>
>>43542493
Holy moley.
>>
>>43542493
>sex hacker
What is that?
>>
Anyone want to help me with a half lewd SFW sex hacker game for possibly this upcoming gamejam? SFW meaning no genitals showing...
https://pastes.io/KJVSMmia
>>
>>43542043
that sequence at the end where all the M6 came together was friggin kino. AI is slowly getting less sloppy with each new release. we will unironically start genning our own MLP episodes in 2027.
>>
>https://www.youtube.com/watch?v=aB8prowklyg
poni ai stuff
>>
>>43544142
noice
>>
>>43533644
>The comfy version that has all the cuda working is not able to run the workflow due to outdated notes
>The new comfy version with updated nodes is not able to run it because of incompatibility with old cuda drivers
>BUT the cpu version can open the workflow just fine
>after 15 minutes the model generation didn't even moved to 1%
Fuck, the python module compatibility are one again kicking my ass. I will need to sit and think about this because having offline music generator sounds rad as fuck.
>>
>>43545000
Are you using the python version or the standalone one?
>>
>>43545000
>https://u.pone.rs/iqoitdyt.mp3
Gentlemen, we got them, the suno at home
>>
>>43545928
>>43533644
This was the initial test, using the lowest quality models from minimax music3, the first clips I made was just 5 seconds and it took 2m50s to generate some instrumental (despite adding some text for it to sing), then I generated some 10s clip and that took about 12s to makes so clearly the early time was just from model taking its time to load up. For the 30s clip it took 39s to generate however I did noticed that the node set up will crop a second or two from the end of the generated clip, so now everything has additional 2 seconds added .
To load the int8 diffusion, text and vae models it took just about 20GB ram (how much vram? no fucking idea, my installation of linux decided to be a piece of shit and break whatever code that was able to monitor the vram use on my end).
Im now testing the mid tier models fp16 and I can see its already too 27.6 ram, its been 20 minutes and it looks like its only just about half way done.
Also one note, for whatever reason the comfyui MM3 nodes will not save the audio automatically, and when connecting to specialised mp3 nodes it will throw a pissy error about arrays not matching, so sadly it looks like I wouldn’t be able to just load up batch of 20 songs and let it run overnight (however, this could be just me problem, because I had yet to see anybody else complain about it)
>>43545343
on the windows I first tried doing some specific version (cant tell you which ones, it was ages ago) , but every time there was some update, modules just got over riden and things just randomly stop working. Above I was trying to use the premade zip environments that come from official website but as you seen that was unsuccessful.
As for my linux set up, I just downloaded the official appimage file and let it install whatever it needed to install to make it work. Other than audio not saving automatically it looks like it all works just fine.
>>
>https://u.pone.rs/eauqqatx.mp3
...and it crashed, I guess using mid and high tier models for music generations will be limited to people with battlestation loaded with 16~32 Vram and 64Gb ram.
Im not planing on dumping all the music slop I may end up making in the thread, however now that I don't have the "you can only make X gened clips per day" hard limitation hanging over my head, I can finally look into some old musical projects.
Also I discovered that the audio files indeed are being saved, however insead of using the "output" file inside of its folder, like a normal program would have it set it, it created a "comfyui-shared/output" in the home directory, to dump all the files as flac. So at least that mystery was solved.
>>
>>43546118
comfy maresic anon
>>
>>43546118
Have you worked with a node based workflow before anon?
>>
So, ElevenLabs V4 came out and I think the instant cloning might be the new SOTA compared to MOSS-TTS-PNY. The voice is insanely accurate (about 95% accuracy compared to MOSS' 100%), with the remaining inaccuracies being slight enough for RVC to be more than capable of resolving. The emotional delivery and control is a substantial step change, too. The only issue rn is that Elevenlabs, even on the >$20 creator plan, still jews you out of lossless wav output format and limits you to 192kbps mp3.

Here's how it handles Pinkie, one of the hardest voices for TTS software to accurately depict: https://files.catbox.moe/fn48ru.mp3
>>
>>43548829
and here's twilight: https://files.catbox.moe/bmtibk.wav
and a very brief audiobook i slopped up with both of them: https://files.catbox.moe/v9wckn.wav
i was able to get wav outputs too, for some reason
>>
>>43548829
okay, listening to it again, i think SOTA is overdoing it. MOSS still mogs for now, but I think a combination of MOSS for more prosaic lines and 11labs for the more adventurous and complex lines is the ideal.
>>
>>43548742
No, this is my first time messing around with confyui (with exception of trying to run some image stuff like a year ago). The only changes to the original that I made was editing the MM3 main note to output the seed, and download some GitHub that converts the INT to text so I could use the seed numtas part of the file name .
>>
>>43548829
>>43549146
one last post. I decided to crank up the similarity slider all the way to 100 (it was at 75% before), and i think it sounds even better than before. also cloned dashie and rarity's voice, and they sound pretty swell, too.

https://files.catbox.moe/85peyw.mp3
https://files.catbox.moe/icynpq.mp3
https://files.catbox.moe/9lajv1.mp3
https://files.catbox.moe/8ddwul.mp3
>>
>>43549657
waow accurate
>>
>>43549657
not gonna lie, this sounds almost perfect if not for some odd focusing on some parts of the words, pauses (or lack of them) and the flow of speech, but still its like 95% sounds like it could be some kind of unused clips from show archives.
>>
mares
>>
>>43551422
This
>>
>>43549657
I don't often check this thread (only when I see it on the front page), but if this is the level what can be achieved then damn. Pretty good. If you ask me to differentiate between this and the actual VAs I would be in trouble.
>>
>>43552205
yeah, now all we need is that anon that was working on a program that converts normal stories to script-like-format to create a version of this app that could work without all the python jank and we are almost near the steps of having an automated pipeline of fics being converted to professionally sounding audiobooks.
>>
>>43548829
>>43549657
These are amazing, way beyond anything I ever imagined would be possible way back when PPP started. Even if it's in the hands of the Jews for now it's super encouraging to hear this level of quality, especially if it doesn't require a reference voice recording as input.

How much effort did it take to put them together? All generated from one take or did you need to edit together several individual gens?
>>
>>43549657
>7 years later the only way AI voices sound good is via a closed paywalled platform
Fuck this shit takes too long, I thought we'd fully own mare voices by now
>>
>>43552909
Just one single take, crazily enough. I did have gemini place inline emotion tags intelligently in the generated story, though, which is partly responsible for the expressiveness.

i'll drop the dataset i used for instant cloning tomorrow once i finish the rest of the M6, if anyone wants to jew out the bux for 11labs
>>
snoofa
>>
>>43515527
>>Do you have a Code of Conduct?
>Of course: 15.ai/code
lol xd lmao
>>
>>43553823
We should REALLY update the thread OP
>>
>>43552949
>I thought we'd fully own mare voices by now
we did
>>
more conversation tests. Twilight and Pinkie streaming Minecraft
https://files.catbox.moe/g9y34c.mp3
https://files.catbox.moe/ekzsb8.mp3
>>
>>43552949
Delta stuff is like 94% there
>>
>>43554149
neat!
>>
>>43554360
yup
>>
>>43552949
*Closed paywall platforms that use our incredible amount of fan labor as a basis
>>
>>43519884
Kickass man! Thanks!
>>
>>43519884
+1 for interest in a smaller tts model
>>
>>43519884
BLESSED
>>
>>43520720
>specialized ai to operate characters inside Blender
Like create animations or use as an avatar?
>>
>>43516297
kino
>>
>>43548829
>1:30
>2:10
holy fuck
Pinkie's laughter is unironically one of the most precious sounds to me and it's been so hard to clone historically
if you're telling me an AI can finally just let me listen to her being all giggly and sweet I'm gonna fucking cry
>>
>https://u.pone.rs/booklrom.mp3
BEHOLD, first /mlp/ song done with minimax 3 (and rvc/sovit voice conversion). Lyrics written by >>43544247 Anon
>>
>>43558702
shit, I fucked up, that was mppp OP, lyrics are here its the >>43552239 poster
>>
>>43558702
>>43558715
Overall it is good!
A few words got pronounced pretty badly (bureaucratic, requirement), and I expected the last sentence to close the song completely (instrumentals already closing at "Raise your glass to Pissfilly").
Maybe some anon could rewrite the lyrics a bit so it flows better.
>>
>>43559281
Yeah, there are some problems with the two generated test pieces were it swapped few sentences around. I feel like its good chance its caused by the use of the low tier int8 model, sadly I don't have a 24gb vram to test out the large minimax 3 music models to compare the quality (some Anons on /g/ stated that it has clearly reach the same quality as suno outputs from last year).
>>
>>43554149
These were super cute. I listened last night in bed with the other ones. Thank you so much for making them, Anon! All of the audio files got put into today's rewatch Cytube by the way, people were overall very impressed with the high quality. Super exciting times.
>>
>>43542043
thanks for the zip anon!
>>
>>43558702
Great stuff Anon!
>>
>>43528650
If I don't get to fuck a real 3d printed llm powered and ai voiced pony by this time next year I will be incredibly blueballed.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.