[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/mlp/ - Pony

Name
Spoiler?[]
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
Flag
File[]
  • Please read the Rules and FAQ before posting.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: altOP.jpg (1.26 MB, 2119x1500)
1.26 MB JPG
Welcome to the Pony Voice Preservation Project!
youtu.be/730zGRwbQuE

The Pony Preservation Project is a collaborative effort by /mlp/ to build and curate pony datasets for as many applications in AI as possible.

Technology has progressed such that a trained neural network can generate convincing voice clips, drawings and text for any person or character using existing audio recordings, artwork and fanfics as a reference. As you can surely imagine, AI pony voices, drawings and text have endless applications for pony content creation.

AI is incredibly versatile, basically anything that can be boiled down to a simple dataset can be used for training to create more of it. AI-generated images, fanfics, wAIfu chatbots and even animation are possible, and are being worked on here.

Any anon is free to join, and there are many active tasks that would suit any level of technical expertise. If you’re interested in helping out, take a look at the quick start guide linked below and ask in the thread for any further detail you need.

EQG and G5 are not welcome.

>Quick start guide:
docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Introduction to the PPP, links to text-to-speech tools, and how (You) can help with active tasks.

>The main Doc:
docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit
An in-depth repository of tutorials, resources and archives.

>Online speech generation
haysay.ai

>Active tasks:
Research into animation AI
Research into pony image generation

>Latest developments:
pastebin.com/4p00iUZM

>The PoneAI drive, an archive for AI pony voice content:
drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp

>Clipper’s Master Files, the central location for MLP voice data:
mega.nz/folder/jkwimSTa#_xk0VnR30C8Ljsy4RCGSig
mega.nz/folder/gVYUEZrI#6dQHH3P2cFYWm3UkQveHxQ
drive.google.com/drive/folders/1MuM9Nb_LwnVxInIPFNvzD_hv3zOZhpwx

>Cool, where is the discord/forum/whatever unifying place for this project?
You're looking at it.

Last Thread: https://desuarchive.org/mlp/thread/43127073/#43127073
>>
FAQs:
If your question isn’t listed here, take a look in the quick start guide and main doc to see if it’s already answered there. Use the tabs on the left for easy navigation.
Quick: docs.google.com/document/d/1PDkSrKKiHzzpUTKzBldZeKngvjeBUjyTtGCOv2GWwa0/edit
Main: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit

>Where can I find the AI text-to-speech tools and how do I use them?
A list of TTS tools: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.yuhl8zjiwmwq
How to get the best out of them: docs.google.com/document/d/1y1pfS0LCrwbbvxdn3ZksH25BKaf0LaO13uYppxIQnac/edit#heading=h.mnnpknmj1hcy

>Where can I find content made with the voice AI?
In the PoneAI drive: drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp
And the PPP Mega Compilation: docs.google.com/spreadsheets/d/1T2TE3OBs681Vphfas7Jgi5rvugdH6wnXVtUVYiZyJF8/edit

>I want to know more about the PPP, but I can’t be arsed to read the doc.
See the live PPP panel shows presented on /mlp/con for a more condensed overview.
2020 pony.tube/w/5fUkuT3245pL8ZoWXUnXJ4
2021 pony.tube/w/a5yfTV4Ynq7tRveZH7AA8f
2022 pony.tube/w/mV3xgbdtrXqjoPAwEXZCw5
2023 pony.tube/w/fVZShksjBbu6uT51DtvWWz

>How can I help with the PPP?
Build datasets, train AIs, and use the AI to make more pony content. Take a look at the quick start guide for current active tasks, or start your own in the thread if you have an idea. There’s always more data to collect and more AIs to train.

>Did you know that such and such voiced this other thing that could be used for voice data?
It is best to keep to official audio only unless there is very little of it available. If you know of a good source of audio for characters with few (or just fewer) lines, please post it in the thread. 5.1 is generally required unless you have a source already clean of background noise. Preferably post a sample or link. The easier you make it, the more likely it will be done.

>What about fan-imitations of official voices?
No.

>Will you guys be doing a [insert language here] version of the AI?
Probably not, but you're welcome to. You can however get most of the way there by using phonetic transcriptions of other languages as input for the AI.

>What about [insert OC here]'s voice?
It is often quite difficult to find good quality audio data for OCs. If you happen to know any, post them in the thread and we’ll take a look.

>I have an idea!
Great. Post it in the thread and we'll discuss it.

>Do you have a Code of Conduct?
Of course: 15.ai/code

>Is this project open source? Who is in charge of this?
pony.tube/w/mqJyvdgrpbWgZduz2cs1Cm

PPP Redubs:
pony.tube/w/p/aR2dpAFn5KhnqPYiRxFQ97

Stream Premieres:
pony.tube/w/6cKnjJEZSCi3gsvrbATXnC
pony.tube/w/oNeBFMPiQKh93ePqTz1ns8
>>
Quick thread to see the state of things.

15.ai is now dead forever.
https://x.com/fifteenai/status/2060098921582772567
>>
Is there a dataset describing the scenes of the show, kinda like a transcript but with detailed information on who does what?
>>
>>43289213
I don't think there would be such thing for mlp (otherwise it would be scrapped for data years ago), I guess the closes would be to try rip off netflix and other tv shows that have the Audio Description/Video Description tracks for the blind people, and see if that could be somehow used as dataset in whatever project you are planing to do.
>>
>>43289085
This MF just abandons every project of his, doesn't he
He'll abandon that marketplace soon enough and then announce a brand new project that he totally won't abandon, bros!
>>
File: coolshit.jpg (1.35 MB, 7096x2020)
1.35 MB JPG
>>43289238
I doubt it would be detailed enough. I heard multimodal llms can take in video, even small ones like the ones I can run on my machine, so one could theoretically tag the entire show like this.

That aside, picrel is a compilation of experiments I did a while ago while playing with show frame compression. I trained a hierarchical VQ autoencoder that encodes 256x384 frames into two maps: 16x24 and 8x12 codebook indices, each codebook has 1024 entries, and then tried to generate larger map given smaller map using discrete diffusion. It's quite undertrained but I got bored of it. Just thought you guys would appreciate the abominations, some of them are even cute.
>>
>>43289447
I think I saw some models months ago that were able to watch 10s video and describe what was happening in a scene along with any interaction people had with the setting, I wish I could remember that the name of it was since that sounds like something you could potentially reuse for your stuff.
Now that I type all of this, I kind of wish there was a program that could easy way to create automated audiobook from a fic, but that would require for tts model to be combined with some llm to understand which characters are included in the story and automatically swap the voices of character/narrator as well as add any relevant background music and sound effects.
>>
>>43289481
>easy way to create automated audiobook from a fic
I made such app a while ago https://files.catbox.moe/cwj64u.mp4
https://drive.google.com/drive/folders/14zMbURz1SuYNMoewX88EjR8sHEkcaKXa
>>
>>43289508
>if yours only supports cuda 11.x but you still want to run on gpu, run the following inside PVT folder: runtime/python.exe -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
Could it be possible to make it work with the cu116? My system+gpu can't get python to work with anything above that, and Im not going to be upgrading my pc for at least next two/three years.
>>
>>43289524
Replace cu118 with cu116 and try it, should work.
>>
Oh hey, the lights are back on here?
>>
does anyone know if there was an upgrade in openly available musical ai models or is the yue music model from last/two years ago still the "best" option?
>>
What would people think about trying to make some kind of collaborative "album", with Anons posting their pony songs made with whatever ai model/service of their choice every Friday?
>>
>>43291450
The idea sounds interesting. I'm not sure if there are enough Anons who are actively genning to make that system work though.
>>
Up.
>>
File: angrycelly.png (447 KB, 880x706)
447 KB PNG
Hello! Maybe someone here could help me find a certain audio file I'm looking for. It's Princess Celestia saying "And so, as punishment for your insolence and treachery, I hereby sentence you to death by decapitation! May the blade fall swiftly and your head be raised high as a warning to all those who would dare to defy the will of the crown." I can't remember whether it was AI generated or audio from Nicole Oliver herself or a fan dub, but I remember that the quality was very good. I know I listened to it in November 2024, but it may have been posted way before that. I scoured the PPP threads on desuarchive and Clipper's master files and did not find it.
>>
>>43293091
>The PoneAI drive, an archive for AI pony voice content:
drive.google.com/drive/folders/1E21zJQWC5XVQWy2mt42bUiJ_XbqTJXCp
Try here, if it was genned in a thread then it'll be found in here.
>>
>>43293096
Thank you, Anon. There's lots of good audio in there but, alas, I could not find the file.
>>
>>43293091
I found it! I can't believe I missed it the first time; it's right there on desu in Nov 2024 just like I said, >>41653178
Direct link to the audio, for anyone interested:
https://vocaroo.com/14eyuFuDu0Zs
>>
>>43293091
>>43293256
That sounds pretty OOC for Celly.
>>
nine
>>
>>43293971
plus one
>>
>cant use raw audio clip because singing reverb fucks with the voice conversion
>cant use the DeReverbed audio because random bits of words are just nuked from the clip
dear song makers, plz stop using aggressive autotune and fake ass reverb, its making finding a good song for ai pony covers way more difficult than it needs to be.
>>
>https://u.pone.rs/basikmdv.mp3
>technobridge523 - BLACK SUNSET (Cover)
What a strange feeling, bumping into a brand new ai cover of old original ai song made in ppp thread from two years ago.
>>
What's the latest and greatest in local voice gen?
>>
we are so dead
>>
>>43289079
We drove this deep into the heart of the Archives. We thought it was dead. We were wrong.
>>
>>43297041
we are so alive
>>
Gpt sovits on haysay is fucked and keeps giving an error. The only decent one that allows text to speech and has all the good characters.
>>
>>43297664
Hmm, Im guessing HydrusBeta amazon server decided to crash out on that specific tts code , which sucks balls. Im guessing it wouldn't be helpful saying there is local ui installation of the Gpt sovits that can run pretty OK on 8gb vram? worse come to worse, you could always request audio clip form anons in the thread here, as long as it's pony related it would be fine
>>
>>43297664
There was a JSON file that got into an corrupted state. I've regenerated it, and I think that fixed the issue. Let me know if the problem persists.
>>
>>43297705
I've been working around it by using the characters on the tts one and converting them into the other voices with mixed since they require an audio input.

>>43297961
This worked, thank you anon.
>>
>>43297961
Thanks for the fix.
>>
>>43297961
Thanks anon!
>>
>>43297961
noice one, m8!
>>
>>43289085
Didn't he say he was going to open source this stuff eventually? Now would be a good time.
>>
>>43299651
I dont think he ever has, and even posting on ppp was very sporadic in 2019/2020. It would be cool if he had, or at least maybe create the documents showing exact steps of how it was made so people with the know-how would be able to recreate it with the modern tech.
>>
>>43299651
>>43300095
I'm certain he at least stated several times that he was gonna release a paper about his research, which he never did, and likely never will.
>>
>>43300117
Damn
>>
>>43289079
hey check out my fuckass tts. I'll release model and stuff for realsies soon
samples:
https://files.catbox.moe/b5asjw.wav
https://files.catbox.moe/2vk68m.wav
demo:
http://198.53.64.194:35029/
>>
File: Mare ooo.gif (428 KB, 235x274)
428 KB GIF
>>43300755
>>
>>43300755

https://files.catbox.moe/hpg3m1.wav

https://files.catbox.moe/xrrt16.wav

https://files.catbox.moe/uz1fyc.wav
>>
>>43300755
Is this your own arch or did you tune something?
>>
>>43301171
Both. I'll explain later.
>>
>>43300755
>https://files.catbox.moe/x2bt7n.wav
hey, thats neat. I can do a tf2 audio shitposting again.
>>
>>43300755
This is really, really good - at least, from the small samples I've generated. And I bet it sounds even better if you know what the hell you're doing with top_p, temperature, and those other autistic sliders. Is there a release window?
>>
>>43301579
>>43300755
Also, just my intuition, but did you use S1 voice clips for your model? If so, bravo.
>>
>>43301579
Tomorrow at latest, tonight at earliest.
>>
>>43301712
Awesome. Tested it a bit more and it'll probably be my daily driver for AI TTS. Will you be posting guides and whatnot, especially for things like interjections (hm, huh, wow, etc.)? Doing those just result in long silences; makes me think it might be bugged or something. And is the 30s audio output limit for the testing phase only?
>>
>>43301727
>Will you be posting guides and whatnot, especially for things like interjections (hm, huh, wow, etc.)
Not even I know how to get those out of the model, that will require further research
>And is the 30s audio output limit for the testing phase only?
Yeah, plus a rate limit so everyone gets to test it. When running it by yourself you can do whatever you want.
>>
>>43301515

https://files.catbox.moe/chaxe4.wav
>>
>>43301745
kek
>>
>>43300755
Heyo, any chance for option to automatically convert the wav to MP3 files when generating them (I'm just but lazy to converting every single file by hand) ?
>>
File: 1663805093515.png (152 KB, 868x920)
152 KB PNG
>>43300755
https://u.pone.rs/hyboomqt.wav
Alright, I've been pretty jaded towards voice AI as of late, but I gotta admit this is pretty impressive.
>>
>>43302597
Yup. It'll still require some audio splicing between multiple generations in some cases for the best possible take, but the quality floor has a notable bump to it compared to past TTS software I've dabbled in.
>>
>>43300755
MOSS-TTS-1.7B-PNY v0.1: a finetune of MOSS-TTS + my custom vocoder for 48KHz audio

HuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNY
Colab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBB
Public demo: http://198.53.64.194:35029/

See pic related on how to run on Google Colaboratory.
For local setup on your own hardware, you want at least 13GB of VRAM. Model runs ~1.5x realtime on a single RTX 5090 with the optimized runner. Download from HF and ask Claude Code to set it up for you.

>>43301171
From a technical perspective, this consists of two models: 1. A finetune of MOSS-TTS with fixed speaker conditioning, and 2. A very custom iSTFTNet2 vocoder that turns hidden states of the MOSS Audio tokenizer into 48KHz audio (which can be also repurposed for singing voice conversion).
>>43301515
The TF2 speakers are a bit lower quality because they were thrown in as an afterthought. Next version will include better emotion control and quality.
>>
>>43303129
Neat, thanks anon, gonna check it out
>>
>>43303129
It works! Very expressive, though it underpronounces some words.
https://litter.catbox.moe/inza6f99g9ioti7z.wav
https://litter.catbox.moe/epx5r2ssfynizrzv.wav
I think it would benefit from running the lm quantized with ggml. It's current ~12 gb vram footprint makes it impractical for most tasks, and it's pretty slow too.
>>
>>43303129
>Delta
>Google Colab
This feels like one hell of a throwback to the very early days of PPP. Welcome back and godspeed.
>>
File: anf shake tail.gif (198 KB, 560x526)
198 KB GIF
>>43303320
Known issue probably due to not much data, try lowering the temperature for now, something like 0.6. Temperature basically controls how "creative" the model is. Higher temperature is more chaotic.
>It's current ~12 gb vram footprint makes it impractical for most tasks, and it's pretty slow too
Very unoptimized, a 1.7B transformer should run much faster. I should figure something out for that soon
>>43303339
Thank you. I'm only getting started. Need to improve this and figure out singing voice conversion and LLM finetunes in 2024 I finetuned Llama on a desuarchive dump of /mlp/ and found out I could ERP with it in greentext
>>
File: 1511063309973.png (344 KB, 685x1024)
344 KB PNG
>>43303129
Cool, also Celestia is best pony.
>>
>>43303320
>~12 gb vram footprint makes it impractical for most tasks
Sadly this, the reason why rvc exploded as ai voice tool was that it was poorfag friendly as it could run in 4gb vram and train any voice on a 8gb vram card.
>>
>>43303320
>litterbox
Damn
>>
>>43303129
noice
>>
Bump.
>>
>>43303129
I was watching family guy reruns on cytube the other day, and I thought what the hell, I'll test the TTS using the show's dialogue. This tool actually rocks. The audio quality is INSANELY crisp and practically artifact-free (at least for the M6 and Trixie - haven't tested anybody else yet) I think once you release better emotion control, it will truly be heads-and-shoulders above 15.ai, because as it stands, I'm not sure if there's a way to force a certain emotion into the line's delivery unless the dialogue itself clearly betrays that particular emotion.

https://files.catbox.moe/u3hjsj.wav
https://files.catbox.moe/ob9myx.wav
https://files.catbox.moe/yurj0p.wav
https://files.catbox.moe/o5olvb.wav
https://files.catbox.moe/bogptq.wav
https://files.catbox.moe/75ttqf.wav
https://files.catbox.moe/fx1zry.wav
https://files.catbox.moe/ifr120.wav
https://files.catbox.moe/isd7iz.wav
https://files.catbox.moe/k9espp.wav
https://files.catbox.moe/vsw445.wav
https://files.catbox.moe/l1ahoj.wav

Godspeed, fren. Looking forward to those updates.
>>
>>43303475
I played around with optimization, q4 quantization doesn't speed it up, the bottleneck is n_vq_for_inference. Halving it (16 instead of 32) produces audio in half the time but it's not as good https://litter.catbox.moe/uqzjohq2d3i03o9k.wav but I assume you already know all this. I wonder if 32 codebook channels is overkill for this task since it only has a limited set of voices to represent.
>>
>>43301515
Good luck, anon - I made the Heavy say
>We are going to destroy Israel AND generate ponies!
>>
>>43305634
Emotion control is in the plans. I'll stick with Cookie's BERT conditioning technique. Audio still has artifacts on the TF2 speakers as those represent only a minority in its training set.
>>43305746
I've got it up to ~3x realtime with some cudagraph and torch compile witchcraft
>I wonder if 32 codebook channels is overkill for this task since it only has a limited set of voices to represent
That I was wondering too. It would probably require training a vocoder specifically for that.
>>43303611
Is RVC still the go-to for singing voice conversion? I can swap out the vocoder in it for mine to improve its audio quality
>>
>>43306298
>Is RVC still the go-to for singing voice conversion?
Yep.
>>
>>43303129
These models sound amazing! What are the limits? Are you able to train a derpy model or is there not enough data?
>>
>>43306298
I'll be waiting for updates
>>
>>43306298
>Emotion control is in the plans.
noice
>>
>fucking around interdimentionally
>https://files.catbox.moe/bbmjgz.mp3
A old audio conversion/remix of the same name panel from last year (or two years ago?). Reposting to see other anons would enjoy this meditation thing .
>>
Did Clipper's 1st master file MEGA get nuked? Whenever I try to access it, it just loads indefinitely.
>>
>>43307177
Seems all there to me.
>>
>>43305634
kek these lines
>>
>>43307177
Nvm its working again
>>
>>43294893
>>
snowpity
>>
File: time to be alive.png (53 KB, 600x280)
53 KB PNG
Im glad all of you Anons are all well.
>>
>>43307109
comfy audio
>>
>>43309365
Sunbeam
>>
>>43309693
thanks, I had an early version were I converted the og audio with just rvc but the audio outcome was very meh, lucky rvc-sovits came out some time around it so I could gen Luna talking in a proper calm yet reassuring emotional style.
>>
>>43310789
ai mare
>>
>>43307079
>>43305634
Working on emotion control. More complicated than I expected because this big ass transformer model is less responsive to conditioning than models of old. Regardless, I'm toying around with an emotion category + an adjustable energy value (from 0 to 100).
For example, Twilight, same sentence, same Happy emotion:
95% energy: https://files.catbox.moe/smt36d.wav
15% energy: https://files.catbox.moe/nl0opy.wav

>>43306329
Not enough data for Derpy, unless we start using fan VAs (and if they have enough clean data for that)
>>
File: 1370206539512.png (31 KB, 181x248)
31 KB PNG
>>43311332
Dang, nice work anon! Thanks for the swift turnaround on the feedback! If you're able to implement this successfully, it would help out a ton for the projects I have planned. Keep it up!
>>
>>43311332
We are so back
>>
>>43311332
Very nice, but is there any reason why she sounds more enthusiastic on the 15% than the 95% one?
>>
>>43312305
Woops, I flipped the labels. They are in the wrong order; 95% is the 15% and 15% is the 95%.
On another angle, I found some NSFW voice pack: https://opennsfw.carrd.co/#vo2 ; I'm getting it transcribed so I can throw this into the training set of the next model. The gist of it is that it will teach the model this stuff and transfer capabilities to the pony speakers too.
>>
>>43311332
>emotions controlled by percentage
fuck yes, this is exactly what I wanted since forever. Please tell me, you are having plans in adding option to mix two emotions together as well?
>>
>>43311332
anon... i litterally cannot think of a way it gets better than this
genuinly what else is there beyond this? it's functionally everything we could need, it's 1:1 like the voices i hear nothing robotic, it has the emotion, and all the little things that make it the pony they are
the one time i keep up with /ppp/ and it's already at the end game
>>
>>43312608
>Woops, I flipped the labels.
Ah yeah I figured.
Good work though! She sounds really happy in the higher energy sample.
>>
>>43312608
>horny 100%
aw yiss
>>
>>43303129
>>43313194
awesome
>>
Up.
>>
>>43313228
Now, this is all I could had ask for and more
>>
>>43303129
if hydrus is still active in this thread, can we get this on haysay?
>>
>>43312608
mister delta please please oh please tell me this emotions update will be coming soon all of dash's lines are screaming and i can't get her to settle the fuck down for emotional lines
>>
>>43311332
amazing work, can't wait to see more updates!
so happy that the ppp is back, it motivates me to gen again.
https://files.catbox.moe/8g084q.wav
>>
hi ponys whats up guess whos back
>>
>>43314144
poopikins thought u were gone
>>
>>43312766
First version will be just select from a set of emotions + energy slider, as it works decently.
>>43312779
There's still a lot of work to be done, like improving voice conversion and LLMs
>>43314139
Yes. Technically everything's in place but I'm trying to bundle NSFW too, which is proving a bit complicated. If I can't figure it out by the end of week, I'll just release as it is.
>>
>>43314222
i love you thank you nonny this is fucking awesome
>>
>>43314144
sunbeam
>>
>>43314222
As someone in the minority who preferred using reference audio, do you plan to add an option for that? For some things, it's faster for me to just speak the lines and get it right once compared to the ritual of spamming generate for half an hour with 15. This certainly outputs faster, but I prefer to fine tune with different takes instead of adjusting emotions with sliders, especially because things like stammering and fluctuating tone of voice can't be changed easily with TTS. But so-vits and RVC are certainly showing their age compared to this.
>>
>>43314207
Don't call it a comeback.
>>
>>43314222
>I'm trying to bundle NSFW too
like
fucking lewd noises and a "horny" meter or some shit for the pony voices?
oh boy
>>
>>43312608
>On another angle, I found some NSFW voice pack: https://opennsfw.carrd.co/#vo2 ; I'm getting it transcribed so I can throw this into the training set of the next model. The gist of it is that it will teach the model this stuff and transfer capabilities to the pony speakers too.
BASED BASED BASED BASED BASED
>>
>>43315015
>>43315021
i'm so fucking excited to put it in my vn mod you anons have no idea
>>
>>43314851
i will call it a comeback if i want to, the reason why i left this forumn was because of idiots like you, and you know ultimately happened the thread died ?, im here to at least provide content and keep this thing alive so no go fuck yourself with a cucumber degenerate fuck face no one fucking likes you
>>
>>43314222
>>43315015
>>43315021
NSFW kinda works but is leading me down a research black hole at da moment(current iteration is extremely unstable) so I am dropping it for now in favor of focusing on emotion control.
Like, these are the only decent-ish samples I could squeeze out of it: https://files.catbox.moe/ehxmzw.wav ; https://files.catbox.moe/ebwxit.wav
>>43314572
Oh yeah, I plan on getting voice conversion upgraded soon.
>>
>>43315106
>dat dashie audio
>cum for me anon
UUUUUUUUUUUUUNFFFF!!!!!!!!!!!!!!!
>dropping it
Fuck my nigger existence
>>
>>43315106
I wouldn't mind some beta unstable testing version being on the page for the time being, but also if that's going to be a giant fucking hassle to have then don't bother yeah. Cool work though.
>>
>>43315158
Also, adding onto this, I think NSFW is less of a priority than enabling further general control anyways. Being able to, yeah, control emotions and general sentiment etc is great, but being able to control intonation by placing some sort of emphasis on certain words etc would do a ton for making there be less gacha in getting what you want.

That being said, NSFW is still peak and something I'd love to see at some point. Some kind of ASMR toggle to go with it to make it sound like it's being whispered in your ear would likely do a lot for some anons too kek.
>>
>>43315106
I hope you will make it a docker image with all the optimizations and shit
>>
>>43315106
Btw, what's the current vram and cuda requirements to run this?
>>
mare bump
>>
>>43316060
Yeah I'm definitely excited to see what's cooking with the new TTS. It's worth keeping the thread boomped for it. The /chag/ usage of it was so damn cool already.
>>
File: tf2 1750730551364987.png (265 KB, 860x1000)
265 KB PNG
>https://u.pone.rs/gykynemn.mp3
>>
>>43317647
and the inside of a horse
>>
>>43316730
>by end of week
just two days away just two days away its like christmas fucking morning
>>
>>43318088 (checked)
Man I completely spaced that. God I'm so excited. I've still been using the demo like daily. I'm holding onto some lines I generated for potential pony shitpost projects and I'm really excited about it. This stuff is an OC inspiration goldmine.
>>
>>43317647
Truth
>>
>https://www.youtube.com/watch?v=45-GDaNgfM4
new BGM kino its mostly an instrumental but there are still some ai voice bits in there so it counts
>>
>>43318783
>pony zone was 5 years ago
i can feel my bones crumbling to dust
>>
>>43318794
so do i bro, so do i
>>
>>43315106
Oh, forgot to ask: is there a way to remove the 30 second output limit? It's still present, even when ran locally.
>>
Pump seed inside mares
>>
File: 3458815 (1).gif (2.51 MB, 1440x1080)
2.51 MB GIF
These big transformer models are pretty hard to condition on emotion; mine was ignoring the labels so I had to devise an output head which runs emotion classification, so that forces the model to actually pay attention to the emotion labels. It works but is a 'lil bit more subtle than in other models. For example, Rainbow Dash:

Angry, 95% energy: https://u.pone.rs/evxtaazv.wav
Neutral, 20% energy: https://u.pone.rs/oquscjwf.wav

But it should get stronger with more training, I'll release this in the next 2-4 days So much technical stuff going on, maybe I should see if I can do a university-style lecture at a future Mare Fair.
Pic unrelated, Lucky Roll makes me hard
>>43319862
Whoopsie. There's a config knob in the code somewhere, it should be easy to find or just ask your favorite coding agent to do so--I'm fully sloppilled and don't read nor write code manually anymore. If you don't have any subscriptions, OpenCode offers free usage of good enough models.
>>43315474
13GB VRAM, any CUDA that can run modern pytorch will do
>>43315122
NSFW will be in some future version.
>>
>>43320423
>13GB VRAM,
Fug
>any CUDA that can run modern pytorch will do
Double fug
>>
>>43320423
Oh shit, was just making a post about how excited I was for the next version. Fucking awesome. I'm really looking forward to the chance to have Ponk actually speak softly in some lines. I was putting in cutesy comfy waifu stuff for her to say, but she would always yell it, haha. Twilight's been the most consistently tonally good voice for me so far from my experimenting.
>>
>>43320427
Yeah I wish it were more compact so that I could load it with a language model
>>
>>43320468
yeah, would be nice if there were voice and text models that weighted under 1GB to make it possible to have a "discussion" with multiple characters at the same time that can't actually read each other minds like the current llms do.
>>
>>43318794
>pony zone was 5 years ago
Good God, how did that happen?
>>
>>43314144
what I've made so far over the last few days, all with MOSS.

Awkward Dash and Trix:
https://files.catbox.moe/8g084q.wav
https://youtu.be/9NO_LqQfNpA?t=10

Smart Dash:
https://files.catbox.moe/p3l2om.wav
https://www.youtube.com/watch?v=SBiXajenKrg

Why would anon do that:
https://files.catbox.moe/nbalhb.wav
https://www.youtube.com/watch?v=m3MaTuv6QHI
>>
>>43322236
kek, nice work.
woundn't mind Trixie sucking on my nose ifkwim
>>
>>43322236
Man, this TTS really is something else. Bravo, anon! Have you ever considered making some audiobooks for a short fimfiction story, perchance? I feel like that would be a great use-case; the TTS is just that good.
>>
>>43322792
>Have you ever considered making some audiobooks for a short fimfiction story,

Maybe, I've thought about it before. I'd rather see if I could animate a short story rather than creating an audiobook, though animation takes forever. There's a few fics that come to mind that I'll like to adapt someday (with the authors permission, if they're still contactable)
>>
>>43322830
> I'd rather see if I could animate a short story rather than creating an audiobook, though animation takes forever.
For your own good, anon, I would advise against that heavily. Start small. Make an animated skit inspired by something under a minute or so, like one of those youtube shorts. The feeling of achieving those small goals consistently will give you the motivation to do something longer. And by animation, I hope you don't mean hand-drawn or flash animation, lol; even just PNGs sliding across the screen like the tax breaks animation would suffice. Don't burn yourself out
>>
Why does dsv4 have such niggerishly slow prompt processing?
>>
>>43323079
wrong board sorry
>>
Not pony related but I just wanna say I am big fan of vibecoding small scripts that I could had write in a day or two but I can get LMM make them in under a minute
>>
File: 3817484.jpg (71 KB, 1090x550)
71 KB JPG
>>43303129
>>43320423
Updated the model with first iteration of emotion control. Links remain the same:

HuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNY
Colab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBB
Public demo: http://198.53.64.194:35029/

You now have 12 emotion classes to choose from, plus an energy slider. They do influence, but it's more of a nudge than a demand. You still have to craft your prompts. Regardless, I hope this makes it easier to get what you're looking for. Nonverbal is an emotion class reserved for NSFW mode sometime in the future.
Pic unrelated. Also, rate limits on the public demo have been doubled as I've now got an optimized runner that does single batch inference at 3.5x realtime.
>>
>>43323889
excite
>>
>>43323889
Thanks for yet another release! Quick question: how do I enable the optimized runner? Just paste the code from HF and then run the gradio? Or do I have to stick to powershell?
>>
>>43323889
i was wondering if you guys could also try adding more ponys such as the student six and the cmc's
>>
oh and thorax
>>
>>43324263
The code from the HF repo and Gradio already has all the optimizations. I think some are turned off by default, because this speed is achieved by using TorchInductor and compile witchcraft to turn the whole inference flow into one big kernel to reduce overhead, but takes 5-10 minutes to startup, which is no problem for a long-running server.
Also, expect the demo to go from a Gradio app to something more refined UI-wise. Maybe I'll give it a real name. Taking suggestions
>>
>>43323889
Oh fucking awesome. I was just using the demo in bed and then I woke up to see the new settings and checked the thread then. Was so damn hype. Only responding now but yeah this is sick.

As a Ponkfag I am especially pleased, as beforehand she would yell 99% of her lines whereas Twilight was fantastically pitched most of the time. Putting "calm" on "0% energy" has led me with a lot of softer lower pitch Pinkie speech compared to before, which I love a ton. Her voice is so cute when it's chill. Extra rate limit is also really appreciated as I love playing with this.
>>
>>43324798
>Putting "calm" on "0% energy" has led me with a lot of softer lower pitch Pinkie
>https://u.pone.rs/krthqfzm.wav
holy fug, yeah this works great, finally a tts Ponk that doesn't talk like she just chug a entire barrel of energy drink.
>>
File: file.png (620 KB, 2516x498)
620 KB PNG
>>43323889
I tried to use the collab and I got this error
>>
>>43324798
>softer lower pitch Pinkie speech
I may be a ponkfag now.
>>
File: 1363919329737.png (196 KB, 830x962)
196 KB PNG
>>43324899
Yeah exactly. She sounds so fuckin' cute dude. Pinkie's always been the roughest one in TTS in my experience, but this is finally starting to deliver something really nice.

I'm glad you found that combo useful too. I've been experimenting with the style text shit and have gotten some interesting results. If I come across good sentences that have really nice results for prompting I'll try and share them here.

Here, before posting this I agonizingly fucked around and managed to get a couple cute ones.


>ponk on the inside
https://u.pone.rs/lubytuhc.wav

>ponk gf (despite her being a pony) (this one required genning it piecemeal and stitching together because it was kind of long, but I hope you guys like it, it took me a minute but I think it's really cute)
https://u.pone.rs/sbdyftft.wav

(also >>43324959 extremely based, hope you like the audio clips I made for this post, you're my bro now if that's true)
>>
>>43323669
Examples?
>>
I'm going to see if OAI Codex can port this model to C++/GGML if I just leave it on a loop.
>>43324950
Fixed. The new optimized path switched to TorchScript instead of ONNX for the vocoder and the Colab demo didn't download that artifact. Also, since Gradio is being weird with share links, the demo uses a Cloudflare tunnel.
>>43325014
Cool stuff anon. If you find interesting ways of using the model, do share them.
>>
>>43325014
non-screechy ponka is best ponka. both are fine, of course, but I like it better when she's less histrionic. extremely cute gens btw
>>
>>43325312
>port this model to C++/GGML
Yes please. Setting this up on windows is a pain and I get random crashes on wsl.
>>
mares
>>
>>43326118
I love them.
>>
>>43325312
>>43325582
Turns out it could. With CUDA backend, practical VRAM requirements with GGML become 10GB with everything in float16, or 8.2GB with the model quantized to 8bit, 7.8 with 6-bit. Working on Vulkan backend which is the most vendor-agnostic then I'll release it.
>>
>>43326545
>7.8 with 6-bit.
Noice! Finally I can have a web-free run tts for whatever project I want without worrying about the net connection randomly going to shit for no reason
>Working on Vulkan backend
not an AMDfag but I would imagine it would be nice to have this for the linux people
>>
hooray for the new era of ai ponies!
>>
https://u.pone.rs/wedifxxb.wav

I'm so sorry but I was cackling like a dipshit edgy 12 year old boy at this one.
>>
>>43326855
kek
>>
File: file.png (324 KB, 1111x503)
324 KB PNG
>>43325312
The collab no longer crashes but now it is stuck on this without spitting out a url
>>
>>43327060
Fixed, refresh and try again. There was some funky shit that meant the polling was stuck.
>>43326588
>>43326545
Works on Windows. This is with the Vulkan backend for the TTS and ONNX/DirectML for the vocoder on 1x RTX 3080 Ti.
https://u.pone.rs/dvhptahc.mp4
Relatively slow because Windows overhead and I have a bunch of junk open, but on a clean Linux and solid GPU it's 2.5x realtime.
>>
>>43327086
Looking forward to using it.
>>
>>43326545
>>
>>43327086
progressbeam
>>
>>43327336
yay
>>
Haysay is dead and I decided to run RVC itself locally, but I can't find the Maud Pie RVC model, any leads lead to dead links, and it seems more by the same person who made that model are also dead.
>>
>>43328369
Hay Say is back up now. Sorry for the unexpected downtime.
Yeah, looks like the original Maud Pie model is a dead link now. I went ahead and re-uploaded it here, in case you still want to run it locally:
https://huggingface.co/hydrusbeta/hay_say_reuploaded_models/tree/main/rvc/Maud%20Pie

>>43313914
Yes, I'll look into it this weekend.

>>43323889
Thank you for doing the hard work of bringing us ponies on MOSS-TTS with emotion control. Good stuff so far.
>>
>>43328468
>Hay Say is back up now. Sorry for the unexpected downtime.
Thanks Anon!
>>
>>43328468
Its down again and thanks. Also, pretty much everything from KenDoStudio is gone.
>>
>>43328896
Still down for me as well.
>>
>>43328896
>KenDoStudio is gone.
damn
>>
>>43323889
I would love a api version of this that can take in longer text I have so many cool ideas I can do
>>
>>43325014
I don't pay much attention to this thread now (not because idc, I'm just too busy with uni), but I remember that australian guy running his voice through vits and thought she sounded adorable as an australian.

>>43324899
>finally a tts Ponk that doesn't talk like she just chug a entire barrel of energy drink
Nooo!
>>
>>43327086
Sorry to bother, but any ETA on a reference audio function?
>>
>>43327086
I'd like to ask how do you train the voices for MOSS? is there a guide somewhere. What are the specs to do train a character.
>>
>>43327086
hey nonny is there any potential for a phenomizer or something of the like to be trained for this model because it gets complicated/uncommon words very wrong very often. it doesn't understand how to pronounce things that aren't in its training data which is all mlp dialogue. these are all just assumptions about the infrastructure but the problem remains the same; and it's incredibly jarring to hear a word like 'erosion' be pronounced 'seyfloerser' like a slurred drunk mare
>>
>>43328896
>Still no HaySay
Is it time to get worried?
>>
>>43328369
i might have it :)
>>
hydrus beta if your out there whats happening with haysay.ai website bud ?
>>
>>43330547
>>43330700
Hay Say went down a couple more times, unexpectedly. I have an idea as to what's causing it but I'm not 100% certain yet. I've added some additional metrics monitoring and I'll be keeping an eye on the server in case it happens again. The server is back up right now.
>>
>>43330247
nvm i fixed it by asking fable to train a LoRA on a dataset that included the voices with less samples as often as the ones with more of them.
before/after a/b test wav:
https://u.pone.rs/tnonbtdv.mp3
lora is included in the:
/chag/ AI VN tts mod and can be used independently of it with some finaggling
https://u.pone.rs/sfwhsrkz.zip
>>
The CIA got Delta it seems. The model was too good.
>>
>>43331172
Nay, I'm just very focused on the Windows scheibe. Basically, the demo works, but with a bunch of quirks that make it not 12GB VRAM-safe if you do something like want output longer than 7 seconds, so have to implement and test streaming.
and was blendering something for second life, but that turned out harder than I thought
>>43330191
Like, RVC-style? Probably 2 weeks
>>43330247
>>43331049
Ah, dataset balanced sampling. Well done anon, I gotta do that.
>>43329611
Yeah. I still haven't tested the model's long context abilities.
>>
>>43331314
>Probably 2 weeks
noice, danke Delta-kun
>>
>>43331314
kino
>>
>>43330883
also you should try to add this tool

MOSS-TTS-1.7B-PNY v0.1: a finetune of MOSS-TTS + my custom vocoder for 48KHz audio

HuggingFace: https://huggingface.co/ZDisket/MOSS-TTS-PNY
Colab Notebook: https://colab.research.google.com/drive/1tDIYCMumcW5w3JWnQ0tBGyAr-ZpaaXBB
Public demo: http://198.53.64.194:35029/

See pic related on how to run on Google Colaboratory.
For local setup on your own hardware, you want at least 13GB of VRAM. Model runs ~1.5x realtime on a single RTX 5090 with the optimized runner. Download from HF and ask Claude Code to set it up for you.

>>43301171
From a technical perspective, this consists of two models: 1. A finetune of MOSS-TTS with fixed speaker conditioning, and 2. A very custom iSTFTNet2 vocoder that turns hidden states of the MOSS Audio tokenizer into 48KHz audio (which can be also repurposed for singing voice conversion).
>>43301515
The TF2 speakers are a bit lower quality because they were thrown in as an afterthought. Next version will include better emotion control and quality.
>>
>>43332450
What the hell is wrong with you?
>>
>>43332477
>he just copy pastes the whole fucking thing
lel incredible retardation
>>
>>43331314
Yeah, looking forward to playing with it
>>
snowpity
>>
>>43332477
>>43332865
retards are talking about themselves again
>>
File: 9374659585773829.jpg (68 KB, 682x682)
68 KB JPG
>>
Windows MOSS-TTS runner. Currently a terminal interface (type in anything to do TTS, use / for commands)
System requirements: At least 12GB VRAM GPU (any brand), modern Windows 10/11, make sure to update your GPU drivers.
Runner: https://u.pone.rs/rlsymqeb.zip
Instructions: Run download-models.bat. Once that's done, run run-vulkan-directml-quant.bat for 12GB GPU, run-vulkan-directml-full.bat for >16GB GPUs. It will open a console.
Source code: https://u.pone.rs/yzragcxy.zip
This uses GGML for inference with transformers (Vulkan backend which is vendor-agnostic) and ONNX with the DirectML backend for the vocoder.
You can find the quantized models themselves under https://huggingface.co/ZDisket/MOSS-TTS-PNY-GGUF ; the .bat downloads from there.

This thing is a research preview for playing around with.
>>
File: twilight embarazada.png (1011 KB, 1098x904)
1011 KB PNG
The MOSS-TTS model seems to automatically understand other languages to an extent, even if you don't change the "language" field. Twilight will speak Spanish with an American English accent. Mispronunciations abound, however, and it takes some finagling.
https://files.catbox.moe/yp2wdl.mp3
>>
>>43334990
>run run-vulkan-directml-quant.bat for 12GB GPU, run-vulkan-directml-full.bat for >16GB GPUs
So which option do I chose for 7.8GB vram? Or is that not yet implemented?
>>
>>43335172
There is none (at least not without quantizing both models to hell), it turns out my earlier 7.8GB VRAM figure was based on only the models loaded without any inference cache. Sorry.
The MOSS team has a nano model, so that could be next after RVC
https://github.com/OpenMOSS/MOSS-TTS-Nano
>>
>>43335303
what gpu are you using to train these models
>>
>>43335314
AMD Instinct MI300X.
>>
>Drop in for my once-in-a-blue-moon check-in at slash em el pee slash
>Check this thread out to see any interesting developments
>Nothing new in OP, scroll a bit
>Delta

https://vocaroo.com/1bvGHJAmkbx2
transcript:
Goodness grace! This is the smoothest I've ever heard of a text to speech! I might actually follow along for this one, given I have a GPU (gee pee you) or two I could enslave for this... Oh, I do hope someone manages to get pure phonetic alphabet input working on these, as those would honestly make control way easier. Þorn and Eð are my favourite letters...
https://vocaroo.com/1jTTlFQudWAy
Another test for fun
As for what hardware I have... Quite the mix:
RX 7900XTX
RTX 3060
Arc Pro B50
I'll try to follow along this thread, but I oughta get to sleep. I'll try experimenting tomorrow with the 7900XTX - the 3060 is currently an SDXL slave
>>
mares
>>
File: 912489237813017851.png (177 KB, 1029x732)
177 KB PNG
>>43335303
Not sure if that helps, but afaik when they quantize transformer models they usually leave attention layers at high quants (like q8) and quant the mlp layers more aggressively, say, q4. MLP makes up most of the parameters while being more resistant to quantization, at least in natural language. But I'm not sure whether it would transfer to this task because sequences are much shorter.
>>
>>43335711
and yes, picrel shows the ggufs from your repo, I indicated what I would try changing. Looking forward to and preserving my cum for the release.
>>
snowpity
>>
>>43334990
uhhh, how do I get it work on linux?
>>
>>43336692
Not officially supporting Linux because I can't be bothered with supporting every distro from Ubuntu to RaritysSmellyFlankOS. However, the source code does compile and works in Linux fine, so you can figure it out with AI.
>>43335711
I should check that out
>>
>>43337124
>Not officially supporting Linux because I can't be bothered with supporting every distro from Ubuntu to RaritysSmellyFlankOS.
Koboldcpp the portable-ish Linux executable has worked fine on both my Fedora server and my Tumbleweed main machine, and likely works on Arch and Debian. It mostly depends on how it's packaged and what libraries it uses,
A vast majority of distros people commonly use have a few base distros as "upstream": Ubuntu, Mint and Pop!OS from Debian, Nobara and Bazzite from Fedora, Manjaro, Endeavour and CachyOS from Arch.
>TL;DR By building for Debian, Arch and Fedora you cover most of the Linux ecosystem.
Still, for wanting people to just compile for their own stuff and not worry, outline _a_ path to make it work on Linux and I'm sure most can follow along with the necessary changes.
Will try compiling the thing. If successful on my Tumbleweed, I'll try to document it and if it does become a single packed executable in the end I'll probably upload it.
>>
>>43337124
>Not officially supporting Linux because I can't be bothered with supporting every distro from Ubuntu to RaritysSmellyFlankOS
Why not make an AppImage?
>>
>>43337245
Soooo
Here's how it went... As Delta said I figured it out with AI. Just plain Claude Sonnet 5 medium effort no reasoning. And Claude also wrote this thingy:
>Got MOSS TTS building and running on Linux with Vulkan on an AMD 7900XTX, no CUDA/DirectML/Windows.
>Deps: vulkan-devel, shaderc+shaderc-devel (glslc is in the non-devel pkg on Tumbleweed, headers in -devel, annoying split), spirv-headers (not pulled in automatically, cmake will just fail on ggml-vulkan's CMakeLists until you install it separately). ONNX Runtime 1.26.0 CPU tarball, unused in the end — go GGUF vocoder route instead, explained below.
>cmake -DMOSS_TTS_ENABLE_VULKAN=ON -DMOSS_TTS_ENABLE_CUDA=OFF, build moss-tts-tui + moss-tts-engine targets specifically (there's a moss-vocoder-onnx target too but it's gated behind a flag that defaults off, don't bother, TUI doesn't need it, vocoding's in-process).
>Two real gotchas if you've got multiple GPUs:

>It picks Vulkan device index 0 by default. If you've got an iGPU or a second card ahead of your main one in enumeration order, it'll try to allocate on that instead and OOM on something that should be trivial. GGML_VK_VISIBLE_DEVICES=N env var fixes it.
>The .onnx vocoder file only works with --provider cpu/cuda/directml. If you want Vulkan, you need the separate .gguf vocoder file and --provider ggml-vulkan. Different file, different provider family, README doesn't make this obvious.

>Once on the GGUF vocoder + ggml-vulkan: hard crash, GGML_ASSERT(src0->type == GGML_TYPE_F32) failed inside ggml_vk_build_graph. F16 vocoder graph hits a Vulkan kernel that only supports F32 input somewhere. Not a build problem, not a driver problem — Vulkan backend genuinely can't run this op combo in F16 right now.
Fix: --provider ggml-cpu for the vocoder instead of ggml-vulkan, keep everything else on Vulkan. Main model + decoder4 (the actually expensive parts) still run on GPU, vocoder falls back to CPU. Works fine, vocoder's cheap enough that CPU fallback doesn't matter.
Back to me, the human:
>TL;DR yeah you can get it working on a setup similar to mine I guess
The Vulkan part was because my main machine has an Arc A310 handling desktop and displays plugged into the first slot. That way, it handles all the browsers and whatnot while the big boy card does everything else.
Dunno what else is particularly iffy about this report or something
Gonna make zip containing the cmake commands Claude made and whatnot
>>
>>43337245
>>43337251
I'll support Linux as in give a "build it yourself" recipe.
>>43337398
Sounds about right. I forgot to say the vocoder is partially implemented in GGML but it's better to use ONNX, the GGML port is incomplete because a ton of operations have to be implemented from scratch (libllama was designed for transformers, not 2D convs). Unfortunately for Linux there is no DirectML equivalent; gotta use CUDA for NVIDIA GPU and MIGraphX for AMD.
>>
>>43337398
https://u.pone.rs/klpogtoq.zip
Here's:
Some cmake commands
A script to run it
A readme
For building and running a Linux version
Hopefully others can iterate on this
>>
>43337439
>Sounds about right. I forgot to say the vocoder is partially implemented in GGML but it's better to use ONNX, the GGML port is incomplete because a ton of operations have to be implemented from scratch (libllama was designed for transformers, not 2D convs). Unfortunately for Linux there is no DirectML equivalent; gotta use CUDA for NVIDIA GPU and MIGraphX for AMD.

sucks given MIGraphX is supposedly deprecated and the new method is python heavy or... something idk
That means what I have in >>43337441 is... Eh, it is what it is.
Hopefully improvements come around, I just did what I could. Now time to figure out how to operate this from the tui
>>
>>43337441
oh
warning:
the build script is incomplete it doesn't pull the onnx runtime
mine and claudes mistake
>>
>>43337451
To fix that mistake, add -DMOSS_TTS_ENABLE_ONNX=ON into the cmake thingy and then in the run commands change the vocoder model to the .onnx one and provider to cpu
it is now very late and the sun is shining brightly here...
https://voca.ro/16NdL4gLBT3m
>>
>>43335368
>that battles atlas generation
fucking kek, I'm not the only one who sees a beautifully demented pro-AI and anti-human message in that song then. Love that track, absolutely did not expect to see it here.
>>
Uhoh, the docs haven't been updated in a hot minute. How would you like to list your new tts model Delta?
>>
>>43337665
Media Molecule did choose some absolute bangers to license...
If I knew how to use music tools and the bare singing necessary I'd be down to make a full cover with one of those rvc things or whatever BGM was using
Though I don't see any specific messaging
>>
File: happyooo.gif (630 KB, 300x354)
630 KB GIF
>>43331314
>2 weeks
Nice, exciting stuff. I can't wait to try it out with speech-to-speech conversion.
>>
>>43337974
You decide
>>43338923
Oh, hi BGM. Still using RVC right?
>>
>>43338953
For now, yes. Your new thing looks very promising.
>>
snowpity
>>
Any chance on including the training process and script? Specially if it includes the emotional control as having that on larger variety of voices would be pretty amazing.
>>
>>43339896
Yeah it would be stellar.
>>
>>43338073
Yeah, my first exposure was LittleBigPlanet back when too. I've since come to love the lyrics for how absolutely diabolical they've become since the rise of AI lmao. An AI pony cover sounds perfect.
>>
https://voca.ro/1kIzNC7MHlGg
REAL AUDIO OF ZECORA SPEAKING HER NATIVE LANGUAGE
(no seriously I have no idea how this happened)
>>
>>43338073
You reminded me of Passion Pit - Sleepyhead, kek. Would be a neat song to do a mare cover of, if not to just do the high-pitched parts!
>>
Oh boy, time to restart some of my older project with the new voice tools.
>>
>>43339625
Very insightful.
>>
>>43341926
mare
>>
Any Panel for /mlp/con?
>>
>>43343294
A little late to bring up. No, nothing has been discussed.
>>
>>43343521
perhaps with all the new development there will be some panel made for marecon
>>
Hello Delta,

While trying to run your Moss TTS model on CPU, I ran into a few obstacles I thought I should let you know about.

First, the default value for the option --decoder4-features-onnx is "ort_sessions/decoder4_features_fp32/model.onnx", but I don't see that model in your repo. Based on the code, I think this option is supposed to point at the vocoder onnx model, so I specified istftnet2_decoder4_50hz/istftnet2_decoder.onnx and that seemed to work. I recommend changing the default.

Second, I was still unable to get the oonx vocoder working because I got this error during the decoding step:
>File "...\moss_tts_torchopt_runner_bundle\portable_tts_runtime.py", line 624, in decode_outputs
>lengths_input_name = self.decoder4_session.get_inputs()[1].name
>IndexError: list index out of range
I'm not sure what's wrong here.

Lastly, lines 177 and 178 of portable_tts_runtime.py enforce the use of a cuda device when the option --decoder4-features-runtime is torch_fp16. However, this seems entirely unnecessary. In fact, I had to comment out those two lines in order to use the istftnet2_decoder_cpu.ts model, and it generated output successfully.
>>
File: discriminators2.png (1.72 MB, 1866x1048)
1.72 MB PNG
I've got my whole iSTFTNet3 thing (what powers the custom decoder that has the audio quality) into RVC and it's training (multi speaker model with all notable characters), after solving GAN instability this will take like a week.
>>43344077
Thank you. I'll get that looked at
>>43339896
Yeah, I plan to. Right now, the codebase is a vibe coded mess.
>>43343294
A bit too late for that, but some other con I want to present some things including this model. Basically, I have a custom vocoder that is highly upgraded iSTFTNet2 including ConvNeXt-V2 blocks in the generator, and a custom discriminator setup: MPD, MS-STFT and MS-CQT-D (from https://arxiv.org/abs/2311.14957). But unlike vanilla implementations, I found out log1p magnitude and instantaneous frequency are more numerically stable, and I have dedicated full-band (sees whole frequency range) and high-band modules. Blah blah.
>>
>>43344198
noice progress
>>
>>43344198
>I found out log1p magnitude and instantaneous frequency are more numerically stable, and I have dedicated full-band
i have no idea what that means but im happy it makes pony voices better
>>
snoof
>>
>>43341475
Whatchya working on Anon?
>>
>>43346012
just wished to get some greens turned into short audios, since talknet tts is dotting into neutral/deadpan note, while the rvc and other voice conversion models are aggressive not compatible with my voice.
>>
>>43346097
Ah, understandable. Text to speech no good?
>>
>>43346429
older models were good when doing few words or short sentences, but when trying to put something longer together I can hear the inconsistencies between each clip (speed, tone of voice, emotion, random word mispronunciations), which leads to generating 5 to 10 times more clips than is necessary just to sort out the good ones. With the Delta new model control makes it so much better at getting the voices follow the directions they are supposed to sound like.
>>
By the way Delta, could you add a toggle to disable the emotion control? I still have the initial release and I feel like the quality of the voicelines were more accurate-sounding (not by much, but just enough to notice).
>>
Hello Delta,

I don't think the "style text" input is used at all in MOSS-TTS-PNY. I'm guessing that's a mistake? Or maybe I am just missing something.

Line 713 of portable_tts_runtime would use style_text, but only if style_features_dim != 2
https://huggingface.co/ZDisket/MOSS-TTS-PNY/blob/main/moss_tts_torchopt_runner_bundle/portable_tts_runtime.py#L713
But... style_feature_dim is set to 2 in the config file and is never overridden in the code as far as I can tell:
https://huggingface.co/ZDisket/MOSS-TTS-PNY/blob/main/moss_tts_local_clipper_checkpoint/config.json#L106
By placing some debug points, I verified that neither line 559 (which loads the feature extractor) nor 713 (which runs the extractor) of portable_tts_runtime.py is ever executed.

By the way, is there another place you would prefer people to report issues and make feature requests, or is it fine for us to just post them in this thread?
>>
>>43346968
Hello Delta,

I just wanted to report another thing I found. It seems that "Audio top-k" is also never used. Line 726 of portable_tts_runtime.py always passes a value of `None` to the model's generate method. I think that should be `audio_top_k` instead:
https://huggingface.co/ZDisket/MOSS-TTS-PNY/blob/main/moss_tts_torchopt_runner_bundle/portable_tts_runtime.py#L726
Also, the CLI version is missing an --audio-top-k option in run_tts_torchopt.py - there are only options for audio-temperature and audio-top-p - and line 214 of run_tts_torchopt.py just passes `None` for its value to the `synthesize` method.
https://huggingface.co/ZDisket/MOSS-TTS-PNY/blob/main/moss_tts_torchopt_runner_bundle/run_tts_torchopt.py#L149
https://huggingface.co/ZDisket/MOSS-TTS-PNY/blob/main/moss_tts_torchopt_runner_bundle/run_tts_torchopt.py#L214
>>
>>43341014
Sounds like simmish.
>>
>>43346930
Emotion control cannot be turned off currently.
>>43346968
Style text is a relic from a previous attempt by me to add emotion control, it was Cookie et al. BERT hidden as conditioning input, which I later found out didn't influence the model
>By the way, is there another place you would prefer people to report issues and make feature requests, or is it fine for us to just post them in this thread?
This thread is fine.
>>43348252
Top-k sampling can be enabled but I keep it off by default because temperature and top-p are enough.
>>
mares
>>
snoofpity
>>
It's funny that he says this, like 15.Ai wasn't only usable for a grand total of one year out of the eight years the site has been up.
>>
>>43344198
sunbeam
>>
>>43350459
Yeah, the fun thing about the technology is that once something is made, somebody can recreate it, even if it take five years to do so.
>>
>>43349081
Ah, no biggie. I'll just alternate between the initial release and the latest as needed.

By the way, I hope you take your time polishing the next release, especially for bug-fixing. I didn't even know the style text function was just window dressing. I think even the list of requirements might need some updating, too, since I recall running into some issues with dependencies that I had to go and solve on my own (something to do with triton IIRC). Plus, there was an issue with the Speaker names only registering as Speaker (ID), so I had to manually find out which character had which ID in order to generate with the corresponding character. I also think the installation requirements bears updating, too: they make it seem as if the CUDA versions for PyTorch don't really matter, but I think I ran into some issues when installing versions above 13.0+ (forgot if it was related to the triton dependency issue or not), which were only fixed once I downgraded to 12.6. Perhaps I'm retarded with computers and fucked up along the way (very likely), but maybe you could consider doing a fresh install yourself to see if the install process is in proper order?

Just take your time to incorporate the feedback in this thread and take care not to rush things. You're doing great so far, but you're being a tad bit sloppy at times.
>>
>>43351025
It's been 5 years and nobody has.
>>
>>43351025
mare snoof
>>
Bumo.
>>
Site down atm for anyone else?
>>
>>43352448
Instance went offline. It was being hosted on a Vast RTX 5090 instance; if it continues being dead by tomorrow, I'll switch to something on RunPod.
>>43351241
Been researching the model itself for a long time, so as soon as I got it working, I got everything else ready as quickly as possible for a release. Next stuff should be more polished.
>>
>>43352630
>if it continues being dead by tomorrow, I'll switch to something on RunPod.
Thanks
>>
>>43351455
The speed of the board has decreased a lot as well.
>>
>>43351455
rip
>>
mares
>>
File: 40k 1780516651858476.jpg (545 KB, 2792x2724)
545 KB JPG
>>43353854
thats the plan
>>
>>43354305
robomares
>>
>>43354305
That's a stallion.
>>
>>43351455
Haysay is a thing. Though it's not quite the same.
>>
File: 260042.jpg (430 KB, 4000x1884)
430 KB JPG
Just finished a cover of ABBA with custom pony lyrics.
https://u.pone.rs/wdeknkzd.flac
>>
File: 4564464.gif (189 KB, 100x125)
189 KB GIF
>>43355660
sounds great anon!
keep up the good work
>>
File: 1701105.gif (452 KB, 297x221)
452 KB GIF
>>43355660
HOOOOLY SHIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIT
>>
File: derpy dance happy.gif (722 KB, 1080x1080)
722 KB GIF
>>43355660
hell yeah, this is the quality of stuff that keeps me coming back. thank you for that anon!
>>
>>43355660
kino!
>>
>>43355660
Lovely work anon!
>>
henlo frens. how 2 MOSS?
>>
File: 6405138.png (550 KB, 5000x5000)
550 KB PNG
https://voca.ro/15vaRks9fDfv
>>
File: 7413403.gif (1.46 MB, 498x291)
1.46 MB GIF
https://voca.ro/1hldXDn4d7ga
>>
>>43355660
ossum
>>
>>43356694
eat the Moss
>>
>>43355660
Well, cool things still happen every now and then.
>>
File: 7270131.gif (547 KB, 498x462)
547 KB GIF
>>43355660
>>
From what I can see in /chag/, the text models are more or less on the same level as good green texts, and the ai voice stuff is pretty much soled, the only other thing left is the pony robot bodies to make waifus truly real.
>>
>>43358904
I could never figure out how to make it work. If it doesn't require me to have a terrabyte of space to run them locally, it requires me to buy fucking Gemini or some stupid shit.
>>
File: yay.png (61 KB, 587x288)
61 KB PNG
Moss TTS has been added to Hay Say. It is, however, veeeery sloooow. Performance has always been an issue with Hay Say, but it's quite noticeable this time. I am looking into ways to make it faster. In the meantime, you can reduce the "RVQ Codebook layers" parameters to reduce the inference time, but at the cost of audio quality. Delta's Gradio demo is still up, too, and it's way faster.

For haysay.ai, I have implemented an output cap of ~20 seconds. On a local installation, you can increase that limit by editing the docker-compose.yaml file; search it for the --max-new-tokens option. Every 12.5 tokens adds about 1 second to the limit.
>>
>>43359159
Welp, that was short-lived. The server crashed 3 times within an hour and a half. Memory usage spiked just before the crash. I've removed Moss TTS from haysay.ai until I can figure something out. It is still available to run in Hay Say locally.
>>
>>43303129
hello delta i was wondering if your able to add more pony's to moss tts like the student six and such others add more voices it would be great to see :)
>>
mares
>>
>>43359159
>For haysay.ai, I have implemented an output cap of ~20 seconds.
You mean for Moss TTS, or does that cap apply to all models now?
>>
>>43359159
>Moss TTS has been added to Hay Say.
sunbeam
>>
>>43359932
Only to MossTTS.
>>
>>43357594
got sick :(
>>
>>43360993
wtf sick mare
>>
>>43352630
>>43355660
>>43359159
Thanks for all your work, Anons.
>>
File: qw3bhis.gif (59 KB, 86x129)
59 KB GIF
>>43355660
lovely
>>
mares
>>
>https://pony.tube/w/6ERFDbLpSCKZv4iWtGu5Zn
darn, wish people were posting poners or catbox mp3 links with their videos.
>>
>>43362835
What do you mean?
>>
>>43363292
>What do you mean?
Like, random people will post ai covers songs as video, but there isnt a link to just get the song. So a lot of time I will need to download the whole video and throw it into audacity to extract just the song, I dont want to sound like /mu/ autist, but most streaming sites will mess around with encoding/format of the creator video file so whatever I download will be downgraded in quality to some degree.
>>
>>43363368
Ah, understandable.
>>
>>43363368
the audio will still be compressed but you can download youtube vids straight to an audio file using ffmpeg or yt-dlp.
>>
>>43362404
Bouncy mare.
>>
Can we get a status update Delta?
>>
Uppo.
>>
>>
>>43352630
>>43364290
would likewise love to hear, if there's anything to share. I'm excited for V2V with this model.
>>
>>43364290
>>43365739
Struggling to get RVC working with my changes, it's a more delicate model than I thought. (Training collapses because GAN instability or fails to deliver quality)
>>
>>43365824
Don't stress yourself out over it too much! It was supposed to be a cheeky lil' request on our part; if you haven't found a viable path towards implementing it, I'd rather you just focus on developing the main TTS function. That's the main draw, anyway
>>
File: Amy's Calling!.png (150 KB, 689x689)
150 KB PNG
>>43365824
Damn, here's hoping you find the secret formula. The T2V sounds great, a similar V2V would be fantastic to have in the toolset.
>>
>>43355660
More than a decade later and the creativity never runs dry. I really love that from this fandom. Very good job and thanks for sharing this with us!
>>
>>43365840
Love me some Sisi.
>>
snowpity
>>
>https://u.pone.rs/ohzbuqxz.mp3
incredibly important message from purplesmart
>>
>>43367736
But they are tasty!
>>
File: smart 1754718043789810.png (3.31 MB, 2792x2959)
3.31 MB PNG
>>43367736



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.