[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109456130

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Wan
https://github.com/Wan-Video/Wan2.2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109457662
Good job
>>
>>109457662
Based moving collage.
>>
>>109457662
not LTX 2.3 in the OP ?
>>
Anybody tried under 20 steps yet? What's the verdict here?
>>
>>109457546
I hear nothing
>>
>>109457702
yes, the movements become stiffer as you go up
>>
Blud this spectrum shit is lowkey crazy
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Circa %30 speed-up. I have only tested a few samples but so far I did not notice any strong deviations or errors in them. I even tried t2v without any reference image and it still spits out close likeness.
I also don't know how it combines with other optimization but you must check it out at minimum. Really missing out if you don't.
>>
>>109457707
UNDER 20
>>
holyfuck did any of the anons download the recent krea2 code lyoko loras for yumi and welita? please share if you have them.
https://civarchive.com/models/2819269?modelVersionId=3183135
https://civarchive.com/models/2819269?modelVersionId=3180138
>>
>>109457725
did i stutter? extrapolate what i said
>>
>>109457705
shes saying gibberish
>>
Hail China for this gift

0 to 3s: the dog on the left which is wagging its tail says "get him, boys". 4 to 10s: as two heavily armed Israeli military soldiers with "IDF" on their chest armor break through the wall on the right side of the room, and grab the man with glasses, taking him out of the room.

https://files.catbox.moe/9n39qt.mp4
>>
File: 8kl8sy.gif (935 KB, 201x202)
935 KB GIF
generate some will smith videos for the next thread
>>
So what date should we hold the official /ldg/ funeral for LTX and Wan? Do all of you guys have your standby obituaries ready?
>>
did I add spectrum correctly?
>>
>>109457776
same day catjack kills herself. we would save so much on the funeral fees
>>
>>109457776
I am writing a heartfelt eulogy for Wan 2.2. Made me coom gallons.
For LTX though, I am only readying my bladder.
>>
Can someone explain to me how IDs work? The Minimax documentation is confusing with this example:

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>


Like, who is S1? the "young woman" or the first of the two children?
>>
>>109457785
I have no idea if it combines well with fp16 acc but otherwise seems so.
>>
>>109457785
>>109457801
Also I know you didn't ask these but comfy native int8 loading runs faster than the extension I believe, and heretic TEs are a meme diffusion.
>>
BFL must feel like shit right now. imagine releasing a cucked model right after H3. if they had any sense they'd just cancel flux 3 and spare themselves the humiliation
>>
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR LLM TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, eliminating the need for multiple models or API handoffs, thus reducing end-to-end latency. It sets new benchmarks by bringing open, robust, and highly natural conversation capabilities. Moreover, NVIDIA NemotronLabs VoiceChat is the first open full-duplex model to support tool calling while maintaining a natural conversation flow during tool execution. For each tool, a specific “on-hold” message can be defined that will be spoken by the agent as soon as the LLM generates the text that will trigger the tool call and response.

The model operates on audio signals, which are encoded using a fast conformer module. The resulting audio tokens are inputted into a Nemotron Nano V2 9B LLM backbone to predict text tokens, which are fed to a TTS decoder [2] to predict audio codes for generating the agent's speech. A separate output channel is used to predict tool calling scripts. NemotronLabs VoiceChat offers an unprecedented trade-off between intelligence and latency in the space of open-source voice agents, as highlighted by our benchmarking results below.
>>
So what optimization do you guys prefer, Easycache or Spectrum? What do you use and why?
>>
>>109457789
i think you are more likely than "catjack" to leave this world considering you spent your christmas eve, day, and holiday spamming 4chan; your lover left you alone at the top of mt fuji to go down with another guy; you've been unemployed for over two years and unsuccessfully begging for money everywhere (including your bull whom you aren't even allowed around!); and you spent 2 years spending your energy samefagging without picking up a single user
i'm sorry for you bro
>>
File: 669.gif (3.86 MB, 400x532)
3.86 MB GIF
>i think you are more likely than "catjack" to leave this world considering you spent your christmas eve, day, and holiday spamming 4chan; your lover left you alone at the top of mt fuji to go down with another guy; you've been unemployed for over two years and unsuccessfully begging for money everywhere (including your bull whom you aren't even allowed around!); and you spent 2 years spending your energy samefagging without picking up a single user
>i'm sorry for you bro
>>
holy meltdown lol
>>
>>109457844
EasyCache rapes it as always. Spectrum feels like magic though.
>>
Should I dl fp8 scaled or int8 convrot?
>>
Does the 19gb INT8 pruned model is enough for everything ? Got a feeling its too small to contain all the datas
>>
>>109457863
int8
>>
>>109457867
gm saar
>>
https://files.catbox.moe/inhcur.mp4
>>
>>109457867
the rest of the data is stored in your balls
>>
Honestly, Minimax is blowing my mind. After being cucked to 5 seconds on wan for so long, it felt like we'd be waiting forever to get natively longer videos. Then we got LTX, but it was audiovisual diarrhea and the prompt adherence was always terrible.
Then suddenly minmax h3 comes along, supporting 15 seconds, outstanding prompt adherence, can run on consumer hardware (16gb vram/64gb dram), very little censorship out of the box and already faster than wan2.2 with lightx2v.
The July-August period truly is like christmas time for video generation.
>>
>>109457867
>Does the 19gb INT8 pruned model is enough for everything
Yes
>Got a feeling its too small
Good news is that AI models don't work on arbitrary vibes
>datas
Data is already plural ESL-kun
>>
>>109457877
>supporting 15 seconds
doesnt it support up to 30s 2k? it just takes a lot of time
>>
>>109457877
LTX is good if you know what youre doing.
t.expert at LTX
>>
>>109457856
>EasyCache rapes it as always
Elaborate?
In gen time or quality of output? I was under the impression Easycache takes a heavier toll on quality.
>>
>>109457877
sorry but I don't believe opinions by nogens
>>
>>109457887
based expert
>>
>>109457885
It theoretically supports infinite length, there's not been any reports of it breaking after any length that I know of.
In any case this is the first model that would be able to gen literal movies anyways regardless of the clip length, the reference model is insane and can extend clips out of the box with vid refs.
>>
>>109457844
Of course Spectrum. EasyCache changes the output too much. Also kills the audio. Spectrum doesbn't.
>>
File: detailed video prompt.png (389 KB, 3060x1174)
389 KB PNG
It's impossible to write all this shit yourself, you need a workflow which creates it for you.
I'm working on doing that for the first time, but I'm surprised it isn't more common. You can just load a giant text model to do it, then unload it, you don't need extra VRAM.
>>
>>109457922
you don't need all that shit, basic 1-2 sentence instructions work, look at comfy example
>>
>>109457887
>t.expert at LTX
thats like being the expert in picking the best indian street food vendor, you'll be eating shit regardless.
>>
>>109457922
Why you couldn't write all that shit yourself?
>>
>>109457888
Your impression is correct, I meant in quality.
Easy Cache is faster (with Comfy default parameters) but it very noticeably degrades quality to a point where I find it unusable. Spectrum is moderately slower but quality is a lot higher. Spectrum (extension default settings, the safe profile) makes very minor, almost unnoticeable, changes.
>>
>>109457936
are you kidding me
>>
the possibilities are endless. this is only at 0.3mp too for faster iteration. if you say they fly through a wall, they do.

https://files.catbox.moe/rbg2ek.mp4
>>
>>109457931
I dont make bollywood and hollywoodslop and only make porn with it. LTX is very capable for doing so.
>>
>>109457940
Maybe almost unnoticeable is an a bit overkill description but the delta is impressively low.
>>
File: 46476.png (519 KB, 949x1564)
519 KB PNG
>>109457922
the kingdom of kino is only available to those who work for it
>>
I'm glad this shit still manages to run on 8gb vram+32gb ram and produce good results but I'm really itching for an upgrade now
https://files.catbox.moe/i034ey.mp4
>>
>OOM when trying 15 secs at 0.8MP

Its Over. Not even 64gb of ram + 5070ti is enough
>>
File: Krea2_turbo_00205_.png (3.6 MB, 1448x2176)
3.6 MB PNG
>>109457922
Regular descriptions work but a local ai model should be fine as well
>>
File: 1782981965686281.jpg (95 KB, 1280x720)
95 KB JPG
I am the only person in this world who amidst H3 release have queued 10 days worth of Chroma v48 gens.

What unique /ldg/-related records does anon probaly hold?
>>
File: Krea2_turbo_00207_.png (3.7 MB, 1448x2176)
3.7 MB PNG
>>109457960
without kroma lora
>>
Does order matter for various optimizations if you're adding Spectrum? I'm doing
> sage attention patch
> minimax mem eff attention
> minimax sigma shift
> spectrum
>>
>>109457964
>What unique /ldg/-related records does anon probaly hold?
>18821 Illustrious gens
>36679 Anima gens
>>
H3 can only do realistic, 3D CG, and animu styles. Change my view
>>
it gets more and more impressive when you realize all the stuff you can do.

https://files.catbox.moe/tg8k46.mp4
>>
>>109458003
Yes, its hard to combine all of those, like smooth animation on 2d animu styles. We need loras for this
>>
could I ask for someone to take screenshot of their task manager when they are generating a video on 3090? For some reason, my gens are really slow...
>>
>>109457964
~200000 ZIT gens with 10% of them kept
>>
File: 4754298.webm (3.1 MB, 448x256)
3.1 MB
3.1 MB WEBM
>>
>>109457940
same experience here also, spectrum is tjhe best
easy cache is total trash
>>
>>109457960
>>109457983
kill yourself
>>
>>109457964
Retrained the same lora 30 times with different settings and the best version was... the first one.
>>
>>109458034
uh oh
you sound buttmad
>>
File: Krea2_turbo_00210_.jpg (3.13 MB, 1448x2176)
3.13 MB JPG
>>109458034
+1
>>
>>109457936
Sir my English is not so good, why would I type many words by myself if I can ask ai to do the needful and write beautiful poem from my rough explanation?
>>
>>109458017
0.4mp vs 0.3 (prior)

https://files.catbox.moe/nykvhx.mp4
>>
Damn spectrum is good, shaves like 30-40% off of gen times for little loss in quality or motion.
>>
>>109457964
Manually tagged with simple tags 40000 images I generated by now.
>>
>>109458071
Youre lying
>>
https://github.com/Haoming02/sd-webui-forge-classic/issues/1380
>>
>>109458067

Yes, Sir,. the prompting is Parfect for gorgous Looks in the AI model.
>>
>>109457995
>sage attention patch
>minimax mem eff attention
These two are exclusive, if both are applied, only one should take effect, remove the regular sage.
I am testing how it combines with attn optimizations, can't report on how well it works for now.
>minimax sigma shift
This is not an opt, just parameter tuning
Possibly worth noting it won't combine with easycache.
>>
https://files.catbox.moe/cp6bp8.mp4
Nipple lora when? I can't take it anymore
>>
>>109458089
>This is not an opt, just parameter tuning
But original comfy workflow doesn't have this node
>>
>>109458092
Just commit suicide already. You've been here for ages and nobody likes you, idiot.
>>
File: MiniMax_H3n_00063.mp4 (1.45 MB, 1376x768)
1.45 MB
1.45 MB MP4
>>
lodestones just announced that he's working on a h3 finetune, cant wait
>>
>>109458092
i enjoyed the demonic piano playing ion the background
>>
Where did it all go wrong for Framepack?
>>
File: 1768895435536793.png (452 KB, 736x416)
452 KB PNG
based China made this possible.
>>
>>109458104
Idk, your mom likes me
>>
Pajeets grab Lara Croft's bobs and vegana
https://files.catbox.moe/9ct3k8.mp4

H3 is still sloppy. any way to improve quality?
this was at 0.6mp used rtx upscaler too
>>
File: 1773243075279951.png (861 KB, 2107x1153)
861 KB PNG
https://i.4cdn.org/wsg/1785851301114689.mp4
>>
>>109458124
bodied that freak
>>
>>109458101
Yes, it uses the default shift values (12 for video and 3 for audio) if you omit it. The node allows you to overwrite them. It has no effect on the speed of generation, just output quality. Read comfy_extras/nodes_minimax_h3.py
>>
>>109458154
MR ADVERTISER GET DOWN!!!
>>
>>109458154
I have hereby decided not to advertise my product on 4chan
>>
File: MiniMax_H3_00005_.webm (2.54 MB, 544x800)
2.54 MB
2.54 MB WEBM
lol I got permabanned from reddit for posting this video (by reddit admins)
>>
>>109458154
free 3 day vacation
>>
File: 5466262.jpg (294 KB, 1026x595)
294 KB JPG
is this normal for h3? generating 10sec 480p video
>>
haha wtf it made the video like a tiktok/youtube clip or something without me prompting the sounds.

https://files.catbox.moe/me1ysa.mp4
>>
>>109458154
are you prompting the boob physics or is it doing that all by itself?
>>
File: download.jpg (46 KB, 600x603)
46 KB JPG
are we behaving ourselves, anons?
>>
>>109458169
>64gb
ramlets get out REEEEEEEEEEEEEEEEEEE
>>
HAHAHA

holy shit, same prompt but diff seed. id watch this show.

https://files.catbox.moe/cmeksa.mp4
>>
File: Krea2_turbo_00225_.png (3.8 MB, 1400x2096)
3.8 MB PNG
>>
Minimax can't generate nice looking panties on anime characters. They need to be part of the ref.
>>
>>109458232
It can
>>
>>109458231
Aren't you the retard that had an existential crisis because some anon baked with will smith eating spaghetti gif was the OP?
>>
Now that the dust has settled, I think LTX won.
>>
>>109458240
Prove it. Been prompting for a bit and the panties look kind of shit and lack cute/sex appeal. Kind of like male underwear.
>>
File: 1782182198963800.png (76 KB, 300x265)
76 KB PNG
Should I be applying the "MiniMax H3 Mem Eff Sage Attention Patch" and "Spectrum Apply MiniMax H3" node to model and inputting into both BasicScheduler and Basic Guider? Or just Basic Guider and have the Scheduler without them? Instructions unclear. Google says do both.
>>
File: Krea2_turbo_00229_.png (3.78 MB, 1400x2096)
3.78 MB PNG
>>109458231
Higher steps give more details lower seems more uniform.
I need to test this more
>>109458249
Have you seen any of my gens?
If not then I wasn't on. I don't spend all my time here. Will you ever get over the fact everyone hates your camp?
>>
>>109458255
genning it right now. Give me 2 minutes
>>
>>109458249
that avatarfag in particular has been melting down for years and is the schizo baker that made the schizo rentries
>>
Fuck, the shit-grade audio I had in videos is literally frpom beta scheduler. Never change it - leave it on simple guys...
>>
>>109458251
...the losing competition.
>>
>>109458089
>>109458140

Noted, thanks anon. Do you (or anyone else) have feelings about good sigma shift settings, I'm just on default
>>
https://files.catbox.moe/i19lgh.mp4
she's having a lot of trouble finding a comfortable spot in her chair
>>
>>109458259
>Have you seen any of my gens?
>If not then I wasn't on. I don't spend all my time here. Will you ever get over the fact everyone hates your camp?
Can someone translate this schizo? I have no idea what this has to do with his meltdown



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.