[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109681974

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 311444707646165.mp4 (3.83 MB, 672x768)
3.83 MB
3.83 MB MP4
Thank you, collagelord
>>
File: ComfyUI1_13931_.png (1.27 MB, 832x1408)
1.27 MB PNG
is this the thread?
>>
File: ComfyUI1_13938_.png (1.39 MB, 872x1440)
1.39 MB PNG
>>109686888
>3 of my gens made it into tha collage
>>
File: miniconstruct5.png (338 KB, 1719x1943)
338 KB PNG
Implemented more creative controls for my Minimax H3 prompt builder.
Can encourage or filter out camera actions. Can customize the tone of the prompt, visual style, how strictly the prompt adheres to the likeness of the reference, and can easily switch music off.
Early testing showing good results.
>>
>>109686921
Mine did too. 3 exactly.
>>
>prompt pure white background on krea
>either grey or filled with noise
this happens even with the default workflow...
>>
>>109686949
have you tried not prompting a background at all? or using the mask output of a completely empty transparent png image as your latent?
>>
File: MiniMax_H3_00664__1.webm (1.95 MB, 928x672)
1.95 MB
1.95 MB WEBM
>>109686936
>>
File: ComfyUI1_13959_.png (1.34 MB, 1408x832)
1.34 MB PNG
freedom ain't free
>>
>>109687002
This is a common zen riddle.
>>
Blessed thread of frenship
>>
>>109686921
>0 of mine did
Do I get half a point for being to post one out of those three?

>pic
Who tops?
>>
File: ComfyUI1_13886_.png (1.27 MB, 832x1408)
1.27 MB PNG
>>109687055
>who tops
white hair
>>
File: 1044995260610001.mp4 (857 KB, 960x960)
857 KB
857 KB MP4
>>
I tested H3 Max, and the results weren't too shabby.
hope they release it soon
>>
File: ComfyUI1_13987_.png (1.78 MB, 1440x872)
1.78 MB PNG
>>109687002
his older brother, mr. worldwide
>>
File: Test 00019(1).mp4 (3.29 MB, 1000x558)
3.29 MB
3.29 MB MP4
>>109687089
I'd really like to see some non-cherrypicked examples of what it can actually perform. Not too hopeful, to be honest.
>>
>>109687089
Is it 50% faster than H3
>>
File: Test 00024(1).mp4 (3.91 MB, 1056x608)
3.91 MB
3.91 MB MP4
>>
File: MiniMax_H3_00780_R.mp4 (3.94 MB, 1184x896)
3.94 MB
3.94 MB MP4
finally got what i wanted
>>
>>109687110
They offer 5 free gen
You can test it yourself. The result are better than Turbo LoRA, even in fast motion

>https://fal.ai/tools/minimax-h3-max
>>
>>109687130
the sky becomes very bright. loud brass reverberate from high in the sky, startling us and causing the lens to shake. our assault rifle lowers out of frame. the lens points up to see the sunbeams coming through the clouds. a celestial being with an impossible form is flying extremely high above the clouds. it has six massive white wings. its top two wings are curled upwards to obstruct its top half, its middle two wings are twirling fast with motion blur, and its bottom two wings are curled downward to obstruct its bottom half. it has a golden halo surrounding it. as soon as a shockwave of light bursts outwards from the celestial being, the loud brass becomes distorted by a low echoing "BE NOT AFRAID.". loud brass continues to reverberate from high in the sky with various pitch changes.
>>
>>109687141
Would be cool if it wasn't gacha shit.
>>
vibe.

https://files.catbox.moe/sfh9ap.mp3
>>
File: 84590.mp4 (1.5 MB, 768x768)
1.5 MB
1.5 MB MP4
>>109687146
Yeah, that's not too shabby.
Compared to H3 local >>109687088
>>
>>109687130
It's cool, you're getting better. Still waiting for Revelation (the book of the Bible) imagery.
>>
>>109687146
Yeah, it's alright.

>>109687197
Not the same guy. Just felt inspired by one of his gens.
>>
>>109687089
>H3 max will most likely be 50-60 of GB's in size for the int8 version
Yeah but are you going to run it though? I think people don't get that if the model is both faster and better there will be a drawback coming from somewhere else
>>
File: 1115352132634244.mp4 (3.55 MB, 552x552)
3.55 MB
3.55 MB MP4
>>
>>109687235
probably. like a MoE thing
>>
>>109687235
Who said it's better than base H3?
>>
>>109687089
>>109687305
I gave H3 Max the prompt "What kind of GPU do I need" and it directly gave me this result
https://files.catbox.moe/jtt6at.mp4

It might not be better, but it will be larger. Why do you think they call it "max"
>>
File: 655734752500047.mp4 (3.95 MB, 600x600)
3.95 MB
3.95 MB MP4
>>109687305
FAL claims it's better, I think it scored higher on the ai video benchmark arena too.
>>
>>109687310
prompting a question and receiving an answer clearly means they are using an LLM to edit your prompt and include the answer with it. that wouldn't be part of the video model
>>
>>109687334
Yeah, those arena scores are pretty much bullshit. Slightly above average video model Happy Horse had Seedance 2.0 for a while.
>>
File: 49926602.mp4 (3.71 MB, 768x768)
3.71 MB
3.71 MB MP4
>>109687355
Maybe, it is at least faster, that's for sure.
>>
File: Test 00027(2).mp4 (2.69 MB, 608x1056)
2.69 MB
2.69 MB MP4
>>
File: 111830119056693.mp4 (3.6 MB, 1184x448)
3.6 MB
3.6 MB MP4
>>
File: ComfyUI1_13408_.png (1.41 MB, 832x1408)
1.41 MB PNG
poorly described a hime cut award

>>109687380
real dead or alive video game stuff, very cute, can send the 3girls pissing image in full if you need
>>
>>109687380
Are you really so fucking incompetent that you used nvidia screen capture to take a video of a videofile you already have ready on hand?
>>
>>109687400
zoomers don't know how to save files
>>
>>109687400
>>
>>109687400
Keep in mind this is your average AI slopper, and yes, they post here.
>>
>>109687400
the original image had a green top bar because it was a comfyui screenshot, wonder if that upsets you
>>
>>109687400
BUT BUTB UT I DUNNO HOW 2 ENCODE 4MB VEDEYO FOR 4CHUN SO I JUST USE RECORD DA WHOLE SCREEN CUHHH
>>
>>109687400
Color me surprised that the same guy spamming the most garabge ai slop also doesn't know how to save a videofile from comfyui
>>
does h3 ref model do a decent job of video continuation if you feed it a video input? havent experimented with giving it video inputs yet since a 5070ti struggles enough as it is and ive seen that video refs are the one thing that actually adds up to a meaningful gen time cost
>>
>>109687452
Giving it a video file as reference is the same as giving it 160 images a reference but in the context that those images belong together.
So no giving it an entire video won't do shit for what you're trying to accomplish
>>
>>109687452
>with giving it video inputs yet since a 5070ti struggles enough as it is
I give video input on 6gb 3060 all the time
just gotta make it 360p or 480p and 5 second long
>>
>>109687411
But he knows how to suck cock and post for hours on the dedicated tech thread even though he lacks basic tech skills despite being able to run ai models be it locally or online.
>>
File: Krea2_turbo_03185_.jpg (1.5 MB, 1776x2368)
1.5 MB JPG
>>
>krea just announced that krea 3 will have editing capabilities and be open weights, yet no one's talking about it in this thread.
what the fuck are you guys doing?
>>
File: 1788033426441940.jpg (534 KB, 1726x653)
534 KB JPG
>>109687397
Nah, I'm good. But thanks for offering.
>>109687400
I just used this image as the input lol
>>
i'm retarded, what's the workflow for using video references for minimax?
>>
>>109687500
what is there to talk about
>>
>>109687504
just add a Load Video (upload) node and connect image to ref_image_0
>>
>>109687452
not really. the fl2va model is better if you use the correct video continuation method
>>
>>109687500
Can't run it
If it's more safetry slop we'll have to turn back the clock for a basic finetune. I need to use a fucking lora just to get massive breast.
>>
>>109687500
We're making fun of the thread schizo for not being able to save video files or take proper screenshots and distracting ourselves from the thread's miserable state, what the fuck are you doing?
Talking about models that are not even out on api yet?
>>
File: ComfyUI1_13410_.png (1.47 MB, 832x1408)
1.47 MB PNG
otokonoko in the front, party in the back
>>109687502
refused my piss poor offer, ok
>>
>>109687461
>So no giving it an entire video won't do shit for what you're trying to accomplish
i know it ingests the images that way, but it doesn't automatically follow that it shouldn't be able to understand continuation as something to accomplish from this. it's not like when you give an only-trained-on-images multimodal LLM a series of snapshots and it muddles its way through and just barely figures out the actions implicit from the frames; it's been trained "on video" in that exact context of being given all the frames so you might as well just mentally model it as it having "seen videos", although in the case of e.g. Gemini it ingests (and so was presumably trained on) videos at 1fps so it's a bit weak for clipslop.
w/ that said, i assume you have either tried this directly or at least messed around with video ref input enough to be pretty certain it wont work, so i mostly trust you, ill probably have a stab at it later regardless and will post here on the offchance it does work for me



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.