[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109722945

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Blessed thread of frenship
>>
>>109730780
Tanks 4 bake
>>
>memorial day weekend
>plenty of weed
>plenty of snacks
>plenty of vram
>plenty of ideas
fellas... it feels so good
>>
>>109730817
>memorial day weekend
bro living in the past
>>
>>109730829
the exact names of holidays have become meaningless to me since i now goon to ai 24/7/365
>>
File: 1763261065223728.mp4 (258 KB, 782x728)
258 KB
258 KB MP4
I got a 5070ti but didn't have the right pcie power cords so it's coming tomorrow.
I have +4gb vram now so can I make videos comfortably? Pic related is what I want, I don't need SOTA minimax stuff if that's more demanding.
>>
>>109730780
>mfw Resource news

09/04/2026

>lightx2v/Minimax-h3-Turbo · FL2V Turbo 4-step v1.2 (768p)
https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/52#6a9a890895a616c64799324f

>ComfyUI NVIDIA DLSS 5 Visual Enhancer
https://github.com/Konohamaru04/ComfyUI-NVIDIA-DLSS-Frame-Interpolation

>Viggle-Animate: Character Replacement in Video from a Single Repainted Frame
https://huggingface.co/Viggle/Viggle-Animate

>DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation
https://robbyant-research.github.io/DSAQuant

>Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation
https://github.com/AMAP-ML/StateAgent

>FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
https://byeongjun-park.github.io/FlashRender

>LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
https://huggingface.co/inclusionAI/LLaDA-Image

>ComfyUI-VDN-H3: v1.4.0 — Faster streaming, VRAM-aware buffer retention Latest
https://github.com/Saganaki22/ComfyUI-VDN-H3/releases/tag/v1.4.0

>AetherScale for ComfyUI: GPU-native NVIDIA video enhancement
https://github.com/vizart-vj/ComfyUI-AetherScale

09/03/2026

>lightx2v Minimax-h3-Turbo ref2v Lora v1.0
https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors

>DreamX-Creator 1.0: Model Weights
https://huggingface.co/GD-ML/DreamX-Creator

>ComfyUI MiniMaxH3 CLIPCached: disk cache for MiniMax H3 conditioning
https://github.com/Mu5hr00moO/ComfyUI-MiniMaxH3-CLIPCached

>VDN-Minimax-H3 (VDN-H3): Hybrid Attention to Speed Up Video Models with Near-Lossless Quality
https://huggingface.co/OpenVDN/vdn-minimax-h3

>SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
https://junchao-cs.github.io/SolarWM-Web

>TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval
https://github.com/sejong-rcv/TAME
>>
>mfw Research news

09/04/2026

>OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping
https://maxtirerror.github.io/octworldpage

>One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
https://plan-lab.github.io/editvid

>Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
https://arxiv.org/abs/2609.03557

>ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models
https://arxiv.org/abs/2609.03688

>SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
https://arxiv.org/abs/2609.03806

>Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
https://arxiv.org/abs/2609.03820

>Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System
https://arxiv.org/abs/2609.04151

>LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
https://arxiv.org/abs/2609.03796

>SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution
https://arxiv.org/abs/2609.03813

>Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
https://arxiv.org/abs/2609.04183

>When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals
https://arxiv.org/abs/2609.03429

>Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization
https://arxiv.org/abs/2609.03158

>The impact of phase information for few-shot fine-grained image classification
https://arxiv.org/abs/2609.03829

>The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs
https://arxiv.org/abs/2609.04110

>Editable Visual Design
https://arxiv.org/abs/2609.04034
>>
>>109730896
Minimax works on low vram setups like yours but it will be hell. Just wait OR cope with some classic seed variation animations.
>>
How does the first frame in minmax work?
if I use different image proportions it tends to make an entirely different composition based on the image, which makes sense
but sometimes even when the image proportions match the video output, it doesn't use the reference image as a first frame
>>
>>109730896
You cant wait 1 day? XD need to goon that bad anon?
>>
File: concepo658.mp4 (3.59 MB, 960x544)
3.59 MB
3.59 MB MP4
>>109730924
>>109730936
>>
>>109730896
You could easily run Minimax with that GPU, but if such simple gens are all you want, you might just use LTX-2.5. It's a lot faster and will probably also give you what you want.

>>109730943
While it's not a huge amount, I think 16GB VRAM is plenty for doing simple Minimax videos
>>
>>109730998
Are you sure you're starting the prompt with the proper syntax per the official skill? You should be giving that markdown file to an agent to have it prompt for you desu.
>>
cozy
>>
When is the new version of GemmaPrompt coming out? no updates for 3 weeks
>>
File: file.png (912 KB, 640x640)
912 KB PNG
>>109731007
I'm cranking out 40 second one shots on the 5070ti, it's plenty at least at 0.3 mp res. https://files.catbox.moe/k93b34.mp4
>>
0.4*
>>
File: ComfyUI_04334_.png (3.14 MB, 1440x1840)
3.14 MB PNG
never prompt photo background, horrifying
>>
>>
>>109731002
Seems like the mental health is having a field day. It's weekend after all.
>>
>>109731201
aww shiet did this nigga lost some words? i am going to bully this guy with my Kentucky dialect!
>>
>>109731137
Psychologists could write volumes about this video
>>
File: output_no_audio_lossless.mp4 (3.36 MB, 1056x608)
3.36 MB
3.36 MB MP4
>>109731216
>>109731201
>upset
>>
File: H3_noAudio__00220_.mp4 (1.98 MB, 1088x1856)
1.98 MB
1.98 MB MP4
fucking h3 loras are so bad
>>
>>109731302
skill issue, been making pov anal wet pussy fingering squirting loud tongue out orgasm (all at once) videos all night
>>
>>109731313
>No feats
Bored tonight?
>>
File: ComfyUI_00826_.png (3.08 MB, 1536x1536)
3.08 MB PNG
>>
>>109731313
silly coomer porn stuff works fine with shitty loras, but they crumble as soon as you prompt something more complex like movement or change of positions
IceKiub = loras by these guy suck ass, I've tried both of his and both have gave me shitty results

slop twerk lora works wonderfully tho (not made by him)
>>
>>109731302
creepy shit
>>
why are this and anime two separate threads
>>
>>109731326
h3 loras results remind me of LTX gens, the people who are making h3 loras are probably using the same dataset and same captions as LTX, hence the same shitty results
>>
>>109731335
Dev schizo wanted try to make a cult following and weebs just took it over and made it theirs.
>>
>>109731319
too old and not obese
>>
>>109731319
I haven't seen fennec girl in ages

>>109731335
a schizo was trying to kill /ldg/, so they spammed a bunch of generals. the anime one stuck around
>>
what we needed, another schizo has arrived, since its Friday night, he's probably drunk as always
>>
>>109731112
Anyone?
>>
>>109731350
>>109731344
nono, there was a legitimate case. during wan2.1 release, the threads were devoid of anime. 1girl realism took over which pushed many posters away.
>>
>>109731354
No it was ani seething
>>
>>109731335
This is the local thread btw no ugly cloud gens plox
>>
>>109731378
He seriously paying to make a gen that shit?
ROFL
>>
>>109730780
>Reading manhwa
>Get idea to train character lora on a character
>Check Civitai
>Several people beat me to it already
https://civitai.red/search/models?sortBy=models_v9&query=Seyoung%20NA

>Get slightly annoyed at myself for not doing it first

Am I stupid for feeling this way?
>>
Extend your clips with motion context, it's fun!

https://files.catbox.moe/kxxd7a.mp4
>>
>>109731393
she's quite popular
>>
>>109731399
Pretty cool.
How many frames did you sacrifice to keep stitching look ok and preserve motion?
>>
>>109731385
He even pays to send his prompts directly to the FBI if you can believe it
>>
File: 1761061170595969.png (1.2 MB, 1273x986)
1.2 MB PNG
>>109724544
>Was the Pepe in this one added in a 2nd step or with the original prompt?
the latter
>>
someone get astra to fix comfyui
>>
FBI handlers don't often come out but I have been tasked to monitor these threads.
>>
>>109731442
What's wrong with it?
>>
File: file.png (675 KB, 640x480)
675 KB PNG
>>109731313
But what about foot insertion?
https://files.catbox.moe/emgq98.mp4
>>
>>109731192
kino
>>
>>109731378
>>109731385
Wow what's this sour grape hostility
I guess I'll take the good stuff elsewhere
>>
>>109731448
what isn't wrong with poothon bloatshit?
>>
What should I generate?
>>
>>109731561
use imagination
>>
>>109731561
Geological events, demonstrating how mountains have been formed.
>>
>>109731555
So nothing
>>
File: naisv5_32_09.jpg (3.08 MB, 2346x1340)
3.08 MB JPG
Really enjoying NovelAI V5~
>>
File: 1757552174699733.jpg (896 KB, 1984x1344)
896 KB JPG
>>
File: localdrama.png (512 KB, 829x425)
512 KB PNG
a chinese created this scene using h3. i’m not sure if he used ref or t2b, but he used local h3. could someone replicate something similar?
>>
>>109731603
you sound like a CEO that enshitifies everything for a quick buck
>>
File: file.mp4 (2.48 MB, 1920x1088)
2.48 MB
2.48 MB MP4
>>109731446
My FBI agent is a hoe.
https://files.catbox.moe/8mbwas.mp4
>>
>>109731688
They are pretty cool guys. Driven by their ranks and ambititions. Only a certain kind desk jockey monitors internet threads.
>>
File: halal.mp4 (703 KB, 736x640)
703 KB
703 KB MP4
>>109731302
it works well for me. but your image is tricky because of the phone. generally, it works great, audio sync and shit
>>
File: 1777849244685796.jpg (720 KB, 2176x1216)
720 KB JPG
>>
File: debo_sw_k2_00039_.png (3.23 MB, 1664x1280)
3.23 MB PNG
>>
>>109731736
Welcome back, master. I have generated these same ships.
>>
>>109731719
thats horrible, low resolution slop

my problem is with the details for example the right hand of the woman on the right is deformed, also the way the woman on the left turns her body is unnatural, looks like LTX tier, dont delude yourself because your coomer brain is telling you that kind of slop is acceptable
>>
File: debo_sw_k2_00040_.png (2.9 MB, 1664x1280)
2.9 MB PNG
>>109731754
we shall assemble a fleet
>>
File: 1773889016850786.png (3.99 MB, 2304x1152)
3.99 MB PNG
>>
File: file.mp4 (1.42 MB, 864x480)
1.42 MB
1.42 MB MP4
>>109731719
https://files.catbox.moe/nkm5lr.mp4
>>
>>109731771
Aye Sir!
>>
>>109731765
it's a 3-second video, you faggot. the movement can't be perfect. i'm not going to generate more minutes for you and this blue channel. at least it works. now kys. Next
>>
>>109731393
Train a better one, the-n.
>>
>>109731796
Then don’t post “it’s works great” when it doesn’t, go back to the cave you came from vramlet coomer
>>
>>109731413
Either 1.6 (~30 frames) seconds or 3.8 (~90 frames) seconds at the end of each clip depending on how good it was looking. 1.6 seconds was usually enough (though if you’re not careful with the prompt the sound might get weird)

Clips are generally 10-13 seconds (minus the carry over amount)
>>
File: 1786900788363763.png (135 KB, 414x464)
135 KB PNG
>>109731854
suck my d, you idiot who doesn't even know how to use a lora, lmaoooooooooooooooooooooooooooooooooo
>>
I thought summer was over
>>
>>109731869
Thanks anon, will definitely try once I'm not lazy enough to create multiple corresponding prompts.
>>
What is the best for make a 1 minute AI video with ref using audio?
>>
I wish someone did hentai moans h3 lora, they're so much erotic than the shitty dirty talk or carrot eating blowjob stuff.
>>
File: 1780716093742161.png (2.56 MB, 2304x1152)
2.56 MB PNG
>>
>>109731870
lol you made me laugh tho, it’s ok lil bro
>>
>>109731880
It’s a real game changer, just remember to use something that actually carries over the latents (I use a modified version of https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef vibecoded into the UI node you see above). The quality of the stitching is night and day
>>
how do you get a workflow from a video? when I drop it into comfy it simply adds it under a load video node
>>
>>109731897
why not just use prerecorded moans as audio input?
>>
>>109731681
Still nothing
>>
>>109732158
you must be a bad CEO if you make nothing
>>
>>109731993
I've seen quite a few context continuation nodes, is this the best one?
>>
>>109732149
A lora would be more varied and correspond better to whoever voice timbre I want to use.
>>
>>109732063
As far as i'm aware, you don't. At least this is how my ComfyUI install works i'm pretty sure workflow. Metadata is only embedded in PNG files, and there's no current implementation to embed that into a video file. I don't even think there's a standardized way to do that unlike there is for PNG files. You could probably create a custom version of the
Save Video
node but then, any metadata embedded into the video what pretty much only work on YOUR local install of ComfyUI ( and also ensure comfyui even supports dragon drop metadata instruction from videos in the first place).
>>
>>109732184
but it would sound tinny. real audio is better quality
>>
>>109732175
Retard good-for-nothing no coder, you're supposed to explain why it's supposedly sucks ass instead of being a crybaby dumb fuck artists who expect everyone to bow down to its opinions.
>>
I am seated, yes, at the kinoplexitorium.
>>
>>109732187
Huh? I can drop videos into comfy and get the workflow every time. It would be extremely annoying if this didn't work.
>>
>>109731302
i used few of such + acts. horror show. avoid them at all costs.

>>109730312
>>109728048
>>109728701
which model and what lora is this?
>>
>>109731771
very cool
model and lora plz?
>>
File: debo_sw_k2_00052_.png (3.19 MB, 1664x1280)
3.19 MB PNG
>>109732345
no lora, base krea2turbo
>>
No other flat chest lora on Minimax ? The one i found in civit is shit
>>
>>109732365
i got kera few weeks ago after not considering it at all during release, it is good for styles indeed it is, has sd sdxl feel in gens

basic prompt for the material look in your prompts plz? has star wars vibe
>>
File: debo_sw_k2_00055_.png (3.09 MB, 1664x1280)
3.09 MB PNG
>>109732464
as with all my prompts, its quite a mess, so its hard to know exactly whats kicking in for the aesthetic. the logline is
>astronomy picture of the day photograph of a spaceship
this seems to have a very nice quality of tethering to near-contemporary tech. I later have a mess of
>film photography aesthetic, photojournalistic realism, documentary photography quality
hard to say whats doing the heavy lifting

I'm happy to give a catbox too if thatd be more helpful
>>
>>109732505
ty notes taken will test
and catbox sample would be much appreciated
that is kick ass desktop tier gens material
>>
File: 1759421665645355.jpg (983 KB, 2400x1600)
983 KB JPG
>>
File: debo_sw_k2_00062_.png (3.28 MB, 1664x1280)
3.28 MB PNG
>>109732559
https://files.catbox.moe/ypt31r.png
hopefully there can be some helpful stuff in there
>>
>>109732594
it werks!
tyvm, yes same style, looks great

wildcards in your prompt, look; not maintained anymore i think,
original (llm stuff added though, pull oldest one if you want basic wildcards + lora loading from encode node this pack offers):
>https://github.com/Tinuva88/Comfy-UmiAI

or check this fork of it
>https://github.com/I-ShadowStar/Comfy-UmiAI-Indemnitate
>>
File: 1763042999745642.png (2.98 MB, 2304x1152)
2.98 MB PNG
>>
File: 1760002037081150.png (2.6 MB, 2304x1152)
2.6 MB PNG
>>
>>
>>109731917
These are all really cool. Make one with Gondola too. Should fit the style nicely.
>>
>>109732857
i like the idea but i dont think base anima knows gondola
>>
>>
Is anyone even interested in music sloppa? I never see it posted here.
>>
note: the 0.1 ref2v lightx2v lora is better at lower res than 1.0, where you need to generate at 0.9-1.0mp or you get more noise.

so 0.1 at 4 steps seems ideal.
>>
>>109732991
music sloppa? funny you should ask
https://suno.com/s/SHLLmAWcdsl0MDtg
>>
>>109733006
for example this is a quick test at 0.3mp with 0.1. it works, now I need high resolution. 4 image references, works fine.

https://files.catbox.moe/za0dp9.mp4
>>
>>109733029
That's not local...
>>
>>109733036
https://files.catbox.moe/f4n6xj.mp4

another test (not the 10s vid)
>>
gem alert?

https://files.catbox.moe/y7cns1.mp4
>>
File: 00031-3382807428.jpg (411 KB, 2624x1536)
411 KB JPG
>>
>>109733045
tough shit
>>
File: 00035-3764416362.jpg (409 KB, 2624x1536)
409 KB JPG
>>
https://files.catbox.moe/0d2w01.mp4
>10 feet behind them is gigachad, a 7 foot tall muscular man in a black speedo who is wearing a black tshirt with the design of <Picture 5>, in a full greyscale style with no color.
>it works

what a fun model.
>>
>>109733171
kill yourself
>>
>>109733176
go ACK
>>
>>109733178
boring
>>
File: 00045-1040069846.jpg (492 KB, 2688x1536)
492 KB JPG
>>
getting closer and closer to the desired result.

https://files.catbox.moe/h4esp4.mp4
>>
File: 00056-3292340106.jpg (537 KB, 1920x2880)
537 KB JPG
>>
1mp was worth the wait

https://files.catbox.moe/uez98b.mp4
>>
i meant what i said
>>
>>109733171
why is gigachad wearing hp?
>>109733280
cute
>>
>>109733347
Welcome back, Master~!
>>
File: 00062-3428528645.jpg (516 KB, 1728x3072)
516 KB JPG
>>109733171
>>109733246
>>109733288
what is the genuine point of these videos. No wonder nobody takes local serious to begin with. At least troons actually make spicy shit worth watching and rewatching.
https://files.catbox.moe/45oeci.mp4
https://files.catbox.moe/6ian1y.mp4
https://www.deviantart.com/shiftingfun/gallery
>>
>>109733383
the point is to make them seethe
>>
>>109733407
>them
Are "them" in the thread with us right now?
>>
>>109733436
given /lgbt/ there are statistically at least a few.
>>
>>109733407
You are suffering from self-reinforcing delusions.
>>
>>109733407
this feels pointless, most troons are on reddit and discord. The pro ai troons are less ideological extremists compared to the anti ai normie zoomer cattle because their programmed default is already pro trans because its the "social acceptable" position. It's the ai haters and luddites you would want to focus your energy on offending and trolling. They are the ones that will likely get open source and chinese models shut down for good.
>>
>>109733485
I dont have a goal, it's for fun. troons getting mad on twitter is just a free serotonin hit.
>>
how good is audio reference in H3
I used a 15-second audio ref and told Clanker to repeat the same phrase. It produced a good 5 second then the rest was just yapping with the same accent
>>
>>109733458
>statistically
Statistically it's less than 0.1% of the population and there's maybe ten people posting here.
>>
>>109733501
use a clip that is like 10s of audio with no music

then just say <picture 1> is the reference for Guy, with the voice of <audio 1>

then when you prompt dialogue it should clone it. worked for my JC Denton test, etc.
>>
File: 00067-392321602.jpg (863 KB, 3072x2048)
863 KB JPG
>>
>>109733518
>then when you prompt dialogue it should clone it. worked for my JC Denton test, etc.
You have to actually prompt dialog? I just tell it to repeat the exact same phrase in the audio. I didn't prompt the dialog
>>
>>109733501
just remember to code all your spoken lines with <d>text here</d> otherwise it will add random gibberish lines.
>>
>>109733520
is this khroma?
>>
>>109733529
oh, you want the gen to copy the audio? in my example it's for voice cloning it, then you prompt what they say and it copies the voice.
>>
>>109733534
yea like, the audio ref is a person singing (no music). Do I have to prompt the lyric? cuz "make video of this person sing <audio 1> " doesn't seem to work
>>
>>109733529
i'm also not sure actually, desu I haven't been able to get the voice cloning to work. I once tried making a clip of Trevor from GTA 5 say something, even provided a voice sample but he kept sounding like Michael. Not sure if the Michael voice is just hard baked into the GTA 5 token.
>>
>>109733549
not sure so far ive just tested voice cloning, but if you supply audio you could just do "sings to the music in <audio 1>" maybe?
>>
>>109733549
You have to prompt the lyrics and reference the audio just like in the guide
>>
File: batman.png (2.26 MB, 1920x1080)
2.26 MB PNG
good talk

batman
https://youtu.be/DxMJ0I_JrTQ
https://suno.com/s/9DrJWWt0Ufajrjf9
>>
>>109733549
I tried this one as well with very poor results. I provided the song a reference image and a voice sample trying to replace the voice in the original song. I had results where it tried to play the entire song in a 15 second clip. Then I manually cut the first 15 seconds of the song, provided it with lyrics but eventually rage quit.
Figuring out the prompt is hard. I'm telling it to use the use the audio 1 as the music being played, audio 2 as the voice to replace the one in the original clip, but it keeps on changing the music or only starting to play it at the end of the clip.
I guess i need to look at the prompt again but figuring it all out becomes quite daunting.
>>
File: 00074-1283841083.jpg (811 KB, 2048x3072)
811 KB JPG
>>109733533
nope its just krea2
>>
>>109733623
looks very clean
>>
cozy breas
>>
>>109733647
t
>>
File: 07108-2875562454.png (938 KB, 896x1280)
938 KB PNG
>>109733647
>>
File: Test.mp4 (3.51 MB, 768x1344)
3.51 MB
3.51 MB MP4
>>
>>109733687
nice
>>
why isn't she naked?
>>
okay now im satisfied. ive also confirmed the 0.1 ref2v lora (lightx2v) at 8 steps, euler/simple, is ideal. the 1.0 one for me has more noise even with euler. I wanted to test multiple things:
>Use <Picture 1> for the physical identity of gigachad.
>10 feet behind them is gigachad, a 7 foot tall muscular man in a black speedo who is wearing a black tshirt with the design of <Picture 5>, gigachad is in a full greyscale style with no color

what a model. ref model in a sense makes it unnecessary to have loras, which are good but in this case the picture reference is the lora.

https://files.catbox.moe/ofa2x0.mp4
>>
File: 00086-3612050057.jpg (642 KB, 3072x2048)
642 KB JPG
>>109733637
its a nice model but its a shame the community at large won't give it a chance. This space has become too fragmented between vramchads, 16gb mid-vramlets and poor vramlets. Krea2 lora section on civitai is depressing to look at. Its truly the human centipede of boring and nasty fetish slop with a very few goodies buried in there. Minimax h3's lora section is even worst to look at.
>>
>>109733716
what do you mean? krea 2 is the standard for image gen, klein edit 9b for edits, and minimax for video. krea 2 is a great model.
>>
>>109733716
idk man to me the krea2 loras look pretty good. Lots of cartoon styles, which is the kind of stuff I like. What exactly are you looking for because to me it seems like the Civitai Krea area is bussin' (no cap)
>>
File: 1766862111699754.png (2.72 MB, 1344x1984)
2.72 MB PNG
hm...
>>
>the camera hard cuts to <Picture 1> showing troon diving at high speed towards the water as the camera tracks them, speed lines show them traveling extremely fast

you can use an image reference as a keyframe. kino

https://files.catbox.moe/4ofydr.mp4
>>
File: monerochan.mp4 (3.32 MB, 854x480)
3.32 MB
3.32 MB MP4
This gen was fun
https://files.catbox.moe/77k3u0.mp4
>>
>>109733821
better phone screenshot:

https://files.catbox.moe/jkyzys.mp4
>>
File: 00075-4097748664.png (2.85 MB, 1248x1824)
2.85 MB PNG
>>109733737
i want more character loras and not just styles and fetish nsfw stuff flood the page on civitai. illustrious page on civitai has some much good character loras to play with unlike krea2. It just agonizing and painful to see how pale the krea2 lora section is. Come on anon, who the fuck wants to download and use a lora of a fake non-real photorealistic ai influencer model OC of another user. May be my expectations are just too high and I'm out of touch with the kinda stuff people like in the space. I know people are hungry for character loras judging by the numbers of downloads for several of the ones I've made and uploaded so far on civitai.
>>
What stuff with AI can benefit from having 2 (or more) graphics cards? LLMs can load models across GPUs to benefit from higher VRAM; can image gen or video do the same?
>>
>>109733926
Minimax supports tensor parallelism and so do many other models as well. Though you probably won't come far if you have to ask
>>
>>109733926
you can gen 2 1girl at the same time
>>
File: 1785830629972134.png (2.56 MB, 1216x2176)
2.56 MB PNG
>>
>>109733926
looking for that one lucky seed out of five hundred that is barely usable in H3 gets faster if you have more GPUs
>>
File: 0ea41fa4d850.png (945 KB, 887x1774)
945 KB PNG



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.