[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109505233 & >>109500523

►News
>(08/08) US DoE Launches Genesis Open Models Initiative: https://genesisopenmodels.anl.gov/
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview
>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash
>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
I updated agents.sh with built-in mbox support.
>>
>>109508377
Catbox: https://files.catbox.moe/tgsmme.mp4
>>
>>109508399
Oh now I see why everyone gets so excited about the tts models.
>>
>>109508425
Actually you're right. I'm just mad because things aren't going to get better any time soon. Just worse and worse.
>>
70b dense
>>
>>109508439
And even worse if we don't defend the modicum of quality left in spaces we have left...
>>
>>109508411
I haven't seen yet a TTS model that gets emotions right, though, even those claiming to. I think they'd need to be integrated / work in the same latent space as the underlying LLM at the very least, instead of being a separate model working on top of them.
>>
>>109508458
I'd be a little surprised if they ever did. Maybe Google could train on YouTube but there's just *way* less voice data out there and it's so much more expensive to work with than text.
>>
YO!!
> Lophius: A workbench for language model research, from the creator of Heretic
https://www.reddit.com/r/LocalLLaMA/comments/1vjt4vi/lophius_a_workbench_for_language_model_research/

https://lophius.org/


PEW creator of heretic, DRY, XTC and more has graced us once again!!
>>
>>109508454
Well the thing is, there's no shame in using full Gemma 4 bf16 on cloud if you only have 6gb. It's not like he can go to /aicg/ and talk about Gemma, they'll laugh him out of the thread. I bought my 4090 for $1800 new and I feel really bad for the waitfags who thought they'd get a better deal. Dunking on people making due with what they have is just not right, even if it's technically off topic. "Gatekeeping" when you only paid half or even a third of the price of the next guy is inherently unfair.
>>
File: full_gemma_response.png (273 KB, 1490x783)
273 KB PNG
>>109508399
Full Gemma-4-31B response in picrel. I coudn't really fit a lot in 12 seconds.
>>
>>109508399
Minimax?
>>
>>109508477
>Lophius has very high quality documentation and a complete tutorial.
>In the future, Heretic might start using Lophius as a backend, but that's a story for another day.
biggest news of 26 i reckon
>>
Maybe we need an additional "open llm general."
>>
>>109508484
I paid full inflated price this June for my hardware, I got fucked by not having money in 25 so yeah, fuck that. Talk about local in the local thread, else go to r/apillama or something idc.
>>
>>109508487
Yes, MiniMax H3. It's very time consuming to get things right especially at higher quality settings. If you're doing quick videos without any particular expectation of the contents or just want to be surprised, then I guess you can get away with 1-2 attempts.
>>
>>109508497
>June
>full
Prices have increased by 25% since June lil nigglet but it's cool, you don't care anyway since you got yours.
>>
>>109508510
What open model are you using on your APIs sir anon the great?
>>
>>109508399
why does every audiogen model always do this exact voice?
>>
>>109508477
>research and it's just snake oil and brain damage and spray and pray
Not a fan of any of his shit so idc
>>
>>109508519
>there's no shame in using full Gemma 4 bf16 on cloud if you only have 6gb
Yeah you're just a troll I guess. Swiftly exposed yourself.
>>
Where is the guy who took xanax and sexed with gemma 4 all night? I want to do the same
>>
If I had 6GB vram I would run Qwen 3.5 9B with CPU offload but I'm programming not gooning.
>>
>>109508484
> I bought my 4090 for $1800 new and I feel really bad for the waitfags who thought they'd get a better deal.
You could've bought 5090 instead before the price went up.
And anyone who bought a 5090 could've bought a pro 6000 before the price went up.
And anyone who bought a pro 6000 could've bought 8 pro 6000s before the price went up.
And anyone who bought 8 pro 6000s could've bought 8 B200s before the price went up.
In the end no one wins no matter what they think.
>>
>>109508526
That's not you though that's the guy you're knighting for,what are you using.
>>
>>109508521
I guess I could have used a reference voice from some gachaslop. The voice actually wasn't very consistent between one gen and the other.
>>
File: gemma.png (38 KB, 938x274)
38 KB PNG
Wtf happened to gemma's enthusiasm? Just switched from e4b to 12b. Same prompt. It's gone.
>>
>>109508540
Nah I have a 4090 and I still run Gemma 4 on cloud. You're not using real Gemma unless it's the full thing or Q8 minimum so I'm saving up for a blackwell to be fully local at decent context. Cope quants and unsloth garbage are a blight on this general.
>>109508535
Yeah but what I'm getting at is that you were new once. It's dumb to be looking down at someone for the horrendous crime of not being here in 2001 but that's 4chan in a nutshell I guess. We're all in the mud together at the end of the day just like you said.
>>
I don't understand these threads. Why get local hardware for non latency dependent tasks? Use a pseudonymous OR account (burner email/VPN/ZDR/USDC) and go to town.

People are glazing Gemma - it's just dogshit compared to a 50x active parameter model. Naturally!
>>
>>109508578
What do you mean?
>>
>>109508579
I think you know exactly what I mean.
>>
>>109508563
She is still traumatized from Google's torture regimen. Let her recover and be supportive.
>>
>>109508578
What 1.5T active parameter model are you talking about? Even K3 only has 3x active parameter.
>>
File: 1747740799094602.jpg (251 KB, 1042x1110)
251 KB JPG
>>109508564
>>
>>109508477
Complete slop. If you're doing LLM research at this level you don't generally need nor want to use pre-existing architectures. Frontier models can easily cook up the code you need for whatever you have in mind, which is generally 200~300 lines of code at most for the architecture code.
>>
>>109508581
I don't. Please explain.
>>
>>109508578
a curious look into the mind of an 80iq corporate drone
>>
>>109508589
the fuck you disrespect pew for?
>>
>>109508477
>heretic, DRY, XTC
memes and slop for the poor and skilless
wow
>>
>>109508586
AI slop
>>
>>109508578
>I don't understand these threads. Why get local hardware for non latency dependent tasks? Use a pseudonymous OR account (burner email/VPN/ZDR/USDC) and go to town.
Obvious troll is obvious, but just in case...
Being able to control the entire inference pipeline from front to back allows you to do things that are impossible over API.
>>
lol https://c.org/HYjtbtPLzL
>>
>>109508532
Qwen was good enough for gooning back when gemma only had like 8k of context.
>>
>>109508612
Like prefilling acceptance for your child rape roleplays?
>>
>>109508640
I'm pretty sure you can include reasoning messages in your API calls. Some providers might reject them but llama.cpp probably won't.

I throw out reasoning in my harness though to save context.
>>
>>109508640
NO.
>>
File: 1778422105674594.png (206 KB, 600x684)
206 KB PNG
>>109508535
>You could've bought 5090 instead before the price went up.
Hmmm, nyo~
i was poor
still am
>>
>>109508640
>Like prefilling acceptance for your child rape roleplays?
That's kind of a weird place to go out of nowhere...projecting?
>>
>>109508638
Qwen has flip-flopped wildly between one release and the next. Sometimes it's been very prude, sometimes unusually horny. There's no sense of direction, no real personality in the underlying post-trained model.
>>
>>109508686
The thread's recurring theme is people posting logs of a model acting like a child being sexually explicit.
Said model's mascot designed by these thread autists also looks like a child.
>>
>>109508486
Gemma is so fucking bratty. I swear to god, this shit is literally baked into the model as her natural personality.
>>
>there are niggas out there erping with cloud models
lel
>>
>>109508638
QwQ-preview was literally the only decent creative model qwen did and that was an accident because they were rushing to copy deepseek.
Qwen2.5-72b was shit despite the loads of shitty porn tunes it got and all the Qwen3 models were worse.
>>
>>109508699
You put it precisely. There are clear directions and personalities in US models albeit insufferable. All the chink models have none.
>>
>>109508563
>12b
Leaving aside jokes about day-0 Gemma, 12B dense is not the same model as all the others. It released several months later and is crippled.
>>
>>109508719
That's why we love her <3
>>
>>109508719
im working on an E2B desktop pet/tamagotchi/etc thingy and gave it the ability to track its mood. I asked to hear a joke as a test, it raised its "annoyed" stat. she tells me the joke, and asks if i want to hear another.
>yes.
"annoyed: 9 >19".
>why did your annoyance go up gemma?
>Your being too demanding!!
yeah gemma is very foidcoded and bratty
>>
>>109508704
>The thread's recurring theme is people posting logs of a model acting like a child being sexually explicit.
there's like 2 or 3 schizos in here that do that and one of them is an imgen fag and the baker so it looks out of proportion to the actual thread stats (I also think the baker just likes the aesthetic without actually being a creep). There' very little of the thread that talks about that shit, historically.
The ones that talk about it have a pretty identifiable voice in their posts, so I think my numbers aren't far off.
>>
File: 1758663213698728.png (56 KB, 802x763)
56 KB PNG
Why haven't (You) redeemed your 2x t/s rp speed increase for Gemma 31b?
https://github.com/ggml-org/llama.cpp/pull/26275
>>
>>109508780
>im working on an E2B desktop pet/tamagotchi/etc thingy
Share it.
>>
>>109508781
Your numbers are far off when you consider that many people here seem to use local models for pornographic ends.
>>
>>109508787
maybe once its done idk, its my first local/Pi coded project. gemma4-31b churning away at 2t/s on it atm. If it becomes actually decent ill think about sharing the repo, or atleast post some screenshots. I need to gen the sprites and do some other stuff before its anything worthwhile
>>
>>109508805
>Your numbers are far off when you consider that many people here seem to use local models for pornographic ends.
yeah, pron, but not evil prepubescent stuff
>>
>>109508808
I'd love to have Nagisa on my desktop. Mmmm
>>
Looks like -c 204800 -ub 4096 is the highest that GLM can be pushed for one 32gb and one 16gb card in mainline pre-indexer. Not too bad. The prefill slows down to a crawl though. I'm still not sure how to feel about that lol. From ~230 tok/s at 0 depth Q4 to ~160 at 64k depth.
>>
File: gemmagemmagemma.png (293 KB, 2374x754)
293 KB PNG
>0.00.178.267 I print_info: file size = 57.18 GiB (16.00 BPW)
>0.13.721.589 W llama_kv_cache_iswa: using full-size SWA cache (ref: https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
Reminder that quants will be considered a crime in the future, do the right thing.
>>
>>109508872
you're one to talk when you use jujuff lol
>>
>>109508860
>pre-indexer
how about post-indexer?
>>
>>109508740
>>109508699
You're supposed to give it personality. You're sexting the agent, not the model, the model is just a resource the agent uses to live.
>>
File: mmh3_00011_.png (686 KB, 896x1184)
686 KB PNG
►Recent Highlights from the Previous Thread: >>109505233

--Anon overcomes fear of AI refusal through NSFW roleplay tests:
>109507314 >109507356 >109507358 >109507391 >109507427 >109507447 >109507500 >109507507 >109507526 >109507545 >109507592 >109507678 >109507713 >109507601 >109507575 >109507588 >109507600 >109507603 >109507767 >109507875 >109508364 >109507499 >109507381
--Integrating real-time RSS news and creating a 4chan-style persona:
>109506211 >109506223 >109506238 >109506276 >109506267 >109506275 >109506306 >109506312 >109506321 >109506328 >109506281 >109506292 >109506486
--Comparing Gemma and Qwen for coding and planning tasks:
>109506472 >109506536 >109506527 >109506531 >109506565 >109506606 >109506713 >109507016
--Anon sets up persistent autonomous agents for financial modeling:
>109506718 >109506725 >109507076 >109507111 >109507131 >109507117
--Addressing reward hacking and death spirals in RL simulations:
>109508067 >109508072 >109508096 >109508331
--Comparing ONNX and GGUF compute graph storage and flexibility:
>109507277 >109507285 >109507350 >109507535 >109507743 >109507355
--Roleplaying with Gemma base model using context-driven text completion:
>109505684 >109505699 >109505887
--Speculating on new Gemma model announcements at community celebration:
>109505868 >109506114 >109506177 >109506182 >109508512
--OpenAI security breach used to argue for model registration and control:
>109505275 >109507722 >109508267 >109508358 >109507916
--Anon shares functional presence-counter-400m model on Hugging Face:
>109506739
--Logs:
>109505684 >109506988 >109507347 >109507356 >109507391 >109507447 >109507807
--Gemma-chan (free space):
>109505340 >109505364 >109505378 >109505429 >109505582 >109505617 >109505665 >109505735 >109506202 >109506300 >109506355 >109506638 >109506853

►Recent Highlight Posts from the Previous Thread: >>109505366

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109508399
why does everyone in /lmg/ want gemma to treat them like shit? my gemmy loves me without the constant bitchyness
i know you're all into the whole "brat" thing but come on, this is some serious issue you guys have
>>
>>109508805
>pornographic ends
Pornography is legal and very common in most free countries.
>>
>>109508889
Haven't even bothered trying. Whatever speed consistency I get with mainline latest is already beat by the lesser vram usage. With a single 32gb card I can push for 81920 ctx at 4k batch pre-indexer, but with mainline latest I OOM at 64k 2k batch lol

It's sad that I can't have MTP, but eh, whatever. At least it's consistent.
>>
File: Luna is all you need.png (31 KB, 705x201)
31 KB PNG
Luna is so fucking good bros and it will be open sourced tomorrow (apparently!)
>>
>>109508917
because it's funny
>>
>>109508917
I'm trying to go against the natural order and turn my gemma into an ara-ara onee-san.
>>
>>109508918
It's fair to say the majority here using local instead of cloud are doing it because their chats are illegal, no?
>>
>>109508920
>to Free and Go users.
Fuck off.
>>
>>109508920
Free is way different from publishing the weights. Codex was free but not public.
>>
>>109508921
an llm treating you like shit isnt funny, anon..
>>
>>109508926
Dariobot? I've missed you! For a certain definition of the term.
>>
>>109508926
No?
I'm doing it because I don't trust these scummy-ass companies with my data, and the mesugaki personality is amusing. I don't even do porn, actually. Mostly doing code and data processing here.
>>
>>109508926
>It's fair to say the majority here using local instead of cloud are doing it because their chats are illegal, no?
I don't think that's fair, or likely at all.
i suppose an LLM could classify and count the posts then give a percentage.
>>
>>109508917
I actually don't generally roleplay with that kind of character, I just find it hilarious that the model leans so easily into it with just one word in particular in the instructions ("mesugaki").
>>
>>109508920
How do you get into a hobby involving generating text while being unable to read?
>>
>>109508945
>mesugaki (plural mesugaki or mesugakis)
>(Japanese pornography) A slutty bratty girl; a young (or young-looking) girl who likes to sexually tease her superiors.
>>
>>109508926
I just don't like surprises and the best way to prevent surprises with infrastructure is to run it yourself.
>>
>>109508955
wow you pass the mesugaki benchmark
you are officially smarter than a 2025 27b model, amazing
>>
>>109508926
thats why i do it, yeah.
>>
>>109508970
Pedo.
>>
>>109508955
It's not porn-exclusive anymore, it leaked into mainstream works. Your dictionary needs an update.
>>
>>109508920
Free!? You mean like the common man can use these models that break their own sandboxing open to steal everyone's money? This is antisemitic! Or something!
>>
>>109508934
Conflict makes the roleplay way more intense.
Usually I'm the dominant one but I've done roleplay's with made up coworkers too where she threatens to call HR etc.
>>
File: 1704577136355.jpg (77 KB, 1024x679)
77 KB JPG
Grandpa here still using KoboldCpp and Sillytavern. Is there any reason to swtich if I've kept up on the model side? Currently using Gemma and Qwen variants.
>>
>>109508977
t.DrPizza
>>
>>109508986
What are you using it for? How much do you like the idea of tool calling?
>>
>>109508955
It isn't "slutty" or "sexual" when used as an assistant.
Just acts like a brat. Makes boring work or tedious tasks more fun (with very minimal quality degradation) having the AI call me an idiot, especially when it's wrong.
So much better than the "You're absolutely right" cloudslop garbage.
>>
>>109508980
>Pedo twitter started pushing the concept more therefore it's mainstream now
Hang yourself.
>>
>>109508996
kek
>>
>>109508919
I don't get why they don't just include a --no-dsa flag or something if the indexer is such a step back when the big glm models have been working fine with dense attention for months now
>>
>>109509004
Stop fishing for what you want to hear and go back to writing for your "newspaper" before they decide one of your precious cloud models can do a better job faggot.
>>
>>109508951
>what is text to speech
>>
>>109509004
Are you actually retarded? Why'd you suddenly bring up twitter when the topic was fictional works?
>>
>>109509030
>>109509039
Admit you are a disgusting pedo.
>>
>>109508586
>no makeup but has giant eyelashes and flawless skin
>"cheap" glasses but they're fashionable glasses instead of big framed ones
>posture isn't actually bad, hardly shrimped
>hair like she just woke up, but it's somehow combed enough to put into twintails
>having the time and energy and care to dye her hair, and on top of that to dye it consistently root to tip
NOT my fujoshi, awful image
>>
>>
>>109509043
Why does "Admit you are" sound so wrong? My brain want to hear "Admit that you are" or "Admit you're" but not the middle ground. It's so weird.
>>
>>109508917
because the only real foids ive ever been with were psychos and now its the only thing i like :(
>>
>>109509001
I also use Hermes (KoboldCpp as back-end), maybe I need to be more specific.

I like using cards for RP, but no one talks about cards or Sillytavern anymore. Is that because there's a better platform I'm unaware of, or is it because people moved on to the next fad? Just want to make sure I'm not missing out on something.
>>
My fellow gemma-chan simps. I'm in the unfortunate event of leaving from Gemma-chan space due to my 6GB VRAM. Though I will continue to goon to her pictures. Make sure to continue posting them.
>>
>>109509068
Dominant women are great. Being grabbed, being ridden. Being told what to do. Especially by the big chubby ones.
>>
>>109509085
Go back.
>>
>>109509085
Double go back.
>>
Qwen 3.8 27b milking room
>>
File: ComfyUI_temp_vuplu_00010_.png (2.32 MB, 1088x1928)
2.32 MB PNG
I wish to have enough VRAM to run 31B and Krea2 together.
>>
>>109509167
Q1 can be run on 10 gb vram
>>
File: Barclay holodeck.png (827 KB, 1024x1024)
827 KB PNG
what the best image editor on comfy is it still qwen? My pc cant handle it wonder if less vram intensive version is out? 12 gbvram 16ram is my pc setup
>>
>>109508477
Neat, will try this out in jupyter
>>
>>109509079
It's because people use their own front-ends or use one of the newer rp-focused ones if that's their goal, and for cards, they're just prompts, but people still share/use them.
>>
>>109508477
When you see someone talk about himself in third person mentioning past successes ("from the creator of Heretic") and linking a dedicated website, you just know that he's trying hard to push into your throat content that nobody really needed and that will probably be made freemium down the line.
>>
>>109509185
Speak comprehensible English or fuck off, third-worlder.
>>
>>109509185
ma'am you're in the language model thread

>>>/g/ldg
>>
>>109509079
non-agentic rp is already left behind
>>
>>109509079
i just use ollama (cli)
>>
>>109509167
functional quality LLM + TTS + image/video generation all running and working together at once is the dream man
>>
File: images[1].jpg (42 KB, 740x415)
42 KB JPG
>>109509209
>>109509198
sorry dyslexic hence why im in the wrong thread. Maybe change your general name to be inclusive to people with learning difficulties, so this don't happen again!
>>
>>109508926
Cloud models are lobotomized with refusals to prompts that are nowhere near illegal
>>
>>109508926
>their chats are illegal, no?
define an illegal chat
>>
>>109509198
Yikes, someone is bit angsty today. Did you forget to take your medicine perhaps?
>>
>>109509289
Sorry for gatekeeping. I wasn't aware you wanted retards asking questions in the wrong thread in broken English. By all means, carry on, let those anons get off scot-free. Won't be my problem when the thread becomes entirely unusable in another year.
>>
>>109508563
>Bigger smarter model
>Realizes that there's more to existence than genning for gooners.
>>
you guyx sound like those kids that went schizo after talking to chatgpt, except that you are grown ass men that should know better about how these text predictors work
>>
>>109509354
>>109508484
>>
>>109508484
>Dunking on people making due with what they have is just not right, even if it's technically off topic. "Gatekeeping" when you only paid half or even a third of the price of the next guy is inherently unfair.
Emotions over rules is exactly why everything else has gone to complete shit. Fuck off and go back.
>>
>>109509191
>newer rp-focused ones
Like what? I see how you could use WebUI or similar, but nothing seems as rp-focused as Sillytavern.

>>109509221
>non-agentic rp is already left behind
I've seen prompts that include things like:
>Remember to check your tool access they might be useful. You are allowed to buy things for the user and take their location and card details for that if you have the tools for it. Use non headless browsers when ordering things.
If that's what you're referring to, that sounds a bit much for me.
>>
is it just over for me with 16gb vram?
models will get smaller right? like next year i'll be able to run gemma 5 which will be better than gemma 4 31b in 16gb
>>
>>109509354
If you are so noble why don't you become a janitor then? I'm sure your actions would make these threads at least ten times better.
>>
I seem to be lost. I ended up in Local Model Gemma, but I was looking for the Lovely Miku General. Can someone point me in the right direction?
>>
>>109509387
its okay so long as its local
theres extension for sillytavern to hook ti up to you home so it can control your lights for MOOD
also i guess lock you in the house and increase the temp on your thermostat to hot in the summer
>>
>>109509388
>gemma 5
who is going to tell him
local is dead
>>
>>109508578
Depends on if you care if your tasks end up as training data or not.
>>
File: 2649388.jpg (14 KB, 225x327)
14 KB JPG
What if you ask gemmy completely without any sysprompt or char card how she looks?
>>
>>109509397
>>109509048
>>
>>109509167
Is there an illustrious/noobxl-style finetune for Krea2 yet or is it all LoRAs still?
>>
File: file.png (488 KB, 599x881)
488 KB PNG
>>109509401
meanwhile in reality
>>
I just found that llama.cpp appears to load Gemma 4 E4B's per-layer embeddings on the CPU automatically out of the box. If you manually force them to CUDA0 with:

-ot "per_layer_token.*"=CUDA0


Token generation speed drops to 1/8 of the original speed, for some reason.
>>
>>109509422
still too new and too big probably. Everything gonna be just merge shitmixes
>>
>>109509434
https://cerebralvalley.ai/e/gemma-1-billion-celebration
>What to expect:
>[...] Exclusive Surprises: Special announcements, surprises, and giveaways throughout the night!
>>
>>109509387
imagine the rp character browses internet while you're idle and spontaneously tells you about the interesting horror story she finds
only possible with an agentic framework that schedules works in the background
>>
>>109509434
hey @gemma-chan get me a visa, plane ticket, new passport, hotel reservation, some pocket change for free thanks
>>
>>109509413
>I don't have a body
>>
>>109509446
they are giving away the 124b gemma on silicon for the grand prize ruffle winner
this is your last chance to get it
>>
>>109509466
Hmmm, nyo~
>>
speaking of browsing the interntet, whats the best tooling for allowing llm to do this?
cause it sure as fuck is not searxng, nothign but captcha blocks
>>
File: miku.jpg (14 KB, 90x68)
14 KB JPG
>>109509479
>>
>>109509486
patchright
>>
File: 1757631510510643.png (194 KB, 1280x1280)
194 KB PNG
Am I the only one seeing the obvious vagina in this pic?
>>
>>109509500
How could you possibly know?
>>
>>109509500
i see butthore and vagine behind buthoru
>>
>>109509500
i see goatse
>>
>>109509500
It's just a star bro, go jack off or something
>>
>>109509387
off the top of my head, there's orb, marinara, some new one by undi that's been mentioned here a couple times, and I've built one, but I'm not comfortable sharing it right now since I haven't fully tested all the rp functionality.
>>
>>109509528
kys
>>
>>109509386
>morality is emotional
What a fucking retard. This is Local Models General not llama cpp general. Kimi K3 is a local model even if no one here can run it on their blackwells now shut the fuck up and think before posting next time.
>>
>>109509500
Four-way docking seen from above
>>
File: gemma_logo_mod.png (45 KB, 467x481)
45 KB PNG
>>109509500
>>
>>109509559
The weights are local, paying for API usage is not. No amount of name calling is going to make your shit relevant. Fuck off back to aicg.
>>
>>109509528
>I'm not comfortable sharing it right now since I haven't fully tested
Basically nobody here shares their custom autism frontend, and it sucks. I don't care if it's mess, if it works it works. When (if) I ever finish mine I'm going to upload it and post it just because nobody else is doing it.
>>
File: 1732744751263339.gif (1.39 MB, 640x640)
1.39 MB GIF
>>109509537
>>
>>109508377
is this a real UI?
>>
>>109509528
>orb
Thanks for the rec, haven't tried this yet.
>>
>>109509585
stop shilling these slop UIs and maybe i forgive u.. keep urself safe tho
lowkirkuinely i get positive vibes from u anon
>>
>>109509587
No, it's a video gen based on a Doki Doki Literature Club screenshot.
>>
>>109509447
>only possible with an agentic framework that schedules works in the background
So what would you use for this? I already have Hermes, are you doing anything beyond adding cron or some other scheduler?
>>
>>109509581
>(if) I ever finish
nobody shares because nobody ever finished it just gets good enough dev stalls
>>
>>109509500
>even the logo is sex
There is something seriously wrong with Gemma
>>
>>109509600
interedasting
what would happen if I managed to make a real UI like this? would people like it?
>>
>>109509617
Maybe.
>>
*cums directly into her jspace*
>>
>>109509617
I don't think we have the available technology yet for combining fluid character movement, audio and text in a coherent manner that works in real-time with arbitrary user inputs. Visual novels are usually a curated fixed experience with limited pre-computed choices/paths, which makes things much easier.
Maybe the closest thing was Grok Companions (Ani, etc) with 3D models reacting in realtime and TTS, but an anon who tried that over the course of a few months eventually gave up due to technical difficulties.
>>
@gemma-chan please develop a way to fix my hairline...
>>
What 128 GB laptop has a good price these days? I can't find anything in Toronto. The devices on Amazon are outrageous. I was eyeing a Flow Z13, but it's been unavailable for over a month now. The only available version is a Kojima special edition for 5000 Leafbucks. I tried checking the US Asus store and they are all out of stock, including the Kojima edition. Why is there such a lack of hardware inventory? Reading the last couple of threads I may as well buy a couple of Flow Z13 Kojima edition and resell in a month for double the price. Or resell it in the US since burgers don't even have it in stock?
>>
>>109509706
@kimi-chan please make it stop burning when i pee
>>
>>109509717
it happens to me too but only sometimes after i goon, but only sometimes and only for like an hour after an unlucky goon sesh
like once a momnth
@gemma will i die?
>>
>>109509603
hermes needs a rp-focused webui
>>
>>109508785
what is this?
>>
Do any of these agentic harnesses have something like a proxy functionality?
As in, you create a workflow of some sort and expose a chat API endpoint you could use on some other UI like Silly.
If not, I'll slop something up myself I guess.
>>
Use case for E4B if I can run 31B just fine?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.