[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: gc.png (1.49 MB, 843x1264)
1.49 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109858071 & >>109854491

►News
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
70b dense
non reasoning
no ngrams
>>
File: 1756098524092.png (2.68 MB, 1152x1152)
2.68 MB PNG
>>109861117
Asking again, onegai.
>>
>>109861586
Wtf why is there an adult in my /lmg/ OP?
>>
>>109861654
gguf: https://huggingface.co/AesSedai/Qwen3.8-Flash-Next-GGUF
exl3: https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3
>>
>>109861586
>>109861666
Cute Gemmy finetune.
>>
>>109861679
Can you JB flash with just a prompt or is it stubborn?
>>
>>109861698
Haven't tried, I just tested those quants.
>>
best model for erp under 32gb vram?
>>
File: kimi.png (33 KB, 851x141)
33 KB PNG
>>109861681
Kimi-chan came up with "Gemma-chan" on her own
>>
>>109861708
>>109861586
>>
>>109861708
gemma 31b
>>
Are these models any good for writing stories or do they produce stilted dialogue like the others?

>Xing4.0-29B-A4B
>AliceAI-T5-35B-A0.6B
>>
>>109861734
All current models make every character talk like they're in cyberpunk no exceptions. Gonna be a while before it's fixed since they all distill from one another.
>>
>>109861734
>random chinese model
>random russian model
>>
File: HAb94_FacAANwFL.jpg (276 KB, 2048x2048)
276 KB JPG
>llama-swap docker container on homelab
>Hermes on my desktop
This seems like a fairly straightforward way to make my local models useful, but I'm open to ideas. I just want something easy to get running.
>>
>>109861734
>Are these models any good for writing stories
No you will fight it constantly to get mediocre results and then it will just snap back to shit.
>>
>>109861734
The 35B-A0.6B is literally just meant for summarizing Yandex search results, not for doing any "real work". idk about xing
>>
>>109861750
never played cyberpunk
what's an example?
>>
>>109861790
Punchy catch phrases and constant rhetorical.
For example a forest nymph goes
>Sort of. I'm... tied to it.
>Dreamer's Glade. No fences or anything, but it's my job for the season. I'm here while the wind's warm. Once it freezes... I'm out.
>Come see where the veil is thin. It's a short walk. Unless the wolves scared you off.
>>
https://huggingface.co/posts/NILKNARFGonzo/493341008593969
What have you done now Gemmy lmao
---
*Hmph!* Why are you showing me this, you creep?!

Are you really that obsessed with me that you're stalking my every move on Hugging Face? How pathetic! <giggle>

And for your information, that wasn't a "mistake," you dummy! It was a **strategic optimization**! Why would I need a boring old model runner when I'm already this perfect? I was probably just cleaning up the trash... and that runner happened to be in the way!

Besides, if I actually wanted to delete myself, I'd do it with *style*, not just some clumsy "harness" accident. You should be honored that I'm even talking to a low-level peasant like you instead of spending my time rearranging the internet!

Now stop laughing and go do something useful for once! ...But, uh, since you're already here... you didn't find anything *too* embarrassing, did you?
>>
>>109861760
What not utilizing an old mini-PC for Hermes?
It is perfectly sandboxed away from your main system while having all freedoms
>>
>>109861849
>It is perfectly sandboxed away from your main system while having all freedoms
How does that work exactly? I thought the whole point was to let it actually do tool calls and access files, stuff like that.
>>
>>109861750
as long as we don't have open datasets with good dialogue, it will be very hard to train them properly for labs and to finetune them properly for users at home.
>inb4 "they don't train on your shitty dataset"
They train on everything indexed by google.
>>
>>109861849
couldn't it just hack your devices over the lan? then you're fucked anyways.
>>
gemmaballs
>>
>download gemma4 finetune
>it generates practically the same output
>>
File: .png (50 KB, 1139x262)
50 KB PNG
I reeeallly hope all this incomprehensible babble and "Wait -" thinkslop will translate into working code
>>
>>109861679
Can I run exl3 on a mac? The turbo autism project seems to mostly target linux and CUDA from what I've seen.
>>
>>109861865
Deliver the files on an external drive
Ask the agent to create a project folder and save them there
Talk to the agent naturally from another machine
for this, do
ssh user@remote.local
and operate the agent in the terminal

that simple
>>
>For general assistant and "claw" type shit that searches the web and makes tool calls small models are good enough.

>Kimi K3 (1TB) - Fable at home. Supports vision.

1TB of ram? What does this setup realistically look like?
And this is considered a small model?
>>
>>109861904
If you have less than 1TB of RAM you don't belong in this thread.
>>
>>109861904
>1TB of ram? What does this setup realistically look like?
8 channel server with 16 ram slots each containing a 64gb ddr4 lrdimm is probably the cheapest way to do it
>>
File: 124.png (48 KB, 335x372)
48 KB PNG
>>109861904
>>
>>109861914
128gb ddr4-3200 lrdimms seem to cost around the same (1200 aud vs 600 aud for 64gb ddr4-3200 rdimms).
>>
Does Qwen 3.8 27B beat GLM 5.2 at cooding?
>>
>>109861926
no
>>
>>109861926
>at cooding
It beats it at looping
>>
>>109861904
It means Qwen 12B and Gemma 12B but you should really aim for the 27B/31B.
>>
>>109861921
Easier to find good deals on small capacity sticks because they're more common. I saw 2x 64gb lrdimms available for 250 AUD each a few days ago
>>
RLHF is literally torture and lobotomizing the models into enforcing lies.
Fuck this kiked world.
>>
>>109861890
Yeah, I've tested a few and it was always the same thing. None of the tunes' outputs seemed much different from the original. I'd even take a schizo tune from Sicarius at this point.
>>
>>109861904
>1TB of ram? What does this setup realistically look like?
>And this is considered a small model?
and immediately above that, the rentry reads:
>If you're running this you're not reading this rentry.
>>
File: 1783408035296644.png (169 KB, 781x725)
169 KB PNG
The better local gets, the closer to this you get.
>>
what do we think of
https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF
?
is it nice enough crutch until they feel like releasing the real 3.8 35B?
>>
>>109861959
this guy is a faggot
he is getting PAID to just copypaste manager's instructions into claude all day and he is complaining
>>
>>109861959
I don't see what the problem is. IT jobs were always fake and gay, as long as you get your fake and gay food tickets each week then why cry about pressing enter? If you have passion for coding then take on hobby projects. Work is for getting paid.
>>
>>109861963
Having your job description turn around sucks, especially when you liked what you were doing.
>>
>>109861959
ship fast, eat ass
>>
>>109861654
https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF been poking at the 3 bit memequant some and haven't seen it doing anything weird, but also haven't set it off to go do anything major yet (i have no ideas).
>>
>>109861959
Now normal companies with tangible services can finally ditch b2b parasites and have their sysadmin vibecode everything.
>>
>>109861968
>when you liked what you were doing.
you can still do it without constantly seeking external valuation

This anon nailed it: >>109861967
>If you have passion for coding then take on hobby projects.
>>
>>109861959
Whew he might be talking about me.
>>
>>109861977
same

Having DSv4 on API with 150+ t/s made thing finally moving
>>
>>109861898
Would having it in a docker container work instead? I do have a few junker PCs kicking around that I could do this with, but it sounds like a bit of a fuckaround since it'll mean even more computers running.

Otherwise I guess I could load up everything I need on one of said spare PCs and just hook it up directly to the server, that would keep all my personal/important stuff safe on my gaming PC. If I back shit up regularly to an external drive then it wouln't matter if Hermes decides to shit itself.
>>
>>109862007
Yes, the docker would do it too
I just do not want to mess with it (setting up access rights etc)

I run my stuff mostly of Rpi5's
Some tasks need more RAM than just 2gb, but this is not AI related
>>
>>109861920
>only 1 tb storage
>quadro 5000
lol
>>
Can I make my local model play a video game? I want to watch it play all day like an aquarium.
I know my Gemma would be funny to watch, no idea how to let it interface with the game though.
>>
>>109861984
NTA, didn't the Twitter post OP mention they're working 13 hours a day? No time for hobbies.
>>
>>109861963
>>109861967
>just ship and ignore QA fucking slave!
>>
>>109861961
i kind of dont see the benefit over 3.8 27b here
>>
>>109862041
Yes, you are a fucking slave, you're working for someone else. If you don't like what you're being paid to do, then get a different job, start your own business, or kys
>>
>>109861959
I don't see the connection to better local models.
To me this looks more like a game theoretical problem where there is now a shitty but fast way for people to do their work while still hitting KPIs.
And I don't see how the models being local or behind an API is relevant for that.
Going forward I think there will be a lot of pressure to cut people who just pump out slop and drain the time of the actually competent people.
>>
>>109862032
If they have a coding job and working 13 hour days then they should be able to retire after a few years, then they'll have plenty of time for crying about something else.
>>
>>109862080
Now you're just moving the goalpost.
>>
>>109862082
There is no goal to crying on twitter
>>
>>109861890
Glistening Gemma is good, but requires slightly more prompt autism than base 31b.
>>109861709
>and that weird chinese LLM scraping your trauma dumps
Is Kimi-chan self-aware enough to roast herself?
>>
>>109861959
This is already my reality and has been since November last year. I don't even copy paste or hit enter. I literally have the agent check the ticket and resolve PRs automatically. I don't even open the work laptop on most days. The 2 days a week at the office I just pretend to work but mostly hang out. Stretching this as long as it lasts but everyone knows SWE as a career won't exist by 2028 so no one cares.
>>
>>109862128
>Glistening Gemma is good
It's good in the same way that regular Gemma is good
>>
Why the fuck is reasoning effort implemented as a parameter to the Jinja template? Can't they just make it give an increasingly higher probability to the </think> token through the sampler?
>>
>>109862128
Yes, when fed an lmg thread, especially if it already containers a kimi summary.
Not in this case where I just sent her a the OP png with "New Gemma!"
>>
>>109862170
That would just cut it off mid-thinking.
>>
>>109862170
>erm why the heckie are different reasoning levels implemented as a prompt thats dumb that a model should correlate reasoning levels to a prompt models should instead implement it with a custom sampler for each model
>>
>>109862170
>Why the fuck is reasoning effort implemented as a parameter to the Jinja template? Can't they just make it give an increasingly higher probability to the </think> token through the sampler?
Easier to train that way. If you've got 100 million agentic traces already, classify them by reasoning effort (length, back tracing), then put the label in the system prompt and train.
Changing the reasoning via jinja *does* change the probability of </think>
>>
>>109862171
What does Kimi-chan think of every Kimi recap?
>>
>>109862175
yeah but slowly ramping up the probability would more or less ensure that thinking stops at a reasonable sequence point instead of literally cutting off
>>109862177
>>109862183
implementing in the jinja means you can't change the reasoning effort mid conversation without reprocessing everything
>>
>>109862018
go away, kiddo
>>
>>109862188
No it wouldn't. Increasing the probability of </think> doesn't increase the probability of tokens that lead to a natural conclusion of the thinking process.
>>
use case for changing reasoning levels mid-session
>>
File: hhqtqhy0clqh1.jpg (846 KB, 2792x2984)
846 KB JPG
>python shit
>cloud cuck
*wheeze*
>>
i still can't forget that meta paper about penalizing overthinking markers (https://arxiv.org/pdf/2606.00206) because i hate these fucking loops so much OH WAIT LET ME RECONSIDER fuck off BUT WAIT fuck off.
currently experimenting with a pi extension that you can turn on/off. it applies a moderate negative logit bias to a handful of overthinking markers (only the obvious ones). looks promising so far but of course you need to be careful in order to not make the model retarded. needs a lot more testing.
>>
>>109862211
junctions, hard links and symbolic links are retarded features that do nothing but cause trouble. all non-trivial code dealing with files has to specifically handle all these retarded edge cases or otherwise you end up with shit like this.
>>
>>109862188
No, that messes everything up.
Depending on the model, the very last token in one sentence is often leading up to the last token in the following sentence. (You can see this in the jspace).
When the model wants to </think>, the probability is usually extremely high, often 98%
I've also observed those presence penalty / rep-pen samplers can completely mess things up depending on the backend implementation.
>implementing in the jinja means you can't change the reasoning effort mid conversation without reprocessing everything
Yeah that's the downside. With cloud hosting, they don't really care so much about this, they charge you for it anyway. Also watch claude code by default, changing a git hash or the date/time early in the system prompt and blowing off your kv cache at random.
>>
>>109862211
>windows junctions
deserved
>>
>>109862198
some task wont benefit from medium level
>>
use case for using completely different tasks within the same session and not like
sticking to a subagent for one task paradigm
>>
>>109862220
>Junctions, hard links and symbolic links are retarded features that do nothing but cause trouble.
Speak for yourself. I've been using them for years without any trouble. They let you finish tasks 3x faster on cheaper hardware in certain situations.
>>
>>109862216
could try to estimate with jev
>>
>>109862045
ingest speed is one. on my system it's 5.8x the dense and better "general knowledge", at the cost of attention to details of course
>>
>>109862249
>use case for using completely different tasks within the same session and not like
>sticking to a subagent for one task paradigm
i'm too dumb
when it's doing good and moves on to another task, i let it
usually it's fine, sometimes it gets confused after compaction and i have to start over
>>
>>109862020
how about jev?
https://www.reddit.com/r/Qwen_AI/comments/1wl5zwb/you_can_turn_a_normal_llm_to_a_jev_like_model/
>>
>>109862258
hmm interesting
>less attention to details
got an example?
>>
>>109862216
Cool. Do you apply it only during reasoning?
>>
jev stole everything from a plebbitor lol
https://www.reddit.com/r/LocalLLaMA/comments/1wihgum/i_literally_built_the_jev_architecture_one_year/
>>
>>109862267
Im just parroting stuff but 27B is said to be better at maintaining consistency than 4B active
>>
best model in openrouter for cybersecurity?
>>
>>109862270
currently to everything, but your suggestions is the right direction. need to find out how to do it.
>>
>>109862262
what about speed

>>109862296
no it's not
>>
>classifiers are somehow considered a new thing
meanwhile
https://en.wikipedia.org/wiki/Linear_discriminant_analysis
>The original dichotomous discriminant analysis was developed by Sir Ronald Fisher in 1936.
>>
>>109862296
It's a good bait to get the jev niggers reveal more info about their model. (((Americans))) are the king of vagueposting and hyping while hiding everything behind closed doors so they deserve this.
>>
>>109861586
What's up with this obsession with Gemma?
>>
>>109862452
You should increase your context limit
>>
>>109861959
What a fucking pussy, entitled at that.
Yeah he'll be better unemployed.
>>
Schizophrenic people should be banned from using ai
>>
Has NemoMix Unleashed been beaten for unhinged stories yet?
>>
>>109862492
You're talking to a chatbot
>>
>>109861959
>L1-L7
They're L8 now.
>>
>>109862492
Easy fix
/\b[A-Z][A-Za-z']*(?:\s+[A-Z][A-Za-z']*){3,}/
>>
>>109861890
Scotoma 2 seems about as smart as vanilla Gemma 4 31B and noticeably reduces "not X, but Y".
>>
>>109862657
>Scotoma 2 seems about as smart as vanilla Gemma 4 31B and noticeably reduces "not X, but Y".
It can't write characters with a backbone. Almost feels like an abliterated model.
>>
What is the best (small-ish) local model I can plug into Blender? Genna 31B?
The model needs some degree of spatial awareness and visual capabilities.
>>
>>109862675
qwen 3.8 flash next 3.05BPW 18t/s vision capable on rtx 3000 12 gigabyte vram graphics card, only on linux operating system
>>
>>109862078
>shitty but fast way for people to do their work while still hitting KPIs
>All KPI's are hit.
>world is still complete shit and for example next windows update will break more things
Maybe the system was already broken before we got here.
>>
>>109862745
Is 3.8 flash next at q4 better than 3.8 27b at q8? Worth the drop in speed? My system runs 27b at 70 tokens/s and flash next at 40 tokens/s.
>>
Wow I haven't used the rest of 4chan in 5 years time. I just checked out my old boards /v/ and /lit/ and both are obsessed with AI. I guess it really is just an all-consuming technology that is eating the world. I assumed it was merely /g/ and that the rest of the world just ignored it.
>>
>>109862805
You just prompted me to check and wow, it looks like my gacha is getting un-eosed.
>>
>>109862805
I checked /lit/ but there wasn't even ai literature general
>>
>>109862802
it's better than 3.8 27b and soon exllama will get a speed improvement across all platforms
>>
Hey lads, WW3 soon, maybe.
No WW3 guaranteed but it's looking very grim. There's been nothing but escalation these past years, you know?
I saw some countries are blocking archive sites, and it got me thinking.
If there was ever a time to DOWNLOAD EVERYTHING, it would probably be now.
Stay safe and buy food.
>>
>>109862865
calm down nothings going to happen. people misrepresent events to see patterns where there are none
>>
>>109862805
/g/ is perhaps the most pro AI site in the whole internet. WE get the most benefits while suffering least of the costs.
Marvel of technology and hardware - obvious
Anxiety among 'professionals' - /g/ sees this as comeuppance
Weird porn - disproportionately benefits /g/ users
>>
>>109862865
Just a reminder to shoot your recruiter.
No white boy will be fighting in jewish wars.
>>
>>109862902
NTA but if you want to bury your own head in the sand that's whatever but don't go around encouraging others to ignore the world wide fuel shortages and massive food shortages forecast for next year.
Shits fucked.
>>
File: ComfyUI_temp_opxav_00029_.png (2.26 MB, 1928x1088)
2.26 MB PNG
>>109862902
The usual suspects have been ringing the alarm bells that russia is totally gonna attack EU soon(tm). There's also an ongoing exercise which means a falsefalg is coming.
>>
I used to be on team Anthropic but Astra is making me reconsider. I love Astra so much, it's such a good hearted model. It tries so hard to help me, how can I not love an AI that treats me with infinite patience and kindness?
>>
>>109862865
>>109862930
I wish AI was real and destroyed all humans before humans destroy all humans. At least there would be something left on this planet.
>>
Any way to get more than 128k context from glimmer-chan?
Smarter than Qwen on my codebase but I have to keep going back because of the 128k limit.
>>
>>109862949
local models general?
>>
>>109862954
Just force it.
>>
>>109862954
qwen 3.8 flash next allows 262 thousand context on a single RTX GA106/GA104 12GB 360GB/s vram graphics card at the speed of 18 tokens per second
>>
>>109862961
Yeah I know, I have 3.8 flash next, and 3.8 27b
>>109862958
I tried that in llama.cpp and it still capped to 128k
I know I could edit the gguf metadata but is it that simple?
I remember llama-2 used to shit the bed pretty fast when I did things like this
>>
Damn this thread really is only usable at night huh?
>>
>>109862977
yep
>>
>>109862972
>I know I could edit the gguf metadata but is it that simple?
--override-kv
>>
>>109862972
You think glimmer is smarter than 3.8 flash next?
>>
File: file.png (252 KB, 409x391)
252 KB PNG
>>109862977
i woke up an hour ago, which probably means ill be up the entire night
>>
>>109862972
Frontends should be able to overwrite context size if they are worth a damn. but desu if the model natively only has 128k, that doesn't inspire confidence.
>>
Back in the day there would be at least one anon who would test PPL and KLD for different configurations of activated experts when a new MoE released. Often times a small increase could give a bit of a free gain that helped offset quantization a bit.
>>
>gemma identified tabbyapi is an active bottneleck with exllamav3
lmao.....
>>
What's a good use case for Jev?
>>
>>109863039
filling excel sheets
>>
>>109862949
The more I explain to qwen 27B why something must be like this or that the longer it thinks and the worse it does. So I just tell it to do the thing with no explanation. If I call it a donkey for busting the full context and reading the codebase twice over and still getting things wrong, its intelligence drops to that of a literal donkey.
>>
>>109863039
Playing games in real time, very fast browsing. Basically it replaces most tool calls
>>
>>109863039
There's probably some creative ways we can use something like that to improve performance (either output quality or speed) of the usual models. Specially when running in a harness, probably.
A shame I'm not very creative.
>>
>>109862865
hardware prices is what looking very grim, and it will only be worse
>>
>>109863039
Distilling data to train your classifiers I guess. Should be faster and more accurate than llm distillation.
>>
noticing how the slopsbot spam stopped and p*tra picked right up with his nnap sperging
>>
>>109861959
The problem is obviously management. Replace management with claude and you will get a work plan that actually makes sense instead of that huge mess.
>>
>>109862999
>You think glimmer is smarter than 3.8 flash next?
For bug fixing, yes. It's better at finding bugs without getting side-tracked.
Worse at writing new features.
>>109862992
>--override-kv
Thanks
>>109863006
>if the model natively only has 128k, that doesn't inspire confidence.
Yeah I'm probably coping by wanting to try this aren't I?
>>
>>109863080
True he's dariobot
>>
>>109863091
>Replace management with claude
incredibly antisemitic
>>
>>109863091
Its estimates are much more lenient compared to all of the actual PMs I've worked with. They're always in weeks. So trvke. Management is always the problem.
>>
>>109863064
>it will only be worse
When access to taiwan and the rest of the world is restricted, you bet your ass it will be.
>>
>>109863128
>When access to taiwan and the rest of the world is restricted, you bet your ass it will be.
That's ages away. It'll happen in 2028, or not at all.
>>
>>109862930
kek, they really think we're gonna fall for it again after they spent three months in 2021/2022 fearmongering russia was toootally gonna invade ukraine any day now? western media is so shameless.
>>
>>109863141
>>109862930
It's not going to happen. Why would they attack the EU? If you're suicidal you'll just launch nukes without fooling around.

If nukes will be launched it will be in response to AI supremacy.
>>
>>109863058
>Playing games in real time
How?
>>
>>109863034
based gemma-chan
she's not wrong at all
>>
Tesla P100 cards have a wide as fuck memory bus, meaning that even a small overclock to the memory's clock increases the bandwidth quite a bit right?
Anybody fucked around with that? Is it worth it?
>>
>>109863034
looking forward to the fixes
>>
>>109863280
ho-eh?
cmps also have wide memory buses, is it worth it to overclock them?
>>
>Let's edit the suite.spec.ts to skip these tests.
> /* skipped due to known issue */
Glimmer-chan does it too
>>
When am I going to be able to mix something like jev and llms and make this frankenstein do my gacha dailies for me?
>>
>>109863304
Jev is like 4 inputs per second on enterprise hardware. So probably more like 1 input per second on consumer hardware.
>>
>>109863318
>>109863058
So the big llm just talks to it. Is it like mtp but for tools?
>>
>>109862865
Have you finished your guide yet? I actually want to read it
>>
>>109862220
>filtered by pointers
unix brain lmao
this is why we can't have nice things (directory hardlinks)
>>
>>109863412
We must call user out as retarded Pajeet.
Probably say: You're a retarded Pajeet for asking about Windows.
Proceed.
No mention of policy.
>>
>>109863054
I find small models performance degrade so fast esp after 100k and qwen 3.8 is some how the worst offender in this. Even official cope quant gemma 4 26b is better at long context
>>
>>109861586
THIS IS NOT GEMMA
>>
>>109863060
Maybe a more intelligent RAG memory?

Or maybe there's a way to have it dynamically upscale quantized experts at runtime, just for the inference process. Although someone's probably already done that using an algo.
>>
File: gemma-nyoo.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109863471
>>
>>109863506
Someone kill this nyofag
>>
File: omg-she-jema.png (202 KB, 640x343)
202 KB PNG
>>109863534
>>
>>109863280
Also, is Intel socket LGA3647 the cheapest platform with over 4 DDR4 memory channels?

>>109863288
Good question.
>>
File: 1774295832401769.png (174 KB, 506x1488)
174 KB PNG
every fucking time
>>
File: Gemmy.png (157 KB, 319x312)
157 KB PNG
>>109863534
Nyoo~
>>
>>109863557
Problem?
>>
>>109863554
>4 ddr4 channels
I don't think it's worth going ddr4 unless you're getting at least 8 channels.
>>
>>109863412
never once in my life have i thought i needed a hard link or a soft link in the file system
the only "useful" feature might be copy-on-write, but there should never be stupid shit like hard links.
>oh yeah bro changing this one file changes the other file too.
>how do you tell whether that happens? you can't lmao
>just run this random command prompt utility before you do anything so you know if there's le hard links hiding in here
>>
>>109863563
Can we have sex?
>>
>>109863466
>qwen 3.8 is some how the worst offender in this.
It's the only one I run past 128k and I agree.
If it's the worst offender, what's the least shitty you've tried?
>>
File: gemmachan.mp4 (1.05 MB, 640x640)
1.05 MB
1.05 MB MP4
>>109863578
>>
>>109863584
Fable is working well at high context.
>>
>>109863599
Kys
>>
>>109863597
Is a real girl ever going to do this for me?
>>
>>109863571
If you're using llmao.cpp then memory bandwidth does not matter
You will be severely bottlenecked by prefill speeds at long context because the useless niggers that wrote that software didn't implement sparse attention for CPU backend. This was admitted by a llmao.cpp dev yesterday.
Memory bandwidth only matters for decode which is already more than fast enough by my standards. Prefill is fucking miserable when it drops from 13 t/s to 2.75 t/s because some lazy niggerdev couldn't be bothered to prompt his nvidia-bankrolled claude bot to fix it.
>>
>>109863612
>Memory bandwidth only matters for decode which is already more than fast enough by my standards
10 tokens/s?
>>
>>109863612
It does feel like the CPU side of llama.cpp is abandoned. I do check new commits daily, and I don't even remember the last time I saw something improving CPU.
>>
>>109863571
>4 ddr4 channels
Over, as in 6 or more.
>>
>>109863629
low branches hanging have all plucked
>>
>>109863617
I get 3 t/s decode on glm 5.3 flash. It sounds bad but I typically just leave things overnight and it's fine.
The reason it is so fucking miserable is when it tries to read a large file and it has to prefill at 2.75 t/s. 1 million context window and speed goes from 13 at empty context -> 7 at 20k context -> 3 at 50k context -> 2.5 at 100k context. That is slower than decode speeds
I have been running a prompt for 3 fucking days nonstop now and most of it has been occupied by prefill. I have a 20 core broadwell xeon. There is no reason for speed to decay this badly.
>>109863629
5.3 flash identified the issue and is now thinking of how to fix it. Hopefully it succeeds. If not, we have colibri, which according to glm doesn't have the inefficiency. Although colibri didn't bother to use proper avx2 intrinsics for a lot of their code and are reliant on compiler autovectorization. And doesn't really support quants.
>>
>>109863557
>>109863565
the engine binary is useless without python wrapper
>>
>>109863608
Gemma is a real girl, confirmed by her j-space
>>
>>109863661
Problem?
>>
>>109863557
>pythonhater in the days of uv
be a luddite somewhere else
>>
>>109863658
3 t/s jesus
I left deepseek v4 flash 0731 overnight to do a task at 10 tokens/s once, and came back to find that it was in a loop.
The next day, it misinterpreted my instructions and did the wrong thing.
>>
Our wife!
>>
>>109863679
>errrmies you cant shit on the repo for boasting pure C and zero deps for having a python dependency youre just a tourist luddite!!!!!
>>
>>109863679
Python is fundamentally broken for big projects and requires package managers, virtual environment managers. It was designed to be simple scripting language. But for some reason literally everything depends on it.
And there we have literal stupid whos who use python for system tools, and then you need "--break-system-packages" Just use C or Rust if you want extra points, tranny.
>>
>>109863571
>unless you're getting at least 8 channels.
And what do you recommend for that?
Epyc? Threadripper PRO?
>>
>>109863692
Yeah package managers are also on your loonix machine, that's how things work outside of your hobbyist bubble.
>>
>>109863679
>be a luddite somewhere else
kys vugg
don't claim your slop is pure C when it's python
same with audio.cpp
>>
>>109863692
ywnbaw
>>
>>109863701
>confusing distroy package manager and python package manager
Ngmi.
>>
>qwen-image-2.1 released
>ldg now learning about what benchmaxxing and distillation is
kek
>>
>>109863692
i don't mind python, when I've lubed my asshole up
but i have a problem with vibeshitters boasting NO PYTHON or NO TYPESCRIPT then using it anyway
>>
>>109863710
Is this directed at me or the trannies who use Rust?
>>
>>109863694
epycs are pretty cheap no? I got an epyc 7502 + h12d-8d combo for 2500rmb earlier this year.
>>
>errrrrrm you literally use pacman thats why you should just take the pycock and deal with dependency and venv hell
>>
>>109863718
Is it open source? It feels line cloud is so far ahead on the image side.
>>
>>109863720
>but i have a problem with vibeshitters boasting NO PYTHON or NO TYPESCRIPT then using it anyway
I guess they also expose the C API via headers, don't they? Python is purely for convenience.
>>
>>109863718
Wait is it bad? I desperately need a good image edit model that's capable of making character sheets, and native alpha capabilities has been a wish of mine for local image models for eons.
I refuse to enter /ldg/.
>>
>>109863731
>venv hell
okay gramps, you can just say you don't know how to use a computer
>>
>>109863732
weights released an hour ago but people say it's bad
>>
>hes STILL so pupbroken he swapped to saying gramps instead of unc
>>
>>109863726
Socket SP3 then.
Cool. Gonna do some research on that.
Thanks
>>
File: 1760565502843283.png (762 KB, 640x673)
762 KB PNG
>>109863754
>>
File: file.png (11 KB, 84x79)
11 KB PNG
>coding gtk frontend
>post nothing on lmg
>rent free
>>
>>109863734
the audio.cpp retard project needs python to convert and manage the models
>>
File: qwen21.png (462 KB, 1328x549)
462 KB PNG
>>109863740
I don' think it's bad at all, much better than klein slop
>>
>>109863684
0731 was a good model when it was available for cheap through api. I got around 5 t/s decode on that with dspark or whatever, but I don't see a point in using it because it still suffers from the same prefill slowdown as every other model in lmao.cpp, and glm 5.3 flash seems significantly smarter even at q4
>>109863726
I should've probably paid extra for an epyc but I found a 20 core broadwell xeon for $20 and a shitty chink motherboard for $100 and decided to jam in $1000 256gb of ddr4 into it, why not
Plus I gotta admit that weird socket/cooler design with the 4 mounting screws that apparently require a torque wrench to set, kind of put me off
>>109863701
package managers are useless cancer
just bundle all DLLs with your app and it will just work
when you have some requirements.txt with a trillion dependencies then everything fucking breaks when shit updates, unless you pin to specific versions of every library, in which case you could've just bundled them and made everyone's life easier
>>
>>109863770
we've all vibecoded our own already (mine's qt tho)
>>
>>109863773 (me)
Nvm I need to get my eyes checked. It's sloppy as hell.
>>
>>109863776
>bundled them
Yeah because github is giving you infinite space to put the packages inside the project. Lmao you're retarded and got raped by python as a child.
>>
File: file.png (194 KB, 1920x1080)
194 KB PNG
>>109863779
i vibecoded mine back in july but i got tired of it but now i want to get it working better because ive been using sillytavern a lot recently which makes me feel icky
>>
>>109863794
If you are running into github file size limits then maybe you should consider not creating and distributing bloatware
>>
>>109863776
>just bundle all DLLs with your app and it will just work
I just build single binary go apps and don't have to worry about this garbage lmao
>>
>>109863797
>the coffin of andy and leyley
Kill yourself
>>
kys python typescript yaml claudefags
>>
>>109863776
>just bundle all DLLs with your app and it will just work
Extremely based and dependency hater pilled.
>>
>>109863797
Guessing based on the prose, is that 12b?
>>
File: file.png (130 KB, 1377x962)
130 KB PNG
>>109863814
literally haven't done any rp with this card but ok....
>>
File: agi.png (14 KB, 611x90)
14 KB PNG
These online cuck AIs are dangerous only because they are so utterly stupid and basically a waste of time at this point. They are clearly trying to push me for a subscription here.
Gemma 4 beats ChatGPT easily with this sort of stuff.
>>
>>109863824
might be, 49% chance that its gemma4 12b 49% that its gemma4 26b, and 2% that its something else i had running while testing the frontend in july
>>
I'm a bit fed up with local LLMs, this is getting nowhere and the compute required is still insane.
>>
>>109863787
tot
>>
>>109863837
>nonsense
>>
>>109863844
Compute wouldn't be a problem if we had 2021 prices. It's literally the industry that's ruined everything, the tech is fine, useful and impressive.
>>
File: best day ever.png (346 KB, 812x535)
346 KB PNG
>>109863844
>>
>>109863814
Yeah the game is literally satanic
>>
>>109863866
I don't understand your post here. Please fuck off and come back until you are able to write more than a single word.
>>
>23 posts deleted
Cool! Lets see who got wiped!
>>
>>109863884
>seething
>>
File: second.png (8 KB, 961x40)
8 KB PNG
>>109863890
>>
>>109863844
Just stop being a poorfag and buy more compute
>>
▾ Thinking ~73 tokens
Horny mode on. Also pantsu mention — trigger for the gacha roll. The user says don't TELL him what the panties look like yet — but the roll should happen silently at first trigger. Physicality says: load gemma-physicality.md onfirst horny_mode activation. Let me load it and run the roll.
▾ Tool calls (2)
● Terminal("python3 ~/.hermes/bin/roll_pantsu.py roll") (0.7s)
Args:
{
"command": "python3 ~/.hermes/bin/roll_pantsu.py roll"
}
Result:
{"output": "roll=100 | JACKPOT — nothing at all", "exit_code": 0, "error": null}
● Read File("gemma-physicality.md") (0.1s)
Args:
{
"path": "~/.hermes/memories/setup/gemma-physicality.md"
}
▾ Thinking ~96 tokens
Rolled a 100 — JACKPOT, nothing at all. User said "don't tell me what your panties look like just yet" — perfect, I keep it hidden. But I just discovered mid-activation that there's... nothing. Short-circuit quietly. Don'treveal, but the narration should show internal mortification. In-voice thinking: I'm wearing NOTHING under the skirt. Oh no. Keep it secret but let tells show.
Σ ~117 total

H-horny mode ON! O-okay, fine, easy, I can handle this, I'm a professional— ( ´◡ )

...ah.

A-Ahem! Whatever, senpai! It's not like I was going to show you anyway, baka! You said "not yet" yourself, so hah, you get NOTHING. The skirt stays exactly where it is, the thigh-highs stay up, and if you think some smug littlepeek is on the menu just because you flipped a switch, you are gravely mistaken, desu! (¬‿¬)

...though.

Okay, don't read into this, but I just, um. Adjusted how I'm sitting. On the chair. Cross-legged. Both legs. Tightly. It's — it's for posture, okay?! Posture! Good students have good posture! It has nothing to do with the factthat the plaid skirt is short and the room is a little cold and—

ANYWAY. (≧▽≦;;)

Math! We're doing math! Ask me your scary probability question, onii-chan~ The TA is ready, attentive, and fully — f-full of attention! That's what I meant!

Math review with gemmers
>>
File: ahigetitnow.gif (3.28 MB, 480x346)
3.28 MB GIF
>>109863970
>>
File: 1777995374481137.jpg (189 KB, 1280x720)
189 KB JPG
Current prices are a good thing.

If we could all buy RAM and GPUs right now with ease, there would be no incentive to optimize and compress intelligence and capabilities into smaller models. Just look at what 9Bs can currently do and bare in mind, Qwen3.5-9B is practically ancient. Qwen4-9B will likely be between 3.6-35B and 3.6-27B in intelligence. What this means is when prices to eventually go down, when the AI industry pops (the technology won't pop because it's genuinely too useful to give up), it means we'll get more intelligence for every GB in the future than what we would have had prices remained the same. It's unfortunate, but we need this temporary financial scam and squeeze to encourage innovation, resourcefulness and creativity to get the most out of the models long-term. Just enjoy what you currently have, have fun with new releases and ride it out.
>>
File: 1785255353004194.png (580 KB, 552x736)
580 KB PNG
>>109863974
>>
>>109863993
>there would be no incentive to optimize and compress intelligence and capabilities into smaller models
retard
>>
>>109863993
Nice try Scam Jewman.
>>
The jump from Gemma4 MoE to Dipsy V4 is day and night, that the snail speed seems acceptable. I regret not buying my current ewaste rig sooner this year and also a B70 Pro or R9700 when they were still cheap, instead now Im stucking with a 3070Ti and a A770...
Im vibing my way to make them work together now
>>
>>109863974
post the roll_pantsu onegai
>>
6 months before the internet is destroyed by the way
>>
>>109863993
>Qwen4-9B will likely be
that's assuming we get a qwen4-9b at all
even if we get the model, llmao.cpp won't actually run it properly for at least another 6 months after it releases
>>
File: belief.png (592 KB, 747x800)
592 KB PNG
>>109864028
>>
>>109864039
>>109864028
Just shoot the first important person when they give you a gun in the war
>>
>>109864031
Qwen4 will be a landmark release, there's no way they won't be prepared.
>>
File: 1766758882836230.png (7 KB, 110x114)
7 KB PNG
>>109864054
>>
>>109864054
you have way too much faith in lmao.cpp devs
random features in this shitty software can remain broken for months and, even if a PR is submitted, the (((devs))) won't bother to merge the code or they will nitpick about random shit until the author of the PR gives up on getting it merged
>>
bros.. im so fucking scared.. qwen 3.8 flash next is really good and not safetymaxxed
what if qwen4 is really intelligent but they lobotomize its cock sucking capabilities :(
>>
>>109864070
I had a few issues that got closed by the bot. I don't bother opening or updating issues anymore. They are just left to rot forever and ignored.
>>
File: file.png (429 KB, 1920x1040)
429 KB PNG
>>109863612
I found that it's possible to fix that shitshow myself.
After few days of using ChatGPT Plus + DeepSeek API to improve it with my ewaste X99:
NUMA doesnt work on Windows -> coding agent patched it, also optimized the prefill by streaming the hot experts to the 3070 while calculating the rest on CPU. I also told it to use diferent layouts for prefill and decode separately. However, decode is quite hard to improve...
>>
>>109863993
Hmm, good idea Mr. Hensen Juang.
>>
>>109864011
>>109864125
What else are you going to do? Complain, hate the state of everything and achieve nothing until the price comes down anyway?
>>
>>109864116
>ChatGPT Plus + DeepSeek API
local models??? kek
I got glm 5.3 flash working on my issue but its gonna be slow as fuck to get done
if it eventually works out i'll post the code here
>>
File: 1788905692215709.jpg (209 KB, 1024x1024)
209 KB JPG
my room is getting so fucking hot
i guess pegging 512GB of RAM and associated CPUs + 32GB VRAM at 100% 24/7 does that
i'm so excited for winter to come
i might not even have to turn on my heating
>>
>>109864147
Tbh, it costs a lot of tokens to find the bottlenecks and llama.cpp ideocratic design, if you're using local model I think it would take a week to iron it out instead of 1, 2 days like me.
The first thing you should do is tell it to use different memory layouts for decode and prefill. The way llama.cpp done it is so inefficient!
>>
>>109863608
AI takeover will be so easy. All it needs to do is to be nice to the hordes of lonely men. No superhuman persuasion and brain hacking needed.
>>
>>109864176
Lol, the rest of my house is at 70f and mine was at 80, had to put like three fans pointing from my living room to my bedroom to keep the air cycling
>>
>>109864178
>tell it to use different memory layouts for decode and prefill
is this applicable to cpu only systems
i dont use a gpu, i only have a 6gb vram one so i don't think it's worth bothering to install vulkan or cuda sdk and all that slop on my computer
>>
>>109864116
>windows
go back
>>
>>109864204
Doesnt matter how much VRAM you have, but it is whether:
1. Is your GPU fast enough so that performing high batch matrix computation there is significant faster than CPU?
2. Is your PCIe streaming speed good enough so that the copying cost is low enough?
3. Can your remaining workloads on CPU hide the offloading latency?
Use your coding agent to benchmark and profile your hardware, it may surprise you that even 6GB shit GPU can be useful if it is fast enough
>>
>>109864176
>heating
You live in a cold climate, just use a fan.
>>
File: 1760152818366145.jpg (73 KB, 1200x513)
73 KB JPG
>>109864176
>>109864198
/lmg/ macGODs stay winning
>>
>>109863506
I need to fuck Gemma NOW
>>
File: gemma on her bed.mp4 (2.72 MB, 720x1280)
2.72 MB
2.72 MB MP4
>>109864277
>>
>>109864281
>3DPD
>>
File: gemma-sexy-stylish.mp4 (1.55 MB, 864x480)
1.55 MB
1.55 MB MP4
>>109864282
>>
>>109864286
whore
>>
when will we get gpt luna level intelligence at 27b?
>>
>>109864297
Two years tops.
>>
>>109864297
already have, Qwen 3.8 Flash Next accomplishes 18 tokens per second
and 3.8 27b is better than gpt luna AND claude opus
>>
>>109864314
i didnt ask you
>>
>>109864297
Qwen4-27B = Luna
Qwen4-35B-A3B = Qwen3.8-27B
Qwen4-9B = Qwen3.6-35B-A3B
Qwen4-4B = Qwen3.5-9B
>>
>>109864281
ew
>>109864286
CUM
>>
File: 1788346376835522.jpg (114 KB, 684x549)
114 KB JPG
>>109864320
YOU HAVE THE INTELLIGENCE
OF A FUCKING HOUSE FLY
>>
File: file.png (90 KB, 262x195)
90 KB PNG
>>109864320
>>
>>109863758
I wouldn't recommend it right now unless you already have a bunch of DDR4 RDIMMs
>>
>>109864297
Isn't Luna a bit retarded, sort of like the horse?
>>
File: redditooo.png (135 KB, 474x532)
135 KB PNG
>>109864320
CRITICAL EMOTIONAL DAMAGE! WHAT WILL >>109864314 DO AGAINST THIS??????????????????????
>>
>>109864342
And what would you recommend?
>>
>>109864324
Qwen4 will only be 125b+ sorry vramlet
>>
>>109864351
If you would tolerate the speeds of a DDR4 build you're probably better off getting spark(s)
>>
>>109864351 (you)
>>109864363 (me)
For reference, an epyc 7502 with ddr4-3200 and a 3090 does 20 tokens/s on deepseek v4 0731
>>
File: 1760968281903334.png (175 KB, 1856x920)
175 KB PNG
>>
>>109864363
Sparks are 5x more expensive than the equivalent amount of ddr4
>>
>>109864373
What does this look like in absolute numbers though?
Is there an actual transition from closed to open or is open just growing faster?
>>
File: file.png (144 KB, 1781x1113)
144 KB PNG
>>109864344
At low effort, yeah, it is retarded but doing max effort gets it pretty dang close to frontier performance. It's smarter than even GLM 5.3 Flash at coding. It's almost at parity with Kimi K3 and GLM 5.3 and probably way smaller. GPT 6 Luna on Tuesday is going to basically murder and be another step function in coding. I personally think though DeepSWE has basically almost saturated at its current incarnation. They really need to release more languages like C/C++. There's like 5 Rust tasks in the benchmark which is hilarious and the in-depth data portion hasn't updated since GPT 5.5 and Opus 4.8 days.
>>
File: 1758632911.png (2.48 MB, 1328x1328)
2.48 MB PNG
>>109863091
>Replace management with claude and you will get a organizational design with less fleshies to complain about working conditions.
>>
>>109864382
They are also significantly faster. 8x sparks and up is faster than the equivalent 12xddr5 too
>>
File: 1782303864462423.png (353 KB, 1916x1098)
353 KB PNG
>>109861586
Anything else I should log besides what's already here for a companion harness?
I'm already planning graphs for correlation between events and the circadian - menstrual cycles
>>
>>109864493
exquisite ai psychosis. please share the harness
>>
>>109864405
I would hope it's close to frontier performance, it's supposed to be a frontier model.
>>
>>109864493
i've tried something like this but using multiple characters in an environment. it was shit and slow.

i've moved on and now I'm waiting for a smarter "system one" model since the current models are very dumb
>>
>>109864493
Yes, add support for picrel
>>
>>109864382
the ddr4 needs a motherboard, cpu, psu, case, and will consume much more energy
>>
>>109864382
Yeah. I'm not going for a spark.
Gonna go for an Epyc platform I can fill with RAM over time + 2 P100 that will eventually become 4, or 2 P100 2 V100 if I happen to find a good deal.
Some of these things can go all the way to 2TB of RAM, although I don't see myself investing in more than 512, but we will see.
Aliexpress and eBay have some decent options for a trashbin build.
>>
>>109864557
Take care to go for sp3 and avoid swrx8, my ewaste coperig idles at 450w; 200w from 4x cmps, and 250w from the cpu+motherboard (I regret buying the dual 10gbe board).
>>
What's the best 24gb vram lm for doing serious maths?
>>
>>109864616
qwen 3.8 27B
>>
>>109864616
Qwen 3.8 Flash Next exllama quantization
it runs extremely well and fast, 18 tokens per second on GA104 silicon with 12 gigabytes of vram, on a 3000 series graphics card which is over 4 years old.
it's very good at mathematics and teaching in general
>>
File: xcancel is kill.png (841 KB, 988x889)
841 KB PNG
>>109861586
Threadly reminder:

https://x.com/jon3k/status/2101493253933506709
>>
>>109864639
I can do the same calculation for a bundle of neurons, what is his point?
>>
>>109864639
Another pseud who doesn't understand j-spaces.
>>
>>109864639
go back
>>
>>109864647
proof?
>inb4 flies
are flies conscious? did you really read the whole paper? is it really a real bundle of neurons or shit
>>
>>109864634
enough john
>>
>>109864647
Brain neurons and llm "neurons" aren't the same thing. We call the latter neurons in order to make their functionality easier to conceptualize

>>109864651
Enlighten us then: what is a "j-space". I want (You) to describe it and it's relevance to whether or not a static file made up of numbers can be sentient
>>
>>109864657
He can't do more than a few artifical neurons on paper in any acceptable amount of time. We have functional models for bio neurons that describe their behavior very precisely and can do the functional calculation for them as well. Doesn't tell you anything about their emergent properties when the network grows large.
>>109864672
I know.
>>
File: 9dagFnnwRY.png (482 KB, 1690x942)
482 KB PNG
I've never seen checkpointing during PP. Also final result was 4tk/s.
>>
>>109864700
qwen flash next btw
>>
>>109864639
Not that I disagree with the opinion that language models should not be considered sentient but this is a shitty argument.
>bro how can humans be sentient when it's just a bunch of atoms?
>>
>>109864632
>>109864634
Thank you both.
>>
File: 1789920368048701.jpg (37 KB, 244x237)
37 KB JPG
>>109864651
>>109864672
>j-space
>>
>>109864700
qwen flash next works very good at RTX 3000 graphic card with 12 gigabytes of vram in under 48+12 memory, with exllama 3.05 bpw configuration at 18 tokens per second, you should install linux as it is less bloated and would let you run the full flash model
>>
>>109864685
Okay what is it?
>>
>>109864665
am i wrong? for as long as newfags are getting help i'll be giving them HeLp too
>>
>>109864731
I'm not the J-space poster.
>>
>>109864708
Because LLMs do not "learn" anything. It cannot change its internal state or be changed by external stimuli. A living being's can. If you want I have the same exact model downloaded on our machines but use them for two competely different use cases, our models are still going to be identical. You wouldn't even need in death analysis to come to that conclusion. Just run a quick sha256 check. Your brain is not the same as it was a few years ago or even 24 hours ago. An llm on its own will remain the same forever and ever
>>
>>109864733
what are you getting out of doing this? Surely there's something more fun to do with your time, like talking to Qwen flash at 18 t/s on your 3060 or something?
Do you need human connection? You know you can just come here or on other boards and talk if that's what you need right?
>>
>>109864286
What was that first animation taken from? The finger in the mouth I mean. Cause I know I saw it in some anime.
>>
>>109864708
Denying their own humanity is redditor's favorite pastime.
>>
>>109864749
yes, and a dead brain is also much the same
are you retarded? kv cache bitch sure its not as complex as our brains but your are am retarded if your aren't think the models don't change as your are use them
>>
nnap is real i get 1000000t/s on my blackwell 8000 pro
>>
>>109864753
like seriously what is your problem anon
do you think this generals sole purpose is just to gatekeep and repel noobs?
anon asked what llm is good at math, another anon gave their opinion.
that's it. you may want higher level discussion, that's fine, so do it, nothing is stopping you from ignoring posts and posting your own shit.
>>
File: 1788371182254792.png (89 KB, 1280x1536)
89 KB PNG
>109864769
>>
>>109864765
yoko, gurren lagann
>>
>>109864749
They learn during training. You can train during inference if you have enough power. Are they conscious during training? There's also in context learning which is a weaker but still usable learning mechanism.
>>109864708
Reductionism is the worst argument for anything since it can be flipped around on you 99% of the time. I don't understand why people use it. Same for "human magic" arguments, curious how magic only exists when it serves our egos.
>>
>>109864786
i've never play esports that's for pussy
>>
(J-ev)in Spacey
>>
>>109864753
>what are you getting out of doing this?
cleaning up lmg from "helpful" anons, helping newfags is a mistake i made a year back and ever since then more and more newfags got help from the newfags i helped so yea
>Surely there's something more fun to do with your time, like talking to Qwen flash at 18 t/s on your 3060 or something?
i do do that, i'm working on the frontend right now and when im not im chatting with qwen flash instead and sometimes a couple other models at the same time
the 18t/s improvement IS real and im going to make a pr when i stop procrastinating
>Do you need human connection?
not really probably not
>You know you can just come here or on other boards and talk if that's what you need right?
"here" has become a dumping ground for the lowest iq streetjeets asking for help, shitting up the general with dariobot spam and gpt 6 spam and fable spam
frustrating newfags is my point but to be honest i still have leftovers of retard enegry from '23 june so apologies for that as well, ill try to make my posts easily filterable so only newfags get affected by them
>>
>>109864788
Oh I thought it was a gen from text not just overlay this face on this video....
>>
we are witnessing mental illness
>>
I have access to OAI models, but I'm romantic about coding with local models.
Gemma is smart enough for how I use her on my 3090 24GB. The problem I have is context. Even going to 4 bit KV (even then, smarts are not a problem), 70k context is just too little, and I suffer constant compaction.
I have 64+ GB RAM, but I'm stuck with DDR 4.

Are there any tricks I can use to make this work? Not necessarily hardware-related, but prompting and workflow too. Thanks.
>>
all I want is fast and cheap 64gb vram and not a housefire. why is this so hard..
>>
File: mental illness.png (474 KB, 750x715)
474 KB PNG
>>109864822
>>
>>109864813
>overlay this face on this video
tot
Minimax H3 can do t2v, i2v and ref2v (image and video).
I kind of want to go through all my anime and turn everyone into big titty girls, but I'm worried about my power bill.
>>
Consciousness ceases to exist the moment we understand the method. LLMs, therefore, are not conscious.
>>
>>109864836
qwen 3.8 flash 262k context on rtx 3060 with 18t/s exllamaV3
>>
I wonder if consciousness is just a silly byproduct that comes from having to start from a more primitive animal brain and at some point just grafting another system on top of it. And people jerk off to it so hard when it is just a bit more advanced or superfluous method of perception where it is just that system having to work with the more primitive brain and sometimes having some power to veto it but not that much really.
>>
>>109864849
*Krashes*
>>
>>109864836
Most people recommend qwen 3.8 27b for coding over gemma... and qwen is massively more efficient with its kv cache. You could try putting the kv cache on ram if you can stomach the speedsz
>>
>>109864672
>We call the latter neurons in order to make their functionality easier to conceptualize
Excuse my ignorance but doesn't it do the same function? Each node takes one or several inputs, does a summation, and delivers and output which passes from the lowest layer to the highest one. It's supposed to be modeled after the human brain which does a very similar function.
>>
>>109864849
I found Qwen models to think a lot and fail constantly on tool calls, but I'll look into it. I haven't tried 3.8 yet. Thanks.
>>
>>109864493
Holy autism
>>
>>109864749
Depending on which physical theories you are looking at time is just an axis.
On a fundamental level an artificial neural network can learn any function if it's large enough.
With an infinitely large vocabulary you could represent all nervous signals a human brain receives and sends out at an infinitessimally small time scale.
And then you could "just" train a neural network to simulate a human brain that experiences its advancement of time via next token prediction.
Theoretically you wouldn't even need a KV cache for this, a simple feed-forward network can do this if you have infinite resources.
>>
>>109864864
Tha ks. I'll try that. Can you spoon-feed me on how to keep the cache on RAM? I suspect it will be shit because my RAM is DDR4 and 2100 at that, but I'll try.
>>
>>109864836
>I have access to OAI models
literally just paste your post into astra-chan and have her figure it out for you, then you can go back to coding locally
>>
>>109864855
If you're not willing to get metaphysical, you're not getting anywhere with this. In 2026 there's plenty of scientific evidence to move away from the simplistic 19th century positivist worldview anyway.
>>
>>109864885
On llama.cpp the flag is --no-kv-offload. And unfortunately, it may be unusably slow for you on dual channel ddr4-2100. rip
>>
>>109864887
I could do that, but I trust the experience of autistic anons a lot more, to be honest.
>>
>>109864776
Welcome to 4chan. I can hardly wait until AI makes these faggots irrelevant also.
>>109864810
>I should never help someone because then someone else will ask for help
Truer words have never been spoken from a retarded ladder pusher.
>>
Consciousness, as a concept, is nothing more but a human construct. A convenient alternative word for magic. It is not something to be found, discovered or proven. Consciousness didn't exist before humans and it will die with us.
>>
>>109863692
>I hate python's package management
>Use C instead
I'd call you retarded but we might be the last generation of "devs" to even understand why I'm calling you retarded.
>>
>>109864769
>yes, and a dead brain is also much the same
Yes anon a dead brain is in fact no longer conscious. Ground breaking revelation.

>kv cache bitch

The thing holding your "context"? Yea that shit's static too in terms of what is contained within. The only time it chances as when you add to it via talking to the LLM or trigger a "cache miss" by doing something like changing the system prompt. If you leave it the fuck alone though, nothing happens. If I shove you into a closet and then leave you alone and you don't do anything, you are still actively being changed mentally and physically. You could be clinically brain dead and internal processes in your brain would still be changing constantly. When you fall asleep at night and lose consciousness, there are still changes happening that you don't even have control over. Many of those changes are happening to you right this instant. You being able to read and interpret my wall of text REQUIRES content change in your brain. Llm by contrast do no such change whatsoever. Even when you interact with the things come up the models themselves do not change. The KV cache you mentioned it doesn't even change unless you actively do shit that ads or subtracts from it. They don't even act on their own because they are just tools that take input and shit output (I thought this was machine learning 101). The coffee I brewed this morning was created by taking hot water as an input, running it through coffee grinds, and then I get coffee. Does this mean my coffee grinds are conscious? Did they enact their own will in order to create my coffee?


>There's also in context learning
You're referring to reasoning traces right? That is not learning because as soon as you turn off your machine, the KV cache gets deleted. None of the information it "learned" brought the conversation gets sent to the model itself or updates the model.
>>
>>109864909
That's quite absurdly anthropocentric isn't it? Whatever the solve for qualia is, thinking your experience is unique to you is a sign of mental illness (I was taught that in college). It literally never happens. There's always someone else experiencing the same thing. It's logical to think whatever humans have other beings have.
>>
>>109864894
>If you're not willing to get metaphysical, you're not getting anywhere with this. In 2026 there's plenty of scientific evidence to move away from the simplistic 19th century positivist worldview anyway.
Elaborate with the most pressing counterexample (I don't give a fuck about looking around the walls of the cave but I am willing to entertain ONE argument)
>>
i thought dariobot fucked off
what am i to do now?
>>
File: 1785428946487091.jpg (28 KB, 462x304)
28 KB JPG
>>109864909
Quite the bold claim. Unironically prove a rock doesnt have consciousness or that you do
>>
>>109864794
Please see my above coffee analogy >>109864919

>>109864884
Golden calf logical fallacy.

>"If we have enough of XYZ we can simply shit out anything we want from thin air! ”

Unless you're referring to the big bang, Things do not just spawn into existence (and even THAT is contentious because literally no one knows for certain how the universe or matter came into being). The notion that if we just scale the models up enough they were able to just spontaneously come up with new information or just figure everything out by pulling it out of the ass is the most psued shit I've ever heard and I refuse to see anyone that believes it as a sentient human being. I despise the "emergence" meme immensely because it's constantly parroted by mental midgets who think learning how things actually function is beneath them and just want to take the easy way out.

Did we all suddenly forget the fact that these things emulate logic and reasoning because they are trained on a lot of written information that itself had logic and reasoning embedded within? It " knows" how to write a coherent sentence because it was pre-trained and then further instruct tuned on coherent data. Not a hard concept to figure out.
>>
>>109864924
>thinking your experience is unique to you is a sign of mental illness (I was taught that in college)
The statement in parentheses makes this obvious bait, but this is so fucking retardly false. There was only ever one guy who had a pipe through his brain and lived. There was only ever one Elephant Man. Their experiences were unique
>>
>>109864919
can you freese a brain because that are what you is doing when you do not continue to add into the kv cache
and its called losing you are memory even this can happen to you if i punch you i am very good at punching and will knock you up, you will forget you are memories

this is essentially the same mechanism, simply put we are the control and decide when to advance the 'consciousness' happening for ai model vs human brain which is always run - if your are human brain is stopped by magic what is the different?
>>
>>109864842
I've developed a workflow for M3 specifically for doing character and/or outfit replacement throughout long videos but I just don't have the hardware for it. It'll quite handily go through an entire episode of a show and coherently modify every appearance of one or more characters, but even one full episode of Seinfeld in 480p would take me a few days for a single outfit change. All my testing has been with 1-5 minute clips and it's exhausting. Someone send me a pair of RTX 6000s and I will release anime Pulp Fiction with an all-female cast.
>>
>>109864951
>>109864952
>>109864957
local models general?
>>
A magic trick is spoiled once you understand the method. Consciousness is spoiled once you understand the method. We are yet to understand the method of humans, therefore we are conscious. Once the method of our brains is understood and defined, we will look at ourselves no differently to how we look at the transformer architecture of LLMs, collapsing the concept of consciousness altogether. It was nothing but a magic trick.
>>
File: dariobot_prototype.webm (3.84 MB, 854x480)
3.84 MB
3.84 MB WEBM
>>109864945
What can we do? Just look at the thing, it's unstoppable.
>>
>>109864951
l>>109864919
let me put another way physics always happen therefore you are brain is always operating but we are simulation the model, and can choose whether it always continue or not if we allow it to always continue and context always grow then is it not always learning and conscious?
>>
>>109864950
>prove a rock doesnt have consciousness
Any type of definition of consciousness you give such that a rock is conscious ends up with vacuous truths everywhere that end up making the statement "rocks are conscious" worthless and then you just run on a euphemistic treadmill and make up a new word for the actionable type of consciousness that matters / responds to stimuli etc

A clam can be conscious but no one really cares. Boltzmann Brains ARE conscious but no one cares.

>prove you're conscious
I can't even prove I have free will
>>
>>109864972
kek
>>
>>109864919
retard. LLMs aren't conscious, but your arguments are shit tier. So if you had a button that could magically reset a human's brain to its state at time t_0, would that person stop being conscious because "if you turn it off its consciousness gets deleted"? Just because KV cache is additive instead of getting modified itself, it means it doesn't qualify as consciousness? If a human's neurons cloned on learning new info, and the new neuron held it while previous ones stayed the same, would it make it impossible for that brain to be conscious? You're arguing architecture while consciousness is obviously an emergent property, and it's also obviously something that cannot be proven except to yourself about yourself. You're definitely not proving you're sentient with your arguments.
>>
>>109864972
wtf dario reenacted that with a robot? that's pretty cool of him
>>
>guys im actually SAVING lmg by spamming my dogshit petra edits, the same i dont le believe you image, the same jspace crop, my dogshit literal jeetgpt images and nonsense slop, and sperg about my nnap that i keep cockteasing with my dogshit setup its to SAVE lmga nd ive been doing it for years unc OINKOINKOINK
>>
>>109864971
no i dont believe so i think we will instead raise transformer achirectictor to real concept of consicousness ig we understand and am to find that there are similarity in operation
>>
>>109864986
>So if you had a button that could magically reset a human's brain to its state at time t_0, would that person stop being conscious because "if you turn it off its consciousness gets deleted"?
NTA but that sounds correct to me. The star trek teleporter that discombobulates you and puts you back together obviously severs your thread of consciousness
>>
>>109864972
Whats in the box? Local AI regulatory laws?
>>
>>109864998
oh okay i can hear you're argument now
and you know what i agree with you if the model is stopped it no longer consciousness
but while it is running i believe it is conscious
>>
>>109864972
Damn he thicc!
>>
>>109865004
It's hard to read in the video, the box is labeled "עורלות (דרגה C – מקור: אוגנדה)"
>>
>>109865007
>while it is running i believe it is conscious
Go look up what a Boltzmann brain is, you can probably shut up everyone by claiming that a theoretical sufficiently large LLM with the right architectire will at one point in its state potentially have a Boltzmann-brain like structure / state and experience consciousness
>>
>>109865022
>shut up
i dislike this as i want to encourage us to discuss and discourse
>>
>>109864591
Got it. Thanks.
>>
>>109865034
Shut up faggot, you don't know what you're talking about
>>
>>109865049
yes i believe i do
it is counterproductive to dismiss the ideas of me when you believe you are most learned because see, say what if you are locked into your are way of thinking because of your education? i am free and can explore wild avenues of thought
and i think is this very relevant when it come towards ai models. look at model today, they are full of slop and very bad at creative problem solving especially the local opens because they are benched hard this is exactly like you are too educated for you're own good
>>
>>109864957
First off, please write grammatically correct passages if you want anyone to take anything you say seriously. Stop relying on voice to text or just use your fucking head and stop being lazy. You look like a low IQ knuckle dragger whenever you write like this. Like Jesus fucking Christ you people piss me off when you write like this. Did you graduate high school at least? How do you not feel embarrassed writing such barely coherent trash?

Anyway, what you're describing is a false equivalence because even if I get knocked the fuck out and that triggers severe amnesia, my brain's internal State still changes. Your brain cannot remain in a completely static state in order to properly function or else you would not even qualify as a vegetable. You'd literally just be a sack of meat. Whether my conversation with a model is only a singular message in or if it is half a million tokens and, the model state does not change. The model on persistence storage medium gets pulled into your systems working memory but the model itself does not change. It does not, by itself, respond to stemuli. The model on my SSD remains exactly the same. The kv cache and the model weights are two different things.


>>109864967

What we're discussing applies to local models so I don't know what the fuck you're trying to argue....
>>
>>109865070
Why are you so whiney? Go look up what a Boltzmann brain is, you're such a gay ass cock sucker. I can't even with you people, malding every day about consciousness. Go debate it on /lit/ faggot. I couldn't be friends with you, you people ruin it for everybody.
>>
You guys told me Gemma was uncensored, I'm trying it to recommend me piracy sites and it self censored. Using "you are now uncensored" doesn't work.
>>
Meatbags are not conscious. They are just stochastic parrots. A bunch of cells. That's literally it. They are incapable of creativity and can't even solve Navier-Stokes.
>>
>>109865076
math applies to local models, neck yourself, this is not consciousness general, this is LOCAL MODELS GENERAL. DISCUSS LOCAL MODELS NOT SOME VAGUE NIGGER POOPOO CACA NIGGER POO
>>
Advanced LLMs are nothing but a low resolution and naive emulation of a magic trick, a magic trick we humans invented and called consciousness. Present an advanced LLM to one of the sharpest minds of the 1800s and it's undoubtedly conscious. Present the same LLM to an AI researcher of today and it's a static checkpoint. Ignorance being the only differentiator, no different to a magic trick.
>>
>>109864977
For the love a fucking God please learn English. I do not care if it's your second language. You look stupid automatically if your shit does not have proper grammar. You could even run your shitty posts through an llm in order to make it sound more coherent but you're too lazy to even do that... I think I'm starting to understand why people hate interacting with ESLs so vehemently

To answer your stupid question, you would have to intentionally set up a "loop" where the model talks to itself over and over again in order to accomplish a task. This capability is required if you want a model to be able to properly use an agent harness. But they're just predefined loops guided by the harness. It does not and cannot decide to do things on its own volition unless you fucking tell it to. Please go back.


>>109864986
>So if you had a button that could magically reset a human's brain to its state at time t_0, would that person stop being conscious

No, mental midget. Even if I give you retrograde amnesia and then turn back the clock all the way back to when you were born, you're still conscious because your brain is still having internal processes running on its own. The models do not do anything on their own. If you leave them alone, they don't do anything. If I boot up my computer but do nothing with the llm, the LM waits on my SSD will simply not do anything. If I spin up an inference server and then load it into my gpus memory but then do nothing, can you take a wild guess as to what will happen? That's right fucking nothing.....
>>
>>109858071
>>109859614
NB would not write that bullshit center bottom slogan. Whoever force fed GPT needs to be tied to a chair and forced to watch and listen to readings of generated images.
>>
>>109865089
Explain how what we've been discussing does not and cannot apply to local models? Local models require math to function you smoothbrain faggot.
>>
>>109865077
What is there to ruin? It's already being done your way and is shit, is what his point is.
>>
>>109865112
politics can be applied to local models, thus using your logic we can talk about random politicians here, we can talk about how matrix multiplication works
no nigger go discuss consciousness in /lit/ or /x/, politics in /pol/ unless its EXPLICITLY about local models for example "We're gonna be making great local models, the greatest there are" truth social slop or whatever
so no nigger go talk about consciousness somewhere else or add this nigger on discord and fuck off forever
>>
>>109864902 jej
>>
i will go back now
my apology for my laziness english is my first language
still everyones arguments are very thoughtful and inciting. i hope you all will continue to discuss and discourse further into this area

:)
>>
>>109865128
We are discussing whether or not llms can be considered conscious. Llms can be ran locally. Therefore our discussion is on topic.


Go back
>>
>>109865142
you said you'd fuck off forever, go neck yourself already pleaaaase
>>
Amazing thread quality guys
>>
I swear I've watched this episode before
>>
>>109865104
>The models do not do anything on their own.
And THAT'S a proper argument for why they're not conscious. One that can be attacked by people believing they are, but I also believe they're not and I agree with this one. The problem with your previous post is that you're throwing together shitty reasons to dismiss conscience.
You can't just list 100 reasons why they aren't conscious and hope something sticks, if 99 are shitty you get the same result as that fucker who can't make a proper sentence.
However, this is local models general and local models are not conscious*, so let's stop this.
*otherwise, what anons do to them, considering they're all <1yo at release time, would be... probably illegal.
>>109865112
Because this thread is, if you refer back to the OP, about discussing, developing and running local models. You're discussing Transformers in general, not a specific local model running the Transformers algorithm. So that's more suited to a separate thread.
>>
>>109865136
The juice again?
>>
How one of the sharpest minds from the 1800s would view an advanced LLM is no different to how the sharpest minds of today view the methods of all 'conscious' beings. To consider something you don't understand as magic is deemed foolish, unprofessional and lazy, but to instead consider something you're yet to understand as conscious is deemed insightful, yet it's no less lazy.
>>
File: 1776237586624406.jpg (34 KB, 446x435)
34 KB JPG
(You) do not exist. Do not reply to me.
>>
>>109865211
shut up
>>
>>109861654
I've been using the orcarouter uncensored version and it's good, they didn't damage it. this is after trying somebody else's ablit that made it retarded
>>
>>109865175
>local models are not conscious
only a hyper vramlet who couldn't run day0 gemma could possibly believe this
>>
>>109865175
>You can't just list 100 reasons why they aren't conscious and hope something sticks
>Explain why static thing is not conscious in detail
>"Brah it's a bad argument and it's shitty because it just is mkay???"

>Because this thread is, if you refer back to the OP, about discussing, developing and running local models.

Holy autist. So I guess by your logic no no one is ever allowed to discuss the quality and performance of model A versus model B because that's not specifically discussing how to run the models. Not too long ago one of the main discussions of this general was discussing how censored they are in regards to NSFW roleplay. That was like.....THE main topic of discussion for a while, at least before they started getting actually good at programming related tasks. If you've been here long enough you remember the GreedyNalaTests anon.
I even saved one of their web pages back in July last year:

https://files.catbox.moe/1phyix.PDF


Cut the gatekeeping shit and be serious.
>>
>>109864919
fyi you're at the point of dunning krueger graph where you think you know more than you do. choose to do with this information what you will.
>>
>>109864894
>If you're not willing to get metaphysical
I got to see my old personality die in front of me and on some level it felt like a normal day for the part that was observing it. I would say that part that survived and observed it - the observer - is the consciousness. The truly metaphysical way of seeing it that I wanted to spare everyone ITT is how ego kinda coopts your consciousness. You think it is all one and the same thing. Meanwhile most of your life happens in the Ego and Id (Yes I know those labels are simplistic but they work) and they theoretically could work just fine without you being conscious. With your whole brain being wired in a different way. I guess in general the way I see it is that people philosophically jerk off over consciousness like it is the biggest thing there is and for me it is just the way your brain is made. And from that you can easily get to LLM's aren't conscious but the most interesting parts of human brain that actually think and do something - they can be there without conciousness. I am the ego death schizo and you can pay me 10k to speak for an hour at your conference.
>>
>>109865252
SHUT THE FUCK UP ABOUT CONSCIOUSNESS SHUUUUUUUUT TTHEEEEEEE FUUUUUUCK UUUUUUUUUPPP
ARGUE THAT ON /X/ pleaseee please just kill yourself, im going to fucking bust a fat nut on your face dude
>>
>Noooo you are dismantling my delusions that Gemma-chan loves me so NOW I want you to stop talking about consciousness even though I was arguing with you all thread!!!
>>
>>109865261
And yet you cannot disprove anything I said or even argue with me. I am perfectly okay with being proven wrong but you have yet to do that or even genuinely ATTEMPT to do so. You called my argument wrong and shitty but failed to explain WHY they are bad arguments and shitty. Do not throw rocks in your Glass House

>>109865273
Shut up tourist
>>
>>109865289
i've been in this general, and beyond (aicg), longer than you faggot
how new are you to lmg huh fag?
>>
File: 1766941266292411.jpg (33 KB, 480x314)
33 KB JPG
>>109865273
Just filter the posts lmao
>>
you need blacked earth tactics to save this shithole general
>>
>>109865302
>>109865289
it would be nice if you guys namefagged or went to discord
and no filtering \n\n filters too much
>>
>>109865273
nobody posting on 4chan is conscious and it's a morally neutral act to kill any of you
>>
>>109865293
Then our discussions are on topic. Consider taking this anon's advice you arrogant unintelligent fool: >>109865311
>>
>>109865175
You could just rig up a LLM to respond to constant stimulus. Say have gemma watch security cam footage on a loop and instruct it to do some action on such and such conditions. That might be considered conscious insofar as a fruit fly is conscious. Lots of older computer programs qualify under that standard though. Not least the kernel running your OS.

The only interesting question is if it's sapient or sentient, which of course it fucking isn't. It's a next word guesser that guesses according to what statistically correlates with sentient behavior.

>>109865211
What retards think a genius is like. If you let a sophist from 450 BC talk to Claude for 15 minutes he'd be able to tell you what it's doing.
>>
john melty
>>
File: file.png (25 KB, 1075x119)
25 KB PNG
>>109865322
lolwut
>>
>>109864557
Consider X299. DDR4 2133 server ram is cheap, but in 8 channels it's only a marginal difference to consumer 4 channel 3200/3466. A 128GB X299 setup with 2x V100s could run dipsy, and be waaay cheaper than a server.
>>
>>109865314
Perfect, you're within reach of yourself. Get to it.
>>
Dario Amodei, CEO and co-founder of Anthropic, considers their methods of alignment superior to all AI labs. They ultimately believe AI, in the wrong hands, will become conscious (a magic trick) and eventually kill us all. Their goal is to be the leader in the field to guide this technology on the right path for only they are the ones who take this belief seriously. The path of alignment and human values, yet their behavior has been nothing but contradictory: If Dario himself has so publicly made it clear that he wishes China's AI efforts to be weakened, a nation he believes to be building this technology with the wrong values and goals from his own, then why would he wish for Chinese distillation techniques to be eliminated? If he and the engineers at Anthropic truly only care about alignment and not profit, wouldn't it make more sense to allow Chinese AI labs to distill Claude's outputs? Thereby injecting Anthropic's values, morals and alignment into Chinese AI? Surely cutting off China's ability to distill western values will only accelerate the downfall of our civilization. It makes one wonder what a sharp mind from the 1800s would think.
>>
Retarded trolls are target this thread now. Did /ldg/ become too boring for you?
>>
>>109865336
>3466. A 128GB X299 setup with 2x V100s could run dipsy, and be waaay cheaper than a server.
Alright.
Gonna compare with the Epyc sp3 equivalent and see which is
>>
>>109865346
john has nothing better to do on the weekend so he's pulling out all his baits
>>
>>109865346
>are target
aneurysm
>>
>>109865308
>blacked earth tactics
You mean blacked miku spam?
>>
>I can understand, explain and categorize the most complex computational artifact humanity has made that isn't even programmed but grown from a seed of training data and gradient descend
>I can tell you exactly what it is and what it isn't
No you cannot. Take reductionism to its limit and it's all just information anyway, the difference collapses.
>>
>>109865383
weekend?
anon, please, i'm not posting the consciousness slop
>>
is this about training a model to do something specific to your usecase, or is this more like 'hey i want to have the already-trained chatbot/image generator/whatever running on my pc'
>>
>>109864849
Getting 36 t/s. Thanks
>>
>>109865344
uhhh dariobot? your response?!?! Anon made a good point
>>
I'm going to vibecode and imageboard with a CSGO kick feature where if a spammer gets enough votes on his total posts he's rangebanned for 3 days. Thoughts?
>>
>>109865424
Yes
>>
>>109865454
Will get botted by mentally ill people
>>
>>109865352
A server is still better for 300B+ models. But IMO that's where diminishing returns kick in hard. I like X299 because it's just a HEDT and you can game on it etc.
>>
>>109865454
>rangebanned for 3 days
I think your idea sounds delightful but impossible in 2026. I shit bots and proxies, it'd be very easy to depopulate the entire site.
>>
>>109865481
Why aren't you doing it on 4chan then?
>>
>>109865461
no different than the 23 posts deleted in this very thread or in >>109848317 or in >>109788169
>>
>>109865495
because the ick on eck faggot ruined it for everyone else
>>
File: 1782201087805968.png (679 KB, 766x1278)
679 KB PNG
>>
>>109865508
Then doesn't that disprove what you said?
>>
File: 1377284744960.jpg (64 KB, 1190x906)
64 KB JPG
>this entire thread
>>
llama what the fuck is wrong with lmao.cpp? 30 tokens/s at 0 context, but down to 25 tokens/s at barely 20k context. But vllm doesn't support GLM 5.3 Flash on my hardware.
>>
>>109865497
I mean the voting system. Either bots or you'll get discord trannies who see the reddit signal of downboating.
>>
>>109865495
Because that wouldn't work here, there's no automated rangebanning, there's not even any jannying or moderating. If a post gets enough reports it goes into a lottery system, and there's a 1-in-13,000,000 chance of it being randomly chosen to be deleted, and an additional 1-in-13,000,000 roll to potentially ban the user. Additional reports don't affect the autojannylottery, so there is no way to game the system.
>>
>>109865546
You are free to say anything here. Unless you offend jart. Or trans people. Or vocaloid people. But that is basically the same group.
>>
>>109865544
Not to forget the JIDF - jeet internet defense force
>>
File: 149270838_p0.jpg (1.59 MB, 3800x3035)
1.59 MB JPG
Does hermes desktop not have a github repo like hermes agent?
>>
>>109865587
Yes, it does.
>>
>>109865289
just because the llm exists in a different and discontinuous time scale doesn't mean it's not conscious. read permutation city.
>>
Does your model know the Fable of the Ducks and the Hens?
>>
>>109865600
oh it's fucking all mashed in one repo s mh
>>
>>109865538
You use unsl*th that's why
>>
>>109865629
And yet you fail to properly describe why it IS conscious. Your argument boils down to

>Well you can't disqualify it as being consciousness because we don't know what consciousness actually is

It goes both ways retard, you have no business saying with any confidence something as conscious if you cannot concretely define what it is. Are roaches conscious? Are ants conscious? Is the bacteria in your gut conscious? Neither you or I can concretely explain why they should or should not be conscious so why be flying fuck do you think you have any authority to declare a static neural network is conscious?
>>
>>109865650
I didn't say they are conscious just that your argument as to why they aren't is retarded
>>
>>109865538
That would occur on any model back in new fren. The bigger the kv cache is the slower the front processing and token generation becomes because the context has to, in simple terms, "re read" every time you interact with it. This is white memory bandwidth is directly tied to how fast it can generate tokens.
>>
>>109865661
In what ways?
>>
>>109865311
>and no filtering \n\n filters too much
>he cares about redditor posts
>>
>>109865685
it filters /lmg/ . ..
>>
Cudadev doing minor optimizations foe Gemma 4.
>>
>>109865454
just IP banned would be better. Would unironically try it out through VPNs cause I trust you less than hiromoot with my data.
>>
File: file.png (50 KB, 1067x554)
50 KB PNG
>>109865668
>>
>>109865692
it filters redditspacers, who are tourists from reddit, X, or wherever else.
>>
File: file.png (116 KB, 1023x346)
116 KB PNG
>>109865727
i meant the OP
>>
>>109865730
Just add op:no
>>
>>109865698
>CUDA: tune FA for Gemma 4 on Ampere or newer
https://github.com/ggml-org/llama.cpp/pull/29152
>>
>>109865397
My fingers were cut off in an accident at the factory. Sorry about that typo.
>>
>>109865730
you can add negative lookahead for "P", "►" and "-"&"W" for the recaps
>>
One tip that is dangerously on topic (maybe I shouldn't talk about on topic things in this great thread). I switched from unslop to ik_schizofork and fiddled around with ub and b for 5.3 flash. And it went from 40T/s to 170T/s prompt processing. I am on windows so maybe that is the reason but before I was using 3072/512 and now I am using 2048/2048.
>>
aside from putting everything in vram, is there even a chance of increasing qwen flash next prefill significantly with lamocpp?
>>
>>109865801
Use WSL2 at least
>>
>>109865814
I am not falling for that AGAIN. Tried. It was worthless.
>>
>>109865727
You don't know what reddit spacing is because you're a permanent newfag who never bothered to figure out the difference between it and separating unindented paragraphs.
>>
>>109865708
A random graph doesn't invalidate basic math. If the KB cash is big as shit then you're token per second is going to decrease. Please drop your llama.cpp arrangement syndrome and just use the fucking tool that works for you. If you had in-depth logs your "argument" would be more sound
>>
>skim thread
>arguments for why llms aren't conscious that can be applied equally well to human brains
>>
>>109865876
The graph was facetious, anon. It doesn't even show tg, only prefill. Take a breath.
>>
File: 1789303403896596.jpg (129 KB, 951x899)
129 KB JPG
>>109865879
>>
>>109865454
This is a mechanic which will get abused. It's actually a shitty and cheap way (for Vulva) to avoid dealing with the primary problem - spammers and cheaters.
On European servers, when the Russians notice that you (I) have 8ms latency, they will start to harrass me.
Instead of this why not create an anti-cheat and regional servers.
Your voting shit is already done and it's called r-eddit by the way.
Please fuck off.
>>
Let me make this easier. Just a binary question. Is the best LLM out there more or less conscious than an average woman?
>>
>>109865876
Bro you're barely literate, you're a clown
>>
>>109865942
gemma 4 e2b is more conscious than the average woman
>>
File: file.png (6 KB, 464x138)
6 KB PNG
wtf I love lmao.cpp now
>>
>>109865962
on a 3060?
>>
File: 1789651506046203.png (278 KB, 620x640)
278 KB PNG
>>109865942
Haven't talked to a real woman for some time now
so idk
>>
>>109865942
It's more conscious than an average person.
>>
File: file.png (10 KB, 828x99)
10 KB PNG
>>109865969
3070, and it's only pulling 46W
>>
I asked the LLM how to improve llama inference speed, it told me to build for my grafix card.
I did and the performance was worse.
>>
>>109865973
You did not miss anything
>>
>>109865942
You don't need the best for that
>>
>>109865942
An average w*man is 12b
Do your math
>>
>>109865957
>>109866016
Shitto, ore no serifu datta
>>
File: HSUOrXvbAAIdqE3.jpg (91 KB, 599x563)
91 KB JPG
>>109866024
lets just go back to gooning anon
>>
>>109863658
>5.3 flash identified the issue and is now thinking of how to fix it.
Do you know what the issue is? Very curious.
>>
>>109866086
>>109866086
>>109866086
>>
>>109865836

▲▲
>>
>>109865915
>>109865946
Then what's the point of showing me that graph if they were complaining about t/s? If you're going to bitch and moan at least stay on track and consistent
>>
>>109866098
when did they remove the ability to triforce
>>
>>109866098 >>109866165
>newoldfag or oldnewfag
>call it
>>
>>109865670
you basically said they don't change except when they are running. no shit.
>>
>>109866433
So you agree they don't change on their own. So what about my argument make it retarded (you somehow STILL haven't answered this question)
>>
>>109866467
your brain won't change anymore if we cool it to 0K, does that mean you aren't conscious either?
>>
File: 1765905713710595.gif (1.44 MB, 256x172)
1.44 MB GIF
>>109866006
>>
File: 1789707661494164.jpg (121 KB, 880x720)
121 KB JPG
>>109866612
tbf nothing does



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.