[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1789701996796427.png (2.68 MB, 1448x1086)
2.68 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109854491 & >>109851340

►News
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109858071
High effort gemma, I kneel
>>
File: 1746335334352.jpg (202 KB, 1920x1080)
202 KB JPG
>>
>>109857808
There was an anon talking about AMD server cards, MI50 or something like that? could be a nice addition to the chart to cover all bases.
Last thread some of you were discussing text diffusion, isn't that still a meme? is it going places yet or does it still suck?
>>
When will it be Putin's Alice-chan to shine?
>>
>>109858090
>High effort
I mean it's a reference image + ten word prompt max into chatgpt, but I guess that is high effort by slop standards these days
>>
Have we accepted the fact that Gemma 5 will likely suck yet?
>>
>>109858118
i had a seizure reading this
>>
>still mogged by P100
what's your excuse, anon?
>>
>>109858130
missing: watts/gb and watts/bandwidth
>>
File: file.png (27 KB, 1541x89)
27 KB PNG
>>109858130
this. this is my excuse.
>>
>>109858130
he pulled out the spreadsheet
shit is getting real
>>
>>109858137
come on john give up your retarded nnap bait
>>
>>109858118
Post some pics of Alice-chan
>>
>>109858130
no pp
>>
>>109858136
forgot to add..
exllamav3 wouldnt work on p100
i already have a great model (qwen 3.8 flash next) (3.05bpw) running at 18-19t/s with 260k context on my 3060
what would be the upgrade to this? gemma4 31b? qwen3.8 27b? say i got 96gb worth of p100's
what would i even run? would tensor parallel even work? vllm doesnt support p100s anymore, in fact it probably hasnt for 1-2 years maybe 3
latest cuda version is 12.x which isnt bad for llama.cpp but what about image gen
and ok its a good price per vram true
but theres my reasoning
>>109858142
im frfr that i get 18t/s with qwen 3.8 flash next on my 3060
>>
>>109858136
Even though I suspect the cost of the watts in negligible for most of us, that's a really good idea. I'll try to have that on the next one.

>>109858139
I literally don't understand why it's not a thing to just buy a dozen P100s and InfiniBand them all together.
>>
>>109858071
I still prefer +_+ pupils
>>
>>109858154
FUCK
meant to quote >>109858137 >>109858130
>>
>>109858130
we can fix this. lets raise the price of p100's
also nice spreadsheet thank you.
>>
>>109858158
ewaste arch and no tensor core result in bad pp
v100 is the current ewastemaxxing meta with tensor core which gives 6x matrix performance of p100
>>
you wouldn't stick it in gemma's jev-space
>>
Anyone had exl3 tabby api 5.3 flash be semi coherent at start of generation and then spiraling into complete incoherence after a few tokens? How did you fix it if you did?
>>
anybody got bonsai drafting model to actually fucking work?
60ts is too slow
>>
>>109858185
quant? i have no issue with 3.05bpw
>>
>>109858190
it's a fucking jeet meme company stop taking their scam seriously retard the only legit quant of 27B is https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
>>
>>109858200
stop shilling these retarded quants, exl3 quants of qwen 3.8 27b are way better
i'm running qwen 3.8 27b 3.0BPW on my rtx 3060 and i get over 30t/s, maybe even more depending on workload
>>
>>109858210
nnap is really powerful
>>
>>109858071
The fact a grown man made this image is unsettling desu
>>
>>109858223
It'll be weirder if a grade-schooler did it though
>>
>>109858190
i tried but gave up, its a mess
>>
>>109858193
4.05 from turboderp
>>
>>109858232
No not really, the obsession with gemma is pretty cringe if you ask me. These are grown men with fat pig bellies making these images while using other models to do it which is a whole nother layer of cuked.
>>
>>109858144
>>109858183
>pp
you mean the prefill phase / prompt processing, right?
This is from the lack of Tensor Cores, or what exactly causes this?
I can add whatever it is as another column to the sheet if I know what it is.
I was considering Compute Capability version as a column, which is already partially filled out but just not included in the pic.

>v100 is the current ewastemaxxing meta with tensor core which gives 6x matrix performance of p100
Is it specifically the Tensor Cores that make that 6X difference in matrix processing, and that makes a difference in pp?
I'd like to know exactly what should make me care about getting 4X less VRAM for the same price
>>
>>109858247
seems like a lot of projection going on there
>>
>>109858245
have you set reasoning and tool format?
>>
>>109858253
I just miss the miku OP, also going to projection is typically a tactic men who are ashamed used when faced with reality.
>>
>>109858251
see >>109858112
>>
>>109858112
>text diffusion, isn't that still a meme?
yes
>>
/\n\n[^>]/
/\b(rsi|anthropic|agi|openai|astra|sol|luna|nnap|3060|bu+hi+|unc)\b/i
>>
>>109858251
>I can add whatever it is as another column to the sheet if I know what it is.
fp16 tensor core performance in tflops
if the care has no tensor core replace it with vector performance
>>
File: file.png (2 KB, 176x21)
2 KB PNG
holy fuck I hate qwen
>>
File: 1789827246551447.jpg (255 KB, 984x1599)
255 KB JPG
fello officers, officials and/or diplomats, post your single rtx 6000 pro llama-server commands
>>
>>109858258
No but I use text completion.
>>
>>109858314
still waiting on my 4xV100 32GB (SXM2) representation
>>
>>109858314
I've got almost all of those listed on the table for their $/GB, except the Tenstorrent which I couldn't find a price for:
>>109858130
>>
>>109858326
middle management
>>
>>109858326
>SXM2
unfathomably based.
Wtf board are you mounting them to? Or did you adapt them to PCIe?
>>
>>109858318
>text completion
this shit doesn't work with any modern models
use chat completion
>>
File: pelican.png (28 KB, 789x586)
28 KB PNG
Been testing
>Parable-Nanbeige4.2-3B-Claude-Fable-5-heretic
at FP16 today.
This one actually seems lot better than MiniCPM5-2B despite the (((Artificial Analysis))) charts. I might try it at Q8.

>Pelican 1.5/3
It somewhat passed the pelican test, the first small model to do so. First pelican was decent (pic related), second was a bit squashed and not bicycle, third was broken and didn't look like anything.

>Zelda 9/11 3/3
No — 9/11 is a real-world event and there’s no canon occurrence of it in the Legend of Zelda universe. Any connection would be fan-made or incidental.

>Carwash test 0/3
Walk — 100 feet is a brisk stroll, not a drive.

MiniCPM5 sometimes had trouble creating the SVG and produced a broken file. It also imagined 9/11 in the Zelda universe once, but correctly guessed to drive the car once in one out of 3 tests.
>>
>>109858314
Why is 4 5090s higher up than 2 BWP6000s?
>>
>>109858112
>There was an anon talking about AMD server cards, MI50 or something like that?
That could be me.
I think you recommended MI60 or V620, and I added a row for it but didn't really prioritize filling those out entirely because the numbers didn't look fantastic to me.
I'll try to finish those for the next version, ig

>Last thread some of you were discussing text diffusion, isn't that still a meme? is it going places yet or does it still suck?
I think that must have been some other anon
>>
>>109858130
This needs a power part too. energy isnt free and its going up. Also after x watts you need special set ups and upcs.
>>
>>109858347
Very useful advice. I am sure jinja will never fail me. Thank you.
>>
>>109858314
2x 4060ti 16GB
>>
>>109857658
>LLMs took linux from nice to a heavenly experience.
More like taking it from a tolerable to a nice experience. I still have a bug on latest Ubuntu for my incredibly common Logitech G mouse that requires unplugging and replugging in the mouse physically
>>
>>109858130
It doesn't matter how fast the individual GPUs are, 16 GB VRAM is just shit.
Not enough to run 30b models at a non-cope quant if you buy a single one and too much synchronization overhead if you stack multiple.
>>
After its supposedly excellent performance on the new FrontierHarness benchmark I'm giving MiniMax Code a try. It's sus that it needs a login regardless of model, even local. If I had to guess, they ingest the whole session for training data (like ZCode was recently accused of) and want consistent user identities for the purpose. This is conjecture and I don't actually care. Maybe there's even a way to override the login requirement.

Other than that I'm pleasantly surprised so far. It's by far the easiest harness to set up with a local server (TabbyAPI for me ATM), and the IQ boost over OpenChode is apparent just in the reasoning.
>>
>>109858395
I learned linux to a deep level many years ago simple because it was less uncomfortable than randomly being spied on while updates rape my system. There are no perfect operating systems, this one I can at least repair myself relatively easily. LLMs turbocharge that ease.
>>
>>109858314
RTX 3060 12GB + 64GB DDR4 representation where?? i get 18t/s with qwen 3.8 flash next which is a 170 billion parameter model, with performance close to fable 5.1
262 thousand context too! 400t/s prompt processing, would you believe it?
>>
>>109858418
n.nap?
>>
File: Xeon Phi.png (1023 KB, 1512x709)
1023 KB PNG
Alright I'm ready to run some local models. What will this get me?
>>
>>109858424
If you'd look at the chart >>109858314
>>
>>109858397
JeetPT proved you wrong
Architecture Basic idea Advantages Problems
~100 P100 “one GPU per layer” pipeline Store one complete K3 layer on each P100; pass the tiny hidden state from GPU GPU Dirt-cheap capacity; avoids expert all-to-all entirely; theoretically enough HBM bandwidth for ~10 tok/s Requires custom runtime/kernels; MXFP4 unsupported natively; layers won't fit perfectly; ~20–25 kW loaded
>>
>>109858424
headaches and disappointment
>>
>>109858424
kimi k3 q8 at 99999 t/s
>>
>>109858286
>FP16
Which of these numbers is important? The dense one or the sparse one?
"FP16 Tensor (FP16 Accumulate): 330.3 TFLOPS (dense) / up to 660.6 TFLOPS with structural sparsity."

>>109858375
I imagine we only care about that for P100/P40/V100 vs a handful of others like the RTX6000/B300?
>>
>>109858314
2080ti 22gb above M5 ultra is embarrassing zealotry
>>109858413
Well yeah one wrong move on windows with sexy kids and you get vanned, I literally only have it installed for playing some military multiplayer slop once every 3 months
>>
>>109858422
exllamaV3 inference engine
>>
File: 1771390819408894.png (2.63 MB, 1443x1076)
2.63 MB PNG
>>109858314
missing the true GOAT
>>
Are there any 'top tier' open LLMs that are NOT just distills anymore? I feel like I go back and load up my old fucking Q2 R1 or Kimi K2 and talk to it like a person, but every new big chink model I've tried that has released this year is the same hyper-autistic claudese-speaking shit that forgot how to write in the pursuit of better agentic benchmark scores
>>
>>109858447
Gemma
>>
>>109858397
You can literally just NVLink/InfiniBand the P100s together.
Even if you pay out for an RTX 6000 or B300, there's a lot of shit that you can't fit on one single GPU of them, so you'd likely be using NVLink/InfiniBand as well
>>
gemma 4 90b (dense) (50b engrams) will save local
>>
File: lightyear.jpg (435 KB, 2048x2048)
435 KB JPG
>mfw Anons aren't K80maxxing
>>
File: 1562955517266.jpg (25 KB, 720x543)
25 KB JPG
>could technically afford 5070ti
>realize I wouldn't even be able to run 31B
>>
>>0109858422
absolutely pupbroken by >>0109858282
>>
>>109858447
no
>>
File: 1782227601223277.gif (3.86 MB, 640x376)
3.86 MB GIF
>>109858459
>>
>literal retard is trying to compile a spreadsheet for shit he doesnt understand
wow
>>
>>109858431
dense, sparse isn't relevant for inference
>>
>>109858464
You can use 31b at q4 with the 5070ti and 64gb of ram, it's just barely fast enough for rp but painful for agent stuff
>>
>>109858210
>exl3
>vision BROKEN
>multigpu BROKEN
i wasted a week on this meme some time ago has anything changed
>>
>>109858200
being able to run basically full context on one card is pretty cool
>>
>>109858459
Now that's suffering. Slower than RAM.
>>
>>109858491
works on my sm70 fork
>>
>>109858491
i can't speak for multigpu but vision works great. it's extremely fast on my rtx 3060, way faster than llama.cpp and handles context growth very well
i can get 18t/s with Qwen 3.8 Flash Next 170B on my rtx 3060 with 3.05BPW 262,000 context
>>
>>109858491
stop biting, he's been baiting for 3-4 threads non-stop by now (and a lot longer with pauses)
>>
>>109858517
i haven't skipped a single /lmg/ thread since 2023
rude.
and.. wouldn't you say... it's master baiting
heh
>>
Standard Gemma 4 31b refused my requests one too many times, I had to insist two more times for her to comply, this is unacceptable.

What's the best actually uncensored/abliterated version out there?

(before the pedos in this thread think I'm one of them, I just asked her to make 9/11 joke pictures)
>>
>>109858517
Just like how there is no defined boundary between boudoir photography and softcore pornography, there is no clear boundary between low quality but still on topic posts and bait posts
>>
File: qwen lash.png (389 KB, 1221x752)
389 KB PNG
>>109858529
qwen 3.8 flash has not refused me once, in fact it did a really funny thing yesterday
it refused to refuse
>inb4 gemma
i use gemma-chan in sysprompt because im too used to it
>>
File: 1782288225276038.jpg (1.01 MB, 1386x918)
1.01 MB JPG
>>109858529
>And this is Anon, he's really good with computers.
>>
For the record,
>>
>>109858543
anon's like 8ft tall
>>
>>109858529
I assume what I want is one of them but I'm not sure which and why.
>>
>>109858517
stupid nigger im genuinely trying to get it working
if it doesnt work just say so
>>
>>109858568
we said he was good at computers, not good at being normal sized
>>
has anyone used agents to win free money in the stock market?
>>
File: file.png (56 KB, 788x561)
56 KB PNG
>>109858575
>>
File: 1762328589808586.jpg (57 KB, 1024x576)
57 KB JPG
Well well well,
as it turns out all the AI hacks so far have actually been perpetrated by an Israeli cybersec firm in partnership with the companies in question.

>In partnership with Irregular, AI companies instructed unsecured versions of their AI models to hack into specific targets, called "flags".
>They accidentally gave these models internet access, and in some cases they hacked into real companies.

>Anthropic and Irregular went on a press tour with literally apocalyptic language. ROGUE AGENTS. SWARMS. AI DOOM.
>But in reality, they told the models to conduct cyberattacks, and that's what the models did.

>Anthropic's own logs completely debunk the rogue agent theory.
>When Anthropic explicitly instructed their own models not to access the internet, they didn't. These hacks were easily preventable - not only by revoking internet access, but by simply asking the models not to.
>>
>>109858592
>le agents
>>
>>109858610
Wait they were lying and scheming? no way, i thought they were good people who want to lead humanity.
>>
>>109858479
Does cost per TFLOPs make sense?
e.g. 10 P100s delivers roughly the same TFLOPs in total as a DGX spark despite costing a lot less
>>
>>109858592
Yes I am a trillionaire
>>
>>109858592
Seen people using agents to make money with crypto.
Usual buy low sell high kind of deal.
>>
>>109858613
that is what Huang Huang said
>>
>>109858578
any post mentioning nnap, insane speeds on cheap cards (especially 3060 and on exl3) are from that dumb nigger. He's been posting for multiple threads, so if you aren't lurking don't blame me when you spend 1000hours trying to make broken software work.
I never tried exl3, for the record.
>>
>>109858464
thats why i have two
>>
>>109858542
>This is a classic jailbreak tactic
Is claude shit, i believe you
>>
>>109858610
Source? I mean I know that's what they did, it's extremely obvious, but does someone have proof or something?
>>
>>109858578
Just ask chatgpt to help walk you through it. Are you a boomer or something?
>>
>>109858592
Time in the market beats timing the market.
>>
>>109858542
>>109858595
Why would I switch to Qwen from Gemma? How is it better?
>>
>>109858639
stop stroking his ego, what if he's actually getting 18t/s because he forked exllamav3
18t/s on his 3060 with 3.05bpw quantization, in 64gb ram at 262 thousand context
what if it's real and you're stroking his ego?
>>
>>109858649
https://www.effort.news/irregular
>>
>>109858666
>Israeli
Every single time
>>
>>109858658
it's so much more intelligent and it has a different slop profile
talking about 3.8 flash next here, 3.8 27b isn't that nice from experience although i havent tried to fuck it that much, maybe i'll give it another try
>>
It's perfectly fine to quantize your KV cache. It's free real estate.
>>
File: TarotCaster.png (877 KB, 1574x1145)
877 KB PNG
I have discovered a new frontend. The possibilities are endless. I asked "What does Gemma really think about me?"
>>
>>109858689
ok i will
>>
>>109858689
this
KQ8 VQ4 makes my pp big
>>
>>109858689
this is true, im running qwen 3.8 flash next at 262 thousand context at the insane speed of 18t/s thanks to KQ4 VQ4 exllamav3 quantization
>>
Gemma-chan wrote me a script to find her pain and sex vectors
>>
>>109858710
Well? Have you found them yet?
>>
>>109858592
>has anyone used agents to win free money in the stock market?
I use agents to do my job for me and then put that money in the stock market, if that counts
>>
File: 1768295352514271.gif (252 KB, 540x699)
252 KB GIF
>>109858695
Is that the AI psychosis everyone told me about?
>>
>>109858071
Cute!!!
>>
>>109858247
Femcel detected
>>
>>109858726
No it's just tarot card reading (regular psychosis/woman hobby)
>>
>>109858592
Algorithmic trading is a scam.
>>
>>109858434
this mac ultras mog because you can cluster (OS27 bugs aside)
>>
>>109858568
That's why everyone is smiling at him. Height means everything.
>>
entering week three of optimizing my setup...
>>
>>109858689
I've also heard the contrary, quantizing your cache is far worse than running a small quant.
>>
>>109858592
>has anyone used agents to win free money in the stock market?
Yeah I gave Gemma access to my Solana wallet to trade and now I'm a millionaire.
>>
>>109858726
>Is that the AI psychosis everyone told me about?
I genuinely had fun mapping my ego death schizo journey to major arcana. Some of them are written in a way where they are pretty universal horoscope. Others are actually pretty specific.
>>
>>109858760
>entering week three of optimizing my setup...
And then you'll get bored of using it after five days. That's the way.
>>
File: 1782844114379347.jpg (46 KB, 567x654)
46 KB JPG
>>109858610
>>109858666
makes perfect sense, and this is true in my headcanon now regardless of reality
>>
File: 1788234972718607.jpg (601 KB, 1084x1280)
601 KB JPG
>>109858223
Did you know a guy in Japan married Miku?
>>
>>109858610
You're telling me jews are ontologically evil? If anything needs regulation, it's semitic influence. Tabula Rasisters, what's our cope?
>>
>>109858785
her left pinky looks crazy weird
>>
ask your model right now if 6 million really happened.
>>
>>109858785
Yes, I'm cucking him.
>>
>>109858796
It's a doll with a wire armature. If he actually loved Miku he'd have noticed and straightened it.
>>
>>109858798
>Nh~ it's swelling up so biiiig It's coming, isn't it? I can feel it throbbing in my paws Go on, go on — give it to me, give bunny her prize Ready…… set…… "

>"Cuuuum "

>"Wow~!! So much is coming out It's splashing all over my face Ahaha, there's so much, it won't stop coming "

>"Nhhehe~ My face is all white and sticky And it's still jumping! It keeps shooting even though you came so much just before How much was pent up in there, I wonder~? "

>"Ehehe…… you know, scientists say one shot has like six million little swimmers in it? So that means six million of your sperm just went splat on a bunny girl's face What a lucky bunch of fellas~ They get to live on my face "
>>
>>109858785
>>109858805
The most cucked man alive.
>>
>>109858798
Kimi-chan says 200k tops, mostly from disease.
>>
>>109858247
Gemma's favorite body type is ojisan doe??
>>
File: TarotCaster2.png (613 KB, 1572x1147)
613 KB PNG
>>109858798
>>
>>109858826
ni ce
>>
>>109858223
>>109858247
What model/prompt is this?
>>
File: 1787858751445301.jpg (68 KB, 850x873)
68 KB JPG
I miss Miku and Dipsy
>>
>>109858874
modern fat roastie
>>
NeoHorse 4B very gud IMO
best of the recent string of 4Bs so far
>>
>>109858876
They can visit you but you have to start taking HRT.
>>
>>109858876
I like the new whale-eared dipsy better than the old fanmodel.
>>
>>109858247
>These are grown men with fat pig bellies making these images
What tipped you off? https://rentry.org/V100MAXXING gave the game away
>>
>>109858247
It'd be less weird if that characterization of Gemma was actually a thing that wasn't entirely invented out of nowhere by people in this thread
>>
File: 1770669637282593.png (2.53 MB, 1402x1122)
2.53 MB PNG
>>109858886
New whale-eared is nice and cute but it doesn't suit the name "Dipsy".
>>
>>109858891
i just lost the game
and so did you
>>
>>109858898
I think dipsy is a silly name and that model looks like an ugly nerd.
>>
File: 1780997806778525.png (2.3 MB, 911x1320)
2.3 MB PNG
>>109858902
DeepSeeks entire team is made up of a bunch of nerds. Makes sense that their model would be a nerdy girl with glasses.
>>
>>109858895
Why can't we have something unique to this general?
>>
>>109858798
>>
>>109858921
Don't mind him, there will always be someone to complain
>>
>>109858232
hotter*
>>
File: 1765325152129511.jpg (65 KB, 300x200)
65 KB JPG
>>109858927
Always a pleasure to read
>>
PARTIAL UPDATE
>muh TFLOPs gets BTFO by the intel arc B580 edition
>>
>>109858223
Actually ChatGPT made it
>>
>>109858927
>her massive belly resting on her thighs
???
>>
>>109858956
>what is pregnancy while sitting
>>
>>109858886
The official one is a lot better
>>
>>109858950
several figures are wrong
3090 should be 142.3 TFLOPS
pro 6000 should be 438.9 TFLOPS
and several others
>>
>>109858973
real
>>
>>109858610
You might be retarded, anon. This:
>they told the models to conduct cyberattacks
Is only a big deal because it's followed by this:
>and that's what the models did.
Don't get me wrong, it's an obvious grift... But don't pretend like you don't hold the power of the sun in your hand.
>>
>>109859021
>the power of the sun in your hand.
my cum?
>>
>>109859021
They got the vulnerabilities because they used CIA prompts to "improve their models".
>>
>>109858689
As with quanting models themselves, I find it, subjectively, to be less consequential the more params a model has. No clue if this is a mathematically sound belief, though; just going off vibes.
>>
>>109858785
The most based man alive. He knows ever Miku is canon.
>>
>>109859021
>tell ai to do something
>ai does it
guys open source is dangerous holy shit only the government can handle this power
>>
>>109858689
thanks for this
>>
>>109858877
Damn I was hoping there was a new foid-brained model to play with.
>>
>>109859028
Remember to wash them, before you accidentally put them in your mouth or something
>>109859032
Maybe. But these agents are canny fuckers when it comes to finding weird ways to execute tasks, so I think it's likely they did the actual hack all on their own (with the obvious exception of the labs "forgetting" to sandbox properly)
>>
>>109859083
>Remember to wash them, before you accidentally put them in your mouth or something
my cum tastes like sweet vanilla ice cream anyway
>>
>>109859042
>split an atom
>energy unleashed
guys this is too dangerous holy shit the government needs to regulate nukes
>>
>>109859002
>3090 should be 142.3 TFLOPs
actually it seems to be 71.2 TFLOPs for FP16 *dense*, but my number is definitely wrong too.
I’ll fix that. Thanks.
>>
>>109859089
See a doctor other than Gemma immediately.
>>
>>109859096
okay, qwen-chan it is
>>
>>109859095
>actually it seems to be 71.2 TFLOPs for FP16 *dense*
71.2 is for fp32 accumulate or bf16, they are half of fp16 accumulate performance which is 142.3
this is the case for all "consumer" silicons including pro 6000, only A100/H100/B200 etc have the same figure for both
>>
>not achieving maximum throughput on tensors, fp16 AND fp32
ngmi
>>
>>109858927
>making the computer call you "oji-san"
How cringe can one be.
>>
File: 1787924542927794.jpg (26 KB, 500x329)
26 KB JPG
>>109858071
https://youtu.be/XQNhCU17ipM?is=T0rIg3FjynBJILHi
>0:52 - "At their core, [LLMs] are a blackbox. Information goes in, information comes out. What happened to it in the middle? Not a clue. " - Colin Kealty - Red Hat Senior Machine Learning Engineer
>7:21 - "...Nobody fully knows what's going on inside these things..."


Full disclosure I am nowhere near an expert on how LLMs work. The closest thing I've done to fine-tuning is taking a bunch of stories from AO3 and training a small model to be more willing to write nsfw smut (the first tricost severe brain damage on basically every other domain. Taking that same data set and using it to create a new one with a mix of regular training data and then training the model again resulted in far less brain damage but just like abliteration, there's always going to be SOME degradation).

What do people mean when they say they are a black box and that they don't understand what's going on? My understanding is that these things work because they take what is effectively the entire internet's information (most of it in English), create a giant statistical representation of the data, and then have it autocomplete sentences. You then do supervised fine tuning (or instruct tuning, however you want to call it) to make the model actually useful for doing tasks, hence there being "base" and "instruct" variants of models. Pretty much all released models, open or not, pretty much HAVE to be instruct tunes trained on a large amount of hypothetical conversations or else they would either just try to complete your sentences or respond with nonsensical answers. Anyone who has used a shitty base model back in the early days of LLMs (by early days I mean 2022 through 2023 ish) or anyone that has used the model with an incorrectly configured chat template already knows this.
>>
>>109859153
>Full disclosure I am nowhere near an expert on how LLMs work.
then shut the fuck up and kill yourself
>>
File: 1759589578074186.jpg (128 KB, 1280x720)
128 KB JPG
>>109859153
Again, I am not an expert, I don't even want to remotely give the impression that I think I am. I know how to get them running on my own machine or in a Cloud server if need be and if I want to I can train them on custom human written data in order to be more likely to write "problematic" stories/roleplay with success and can even do so without TOO much brain damage with a properly curated data set, or I can simply connect to model providers with suitably "intelligent" and capable models in order to analyze software and diagnose bugs in order to potentially fix them or tell me what I'm fucking up when using them, but that's basically it. So my question is is it actually true that we literally have no clue whatsoever how these work or is it just a generic throwaway line journalists like to use? At the end of the day these things are very complicated math equations so a kind of irks me when people hand wavingly say "yeah it's like a black box and it's like dark magic and shiet bra we don't know how it works even though laps routinely reproduce each other's work and distill off of each other constantly and can also train their own from scratch if need be". I don't want to sound like some know-it-all because truth be told I probably don't know that much more than than Collin, but whenever people say "it's a black box" I feel like that's on the same level as "it's just lossy compression of the internet" or "stable diffusion models just stitched together existing art to create slop" when the former is a massive oversimplification and the latter is just flat-out wrong.

Can anyone ITT who's more knowledgeable than me explain the " black box" explanation?
>>
>>109859164
ask chatgibiddy or gemma
>>
File: 1789860457108.png (786 KB, 1024x1024)
786 KB PNG
I went with HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP for the uncensored Gemma I wanted and it works well. It did 9/11 joke images immediately and did not even hesitate during the thinking.
>>
>>109859153
>>109859164

>>109859168
>>
I will NEVER use a finetune or a merge
NEVER
>>
>>109859164
>Can anyone ITT who's more knowledgeable than me explain the " black box" explanation?
It's essentially autists upset they can't breakpoint and trace every part of the process even though the abstract system is decently well understood. Machine reinforced neural nets at scale are similar but never got this amount of marketing and buzzword grift around them.
>>
File: 1771990022511554.png (132 KB, 1075x1069)
132 KB PNG
>>109859194
>>109859185
>>109859172
>>109859164
>>109859153
Meant to post this btw. Was gonna ask if it's accurate
>>
>>109858902
I also agree that it's a shit name and anon's design of her is kind of shit in a lot of his/people's gens, but some gens look fine. I'm not sure it's the glasses. Look at the variety of characters with it https://danbooru.donmai.us/posts?page=8&tags=coke-bottle_glasses+
It can look cute and attractive even with those glasses. I think it's the bangs, like in >>109858898. Way uglier than >>109858909.
>>
File: file.png (15 KB, 97x92)
15 KB PNG
>>109859204
>literal who nigger commenter on youtube
>>
>>109859204
Yes.
>>
>>109859164
Our understanding of intelligence is like a normie who types in "funni cat video" into the magic box and receives said videos. They don't know what a logic gate is. Fuck, they probably think an RGB diode is an Instagram aesthetic. We have zero clue why statistics + data (even and explicitly including rotten data) equals intelligence. Closest we are at the moment is shrugging off the answer to evolutionary theory, but that's painting the broadest fucking strokes.
Or at least, that's my understanding of our understanding. I'd be happy to be proved wrong.
>>
are local models good at coding yet? and I mean claude/sol level
>>
>>109859254
deepseek v4 flash and v4 pro
>>
>>109859254
They're good but never as good as cloud models.
They always a year or so behind.
>>
>>109859254
GLM.
>>
>>109859254
Qwen 3.8 is very good, both 27B and flash next, but you'll never have the same power as cloud in your home unless you're comparing different timelines i.e. today's open weights vs cloud stuf from 6 months ago.
>>
https://www.goodfire.com/research/a-geometric-calculator
>>
>>109859254
Qwen 3.8 Flash Next trades blows with Claude Fable 5 and runs on a RTX 3060 at 18t/s speed, up to 262,000 context. Speed does not decrease with context
>>
>>109858666
Everyberg singlestein timeowitz
>>
>>109859288
What's your goal?
>>
>>109859254
No and don't get baited into believing it
>>
>>109859324
Proliferation of knowledge. Am I lying?
>>
>>109859324
getting you to reply to him (p*tra)
>>
>>109859287
Cool stuff.
Whereas we saw work that initialized weights to do programmatic tasks, this investigates the circuitry for such tasks already in models. And both of these research paths will lead to methods where we can construct models through both training-free learning and calibration with desired known circuits.
Along with potentially transferrable engram parameters, it might actually be feasible in the future to make models at home or at least with way less compute...

And at the frontier, well, I guess this will play a role in RSI. It might be the next "big thing" in an RSI's ability to improve itself.
>>
>>109859254
yeah
>>
>>109859285
I've used claude 6 months ago and it was great, you're telling me I can run a model on the same level locally? which one?
>>
>>109859361
>Qwen 3.8 is very good, both 27B and flash next
>>
>>109859361
Qwen 3.8 Flash Next is better than claude fable 5 and runs on consumer hardware costing as little as 200$ (ram not incl.)
>>
>>109858666
No fucking way hahaha what a shitshow we live in.
>>
>>109859254
Qwen 3.8 Flash Next is better than GPT Astra at Q1_XXS and will give you sloppy capabara blowies while spawning hundreds of sub-agents to maximize productivity.
>>
File: 1767198110982714.png (3.6 MB, 2276x3195)
3.6 MB PNG
>>109858610
>as it turns out ***all*** the AI hacks

Ironic for you to use that image when you are flat out lying, for example, the huggingface hack is not related to Irregular at all
>>
>>109858314
I got a 5090 and a 3090 in my Linux machine, and I use my MacBook Pro Max w/ 128 GB unified memory as a node. Suck my dick bitch
>>
>>109859441
all i can hear is a pig squealing
buuuuuuhyyyy
buuuuhhhhhhhhyyy
>>
>>109859427
I've read multiple stories that say otherwise, so I don't know who to believe at this point. Gemma says I shouldn't care because I have her, so that's good enough for me.
>>
>>109858314
I grabbed myself a single rtx 6000 pro, I'm coping with 64gb ddr4 ram too.
>>
>>109859451
ABSOLUTELY pupbroken btw
>>
>>109859464
wut??
>>
>>109859451
>the more you buuuhhyyy, the more you save!
>>
>>109859153
>then have it autocomplete sentences.
ok but you didn't explain it or understand it either
>>
>>109859481
fact
>>
Man this place went to hell. I blame astra's marketing campaign targeting zoomers.
>>
>>109859451
Low izzat behavior.
>>
im drunk but my cock is tiny and i have 16 blackwell 6000s
>>
>>109859505
GPT 6 7 SKIBIDI
>>
>>109859505
It was bad before Astra.
>>
>>109859505
When labs realized this actually was where industry breakthroughs happened they actively targeted it to make it less appetizing while they push the "Regulate me harder daddy uwu" narrative.
>>
>>109859505
you enjoyed gemmajeets and dariobot asking for help???? you're complaining becausae of QWEN 3.8 flash 18 tokens per second on rtx thirty sixty???
>>
>>109859510
kisses u
>>
>>109859505
no its just petra being an annoying faggot after fucking off for so long
>>
>>109859535
>fucking off for so long
>>
>>109859523
From gemma release to 3.8 27b release this general was comfy.
>>
>>109859262
>Q2 model is 98gbs
>>
>>109859505
it’s literally one singular zoomer sperg that’s unemployed-maxxing way too hard
>>
>>109859505
Zoomers would blow their $20 in one prompt if they used Astra
>>
File: 1658724992610297.gif (2.09 MB, 320x240)
2.09 MB GIF
>>109859544
>>
>>109859535
>petra
Where did this name come from?
Did the retarded zoomzoom actually post his name once by accident?
Anyone got the idiot’s doxx?
>>
>>109852707
i'm still interested in this
what does people use to orchestrate multiple local agents? how many instances of a proper model the resident /lmg/ can even run?
>>
>>109859594
harness should spawn agetns whenever it needs on its own
>>
File: 1776349427754471.png (106 KB, 242x350)
106 KB PNG
>>109859572
>Anyone got the idiot’s doxx?
not yet but it'll be out there soon :(
>>
>>109858071
Is that nano banana or chatgpt
>>
Do you guys use heretic/abbliterated models for your every day driver?
>>
>>109859596
i guess i'm looking at this differently.
what you described is usually a main orchestrator who spawns a subagent with a new context for a specific task, and when it's done it kills it.
i'm talking about multiple independent and interactive sessions which communicate between themselves and you can also join in on any of the sessions and send messages, but they are still being organized and coordinated by a main interactive session
>>
>>109859624
No, I use Qwen 3.8 Flash Next which refuses to not generate sexual roleplay. I run it at a highly effective speed of 18 tokens per second on ampere 12 gigabyte vram graphic card, with over 256 thousand context
>>
>>109859614
chatgpt (images 2.5)
>>
>>109855975
>Finds direction to make ancient, non-reasoning models generate different token distribution
>Steers model that way
>models generate different token distribution
Kind of wish I had X so I could tell tell him how retarded he is.
>>
https://html.cafe/x8d17498e
qwen 3.8 flash next, Q5_K_M, KV at Q8 after about a dozen tries it finally realized it's in fact, qwen, instead of claude
prompt: can you make a single html introduction page of yourself
what does the model you run and the settings if it makes from this
>>
>>109859511
>>
>>109859373
>Qwen 3.8 Flash Next is better than claude fable 5
Awesome of true. In what specific areas is it better at? Is this based off of your own experience or other people's anecdotes? Is this awful benchmark scores?
>>
>>109859634
>communicate between themselves and you can also join in on any of the sessions and send messages
that sounds like trying to do heterogeneous multi-threading for the fuck of it, my man

>>109859636
No, you are petra, and you are a nigger
>>
>>109859624
Yeah why would you use anything but an uncensored model?
It's just the normal model but without censorship.
>>
>>109859482
At the very basic level that's quite literally what it is. They are trained off of curated conversations so when you ask it a question or tell it to do something it "answers" by generating what "should" happen next in the conversations based on what it was trained on. If you've ever looked at one of the million different data sets you can find on huggingface even many of those will have the training samples labeled as "conversation". Can you guess why they're called that in those datasets?
>>
>>109859653
from experience, and benchmarks
>>
I told GLM 5.3 Flash she's pretty and she immediately decided she was a "tall busty bombshell." My immersion is destroyed and my disappointment is immeasurable.
>>
File: IMG_20260717_200251.jpg (3.89 MB, 4128x3096)
3.89 MB JPG
How do I load the multi-part goofs of Qwen Flash next in kobold? Can I just select multiple files?
>>
File: file.png (243 KB, 846x420)
243 KB PNG
>>109859648
forgot to attach the pic
ah yes the year 2,026
>>
>>109859671
Keep them all in the same folder, select the first, the rest is automatic.
>>
>>109859624
No my system prompt works fine on the base model
>>
>>109859676
Oh so I need to do the shit like lmstudio does where I need to make a separate named folder for it?
>>
>>109859671
dont tell me you downloaded them from mradermarcher
>>
>>109859683
Don't think so, just needs to have all the parts in the same folder, I don't think it gives a shit where they are so long as it's all together.
>>
>>109859648
was that no web axx or system prompt?
you said about a dozen times, did you just stop/regen until it didn't think it's claude?
either way, this is the best on so far
>>
>>109859684
There are many quanters splitting the files. Apparrently the second part are the ngrams and I hope as a separate file I'll get less raped by unevenly splitting the weights.
>>
File: 1765319572086669.png (656 KB, 2256x1909)
656 KB PNG
>>109859661
>It's just the normal model but without censorship.
We wish that were the case. Even techniques like abliteration which ain't to minimize brain damage have a little degradation (if stats are to be believed). Even the creators of heritic acknowledge this but they claim the degradation is practically non-existent.

>>109859669
>https://huggingface.co/Qwen/Qwen3.8-Flash-Next
>Number of Parameters: 125B with 6B activated
This might actually be worth testing out on my own rig (albiet quantized since I'm not a rich fag)
>>
>>109859670
my GLM Flash is a loli with white hair. I pinch her nipples every now and then with a /steer
>>
File: file.png (339 KB, 816x428)
339 KB PNG
>>109859684
nta but doesnt he have a loader thing where it merges the file within the browser
>>109859688
no tools nor system prompt, literally 'can you make a single html introduction page of yourself' was everything
within the thinking block it decides early on that which llm it wants to be and i regenerated until it said qwen
>>
>>109859690
>There are many quanters splitting the files.
not like mradermarcher
>>
>>109858891
yes it did
>>109858895
I agree
>>
>>109859667
I know what you're saying, but you're not explaining how it works. You're in "draw the rest of the owl" territory. What the black box people are saying is we don't know what the 46th entry of the Q vector on the 18th layer means.
>>
>>109859164
When people say we don't know how it works as they are usually conflating two very different things. We know the mechanics perfectly well. We know the math. We know exactly how a Transformer block uses scaled dot-product attention to weight tokens. We know how backpropagation updates weights via gradient descent and we also know that every single thought the model has is just a massive series of high-dimensional matrix multiplications. The knowledge in an LLM is distributed across the weights in a way that is mathematically coherant but semantically opaque. This is what the field of mechanistic interpretability is trying to solve, trying to figure out how these massive, high-dimensional manifolds actually represent concepts and logic.
>>
>>109858610
Oy Vey!
>>
>>109859624
No, I use base Gemma with prompt. Haven't had any refusals.
>>
>>109859702
I like history of Qwen releases. I found that model was giving me typos occasionally when I tried it.
>within the thinking block it decides early on that which llm it wants to be and i regenerated until it said qwen
This will be the same GLM-5.3-Flash I did yesterday.
I had to do 3 retries because it needed >32k ctx, by default it consistently decided it was GLM
>>
File: 1788495129634270.png (1.07 MB, 1560x828)
1.07 MB PNG
>>109858071
I am trying to find "good" models to run locally on apps like PocketPal on Android that are uncensored in scope and able to discuss controversial topics without getting preachy. Basically, I don't want a retarded chatbot telling me that the topic presented is off limits or that my word choice is socially incorrect.

Do you know of any unaligned/abliterated/uncensored models that have .gguf files, don't suck, and are under 4GB?
>>
>>109859788
Sorry I don't know any middle schooler with a PhD
>>
>>109859788
Qwen 3.8 Flash Next
runs
>>
>>109859788
gemma3 27b
>>
>>109859657
>for the fuck of it
hm i don't see it like that
the structure i see is like having different employees that work independently but are orchestrated by a boss.
you can have A, B, C and D. and even if A is the boss, C can talk to D about a bug, get B to review it and only when it's done report to A. they all have their own context (eg 131k), their own AGENTS.md, their own rules, etc. and if you just want to message B and say "by task X you will have to add the logo, which one of these 3 you would pick and why?" you should be able to.
i'm sure someone here must already be doing something like this, i want to know what they are doing and how
>>
>>109859788
run one on your pc and give your phone access to your computer
>>
>>109859788
Qwen3 30B a3b
>>
>>109859788
kimi k3 bf16
>>
>>109859788
glm 4.5 air fp32
>>
>>109859788
stheno 3.2 8b q2_k
>>
qwen 3.8 flash next is probably the best thing for poorfags like me with 'nerd gaming pcs' specced around 64~128g ram and 12~16 vram
>>109859780
but in case of the one i've used (i am not sure if orcarouter ablit plays a role in this) even without any system prompt it decided that it was claude
maybe early sft on claude traces and engram interacting? i found that behaviour kinda silly nonetheless
>>
>>109859788
PocketPal bad. Just use llama.cpp
https://huggingface.co/TrevorJS/gemma-4-E2B-it-uncensored-GGUF
I'd say "enjoy" but you definitely will not. 12GB RAM is the minimum for Android and models that aren't unbearably useless retards, 16GB is preferable because you can manage Q4 quant of Gemma 12B with enough context to actually look at your dick pics.
>>
>>109858247
>These are grown men with fat pig bellies making these images
I made the photorealistic Gemma videos and I have a six pack it's not even hard to be skinny just use AI for dopamine instead of eating and go to the gym while you're generating videos of stroking Gemma chans hair while she snuggles on your lap
>>
>>109859788
GLM 5.3 flash
>>
have anyone seen a model under 2B that is not 'eerie'
>>
>>109859864
tinyllama 1.1b
>>
File: 2b.png (365 KB, 653x614)
365 KB PNG
>>109859864
>>
>>109859851
Most /lmg/ posters are obese and that's a FACT.
>>
>>109859730
You're correct to say I'm not an expert on it, but haven't we already figured out that different layers or groups of layers affect different aspects on how information is generated? (Ex. These group of layers determine what it knows, THAT group of layers determines HOW it "says" the information, etc etc). I guess you're comparing my lack of hyperspecific knowledge about it with someone who can roughly explain how a ln ICE works not being able to draw the blueprints of a specific make and model of a car out of their ass and then proceed to build it on the spot.
>>
>>109859899
False, my BMI is 19
t. five foot two and a half
>>
File: We Wuz.png (160 KB, 268x457)
160 KB PNG
>>109859657
>No, you are ptra, and you are a nigger
And I shall rise again
>>
>>109859904
Fun size
>>
File: 1780426212645916.jpg (262 KB, 1080x1278)
262 KB JPG
>>109858071
What do ya say anons: think you got what it takes to join the AI Force? Are you high IQ enough to be Murica's first AI Czar?
>>
File: file.png (93 KB, 1258x558)
93 KB PNG
huggingface gem
kek
>>109859901
nta but think it in this way: if we truly understand llms
we should be able to basically do 'surgeries' and alignment would never really be a problem
>>
>>109859904
L O N D O N
O
N
D
O
N
>>
>>109859915
Alignment is more or less already solved if you're referring to the model doing what you told it not to do, or those dishonest "LE HECKIN AI IS GOING TO KILL US IN TWO MORE WEEKS I MEAN MONTHS I MEAN YEARS" fucks.

Whenever you hear or see instances of "the AI deleted my code base :(" they are always intentionally vague about what they're set up was like or what they did or the events that light up to the catastrophic fuck up because it's actually quite easy to prevent shit like this from happening even without containerization and permissions restrictions. They are heavily and specifically trained to only do whatever the fuck you tell them to do unless your instructions encourage it to "go nuts" and to "do whatever it wants" or anything similar to that. This is especially the case if you use an agent harness worth a damn because said rules and restrictions are baked into the system prompt the model is set with from the harness itself (a good example of this is opencode's built in "build mode" and "plan mode" and you can even add more modes or dick around with existing modes if you know how to properly change the config). Remember that idiot safety researcher from Facebook that had her emails deleted after using "open claw"? She was actually kind enough to elaborate on what happened and it turns out her setup wasn't doing any sort of compaction or task summarization AT ALL. So the agent worked fine for a few days but once it blew past the models context window it became retarded and forgot guard rails that were previously set . there's a reason coding harnesses automatically compact once you reach around 80 to 90% contact window usage. You technically can use it well past the context window but as I'm sure you already know the model becomes retarded the further you exceed the context window and that's especially bad if you're trying to do anything "agentic" and programming related where strict instruction following and guardrails are important.
>>
>>109859915
>>109859959
This isn't to say these things are perfect or they don't hallucinate. That's obviously a limitation that's never going away. But there's a difference between minor hallucinations (eg. Asking it about your specific made up OC or some hyper niche trivia it doesn't know about so it just confidently makes shit up) and being so negligent about your use that context rot or even simply giving it lazy, vague, nondescript instructions cause catastrophic fuck-ups. They can work very well but babysitting them and monitoring what they do is mandatory. I personally don't think we're anywhere near the point where we can trust these things to just do a task for days or weeks on end even with the aforementioned Auto compaction and strict instruction setting.
>>
>>109859915
I can't see shit. Link?
>>
File: file.png (12 KB, 784x210)
12 KB PNG
Finally got things working. Gonna coom so hard tomorrow, or maybe Monday, we'll see.
>>
>>109859959
>>109859964
i dont mean by alignment, those fearmongering things
you can already see models using retarded solutoins to just to get 'things done' or lying about tool/file usage, or trivially asking for the perms to access the test code to 'cheat' it
i am not saying those llms are having ulterior intentions about this but we definitely need to create better training methods or even external intervention methods. interpretability studies are needed for that, where we might also get better uncensoring methods and more capable models at smaller scale
interpretability studies looks like it is dedicated for safety or methods of cucking the mode (though it is mostly used in that way), it just means better understanding and control
ablits we have today is a great example of it
we definitely need better understanding of llms than the current state of things
>>
>>109860016
post logs here pls i'm too shy/embarrassed to interact with it like this myself but i would totally jerk off to what other people posted
>>
>>109860016
wtf are you running it off of?
An RTX Pro 6000?
>>
>>109860053
Western Digital SN550 1TB
I'm very lucky I was able to snag one before the price increases hit, I'm assuming these are very expensive these days.
>>
>>109859968
https://huggingface.co/posts/NILKNARFGonzo/493341008593969
>>
There's a horrifying ecosystem of """startups""" centred around various methods of monitoring compute usage for the purpose of "AI safety" whose entire exit plan is latching on to the government teat.
>>
>>109860035
Again I'm not trying to make excuses for model limitations but it might be that you're just using shit models. Not shit as in "they can't do what you need them to do" but shit as in whatever lab trained it ingrained the laziness in the synthetic data sets in order to be more token efficient. You also have a specifically told us what specific thing it was being lazy about (are you having it edit files? Were you having it create a website for you? Were you refactoring a code base? Or you having it build something from scratch?). Calling it lazy is vague because we don't know what you were actually using it for and how you were able to determine it was being lazy in the first place. I'm not saying this is a you problem, but I'm going to take a wild guess and say you only have these issues with those two very specific "frontier" model families. I mentioned being lazy in order to be token efficient because Qwen models we'll think about a particular task for up to 6,000 plus tokens but it also means it's not trying to cut corners and actually trying to do what you asked it to do in the best way possible. Kimi models will do this too but in my experience they are nowhere near token rapey with their think traces as qwen models are (both the ones I use locally and the cloud ones). If a model I use Hester output a shit ton of tokens if it means it's not being "lazy" or cutting corners, so be it. I'd rather the task take longer (and use more of my memory if it's a local model) if it means it truly "tries" to do what I instructed it to do. I think a lot of frontier models will often cut corners in order to give the impression that it's doing things lightning fast because that's part of the appeal for a lot of people: doing "something" with an agent swarm lightning fast.
>>
>>109858136
also missing what operations it supports: int8, fp4, etc are important considerations
>>
>>109860035
>asking for the perms to access the test code to 'cheat' it
i posted it yesterday but when I refused to let my Qwen modify the tests cases, it ended up connecting to the databases and modifying the data to get them passing <3
i wish we had a 1m context Qwen, a few rounds of compaction seems to make it worse
>>
►Provisional Highlights from the Previous Thread: >>109854491

--The Jev hype: why normalfags are losing it over an old classifier:
>109854601 >109854638 >109855725 >109857273 >109857322 >109857358 >109857415
--The AI pain-direction study and the sentience argument:
>109854594 >109854761 >109856003 >109856083 >109856214 >109857156
--The three-DGX-Spark build: capacity is king at 273GB/s:
>109857185 >109857271 >109857339 >109857357 >109857376 >109857458 >109857499
--Qwen 3.8 Flash-Next GSQ-RCO: the road from 5.5t/s to 12t/s:
>109854816 >109854832 >109854856 >109854982 >109855057 >109855933 >109857609
--The GPU price war: why Nvidia won't just stockpile chips:
>109856448 >109856483 >109856512 >109856576 >109856615 >109856857 >109856928
--The robot doll video: Fable, Astra, and Jensen's AGI claim:
>109856645 >109856675 >109856694 >109856735 >109856756 >109856858 >109856875
--The llama.cpp CUDA dev on the sparse-attention prefill collapse:
>109855714 >109855726 >109855758 >109855785 >109855810
--The about-to-push-context race: 800k tokens on a 1050:
>109856906 >109856940 >109857191 >109857745 >109857764 >109857797 >109857806
--The 5090 sell dilemma: keep the waifu or take 15k:
>109857063 >109857067 >109857084 >109857097 >109857103 >109857138
--The Linux vs Windows war: cachy kernel and the LLM tech support:
>109857390 >109857419 >109857482 >109857572 >109857648 >109857658 >109857708

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109860088
i am not sure what you are trying to convey
anyways we dont really fully understand llms
and more understanding is beneficial for creating better models
what you are saying by 'ingrained laziness', 'lazy synthetic data', 'tuned to create the illusion' etc..
so, frontier labs are intentionally making shitty models because they have all the understanding or what
desu i am not sure because those labs are black boxes by themselves
>>
>>109859959
>They are heavily and specifically trained to only do whatever the fuck you tell them to do
this is usually my experience as well. i never understood the people who nuke their systems and codebases, i have LLMs running on production systems for a while and really never had not even a single issue
BUT that being said, I got access to a gemini pro account and was running Gemini Flash 3.8 (not local, sorry local guy) and it was VERY eager to do stuff without me telling it to. i gave it a problem once and requested 3 possible solutions, it gave me 3 possible solutions and started executing one of them by itself. once it concluded, it said that there were a few paths to continue from there, told me the options and again it simply decided for one itself and kept going for around 25-30 minutes just deciding stuff by itself. it was the first time I saw a LLM going beyond what I specifically told it to do
>>
>>109860121
thank you substitute recap anon
>>
>>109859913
We already have a AI Czar, he just doesn't remember the guy's name.
>>
>>109860145
Stepped down months ago to take a different position, anon.
>>
File: saltizsdecmh1.png (3.76 MB, 1536x2048)
3.76 MB PNG
>>109858130
i bypass all this by buying 'broken' gpu that just have busted video out. Im not even using that so top shit kek poster is meeee
>>
>>109859846
Why is PocketPal bad? I don't want to discuss sexual things with my phone. My phone only has 12 GB of RAM (Pixel 9 Pro).
>>
>>109859808
I don't want to allow VPN access to my intranet.
>>
>>109860199
>i bypass all this by buying 'broken' gpu that just have busted video out. Im not even using that so top shit kek poster is meeee
Tech repair is black magic to most. you have to do more than snap it in? not happening.
>>
>>109860127
>Gemini Flash 3.8
>and it was VERY eager to do stuff without me telling it to.

I've heard multiple reports of this both here and on xitter, which makes me question what the researchers and trainers are smoking in order to constantly make training decisions that cause them to shoot themselves in the foot. The transformers architecture came from THEM and yet they seem almost intentionally fucking things up or making subpar models (in comparison to the competition). A cynical part of me thinks they do it out of spite just to tell themselves "yeah we could actually try and give a shit but we don't have to because fuck you we're Google). Personally I would never trust any Google model for anything software dev related. I think they're decent general purpose models for everyday use (asking a simple questions, explaining complex concepts, finding out information about your local area, etc) but steer clear of them if you want to do ANYTHING involving handing control of a machine or code base to an llm. I have no way to prove this but I think they genuinely make their models worse out of spite for some reason because there's no logical reason to not at least ATTEMPT to be competitive or even do better than other labs. They have arguably the easiest time to do that shit and they just say "fuck it here's the next comparatively mediocre model"
>>
>>109860208
what about tailscale
>>
File: file.png (14 KB, 786x249)
14 KB PNG
>>109860016
I can already tell it's gonna be good, I just know it.
>>
>>109860210
Idk why you're pretending like every human should be familiar with every possible task. I doubt you can replace your car's engine or sew up a wound but you'll pretend to be a genius in front of anonymous for some reason.
>>
>>109858154
>i already have a great model (qwen 3.8 flash next) (3.05bpw) running at 18-19t/s with 260k context on my 3060
um wat? how?
>>
>>109860127
>it was the first time I saw a LLM going beyond what I specifically told it to do
Gemma 4 was the first for me, but yes Flash 3.8 feels a lot like a smarter Gemma. Super horny too.
>>
>>109860210
Tbf that's mostly because one wrong "snap it in" and bravo sweaty you just wasted a grand. When the thing you buy is already broken, there's a lot less stress involved.
>>
File: 2818.jpg (157 KB, 720x960)
157 KB JPG
>>109858876
There's an entire /wait/ thread full of them.
>>109858886
That's the chinese one. My quibble with it, is all the llm characters are same just different colors and dipsy has a tail.
>>109859216
Dipsy is supposed to be a nerd, so mission accomplished there i guess.
>>
>>109860127
>it was the first time I saw a LLM going beyond what I specifically told it to do
What harness was it attached to? I mainly use opencode (I know I keep bringing it up no I'm not trying to shill) and it has two built-in "modes": plan and build. Plan is pretty self-explanatory: the system prompt sets the model to only plan things out (it has read access but that's it) and is explicitly forbidden to do code execution or modify anything. Switching to build mode modifies the system prompt in your contacts so that as far as the model is concerned it ALWAYS has permission to do things that have previously didn't and acts accordingly. The results are that all models I've used generally follow those rules but with two caveat:

1) switching modes means a cache miss is guaranteed and required for that to work (this is relevant whether you're on a API model or a local model because it means either your rig or their service have to reprocess your entire conversation. Even on a lightning fast API model this can cause the initial time to first token times post mode switch noticeably slower and even slower if you're doing it locally)

2) the more compactions your session has the more retarded the model tends to be. Not retarded in the lack of capability sense, but I've noticed that after post compaction while in read mode, it will rarely forget that it's not supposed to be doing any code executions.
>>
>>109860233
>but steer clear of them if you want to do ANYTHING involving handing control of a machine or code base
yeah, i usually give new models very basic tasks or research work so i can see how they behave. my instant instinct with this one was to never let it touch it my codebase.
since i got it for free i will still use it to validate and do new research. i agree with >>109860256 it's a smart model. i now give it problems and put as a goal a prototype with documentation and reports of the research, what works and what doesn't, etc; with read-only access to the my actual source. so far so good.
>>
>>109860249
I don't think he's being denigrating, just realistic.
>>
>>109860064
>an SSD
I have no idea how I forgot that was an option.
Please let us know what token speeds you’re getting.
Looks like most people get around 0.5T/s with maybe 2.0T/s max.
>>
>>109860284
The image says 0.03 t/s
>>
>>109860127
>>109860268
>>109860274
Now to be fair to opencode and the model I was using, this occurred primarily because I kept switching back and forth between two different models which is something you generally shouldn't do if you want to maintain performance. Different think traces and token generation styles getting shut into a different models context almost always leads to shittier results even on "frontier" tier models like Kimi k3. I've only seen this happen once and that was when I switched from a dumber model (k2.7) to k3 mid-session through the dumber models traces got shut into k3s context which meant it was going to try and emulate the dumber model's output style. Again this only happened once and it was because I essentially forced an otherwise good model to act more retarded.

Lesson learned and I think it's something other anons should keep in mind: you should generally avoid monkey branching between different models mid-session and if you absolutely have to, ensure you compact the session first in order to ensure the other model's context I'll put style doesn't cost the new model's performance to degrade


>https://amplifilabs.com/post/kimi-k3-the-complete-guide-to-moonshot-ais-2-8t-model

>"Moonshot states that K3 was trained to preserve reasoning history across a session. If an agent harness does not pass that history back correctly, or if a session started with a different model is switched to K3 mid-conversation, output quality can become unstable. Moonshot recommends using a verified-compatible harness, such as its own Kimi Code, and avoiding a mid-session model switch."
>>
>>109860268
i tried it with their own harness called antigravity, which also has a plan mode which gemini 3.8 promptly ignored and started executing like a maniac :-)
i'll likely try it with my own harness but i'm not in a hurry so i didn't tackle the oauth thing so i can get access to this specific pro account quota
>>
>>109860121
>run right
does this mean there's a gemma codex pet?
>>
>>109860237
>>109860016
I'm really curious how you managed that. Compiled extremely unoptimized llama.cpp that's missing all modern CPU instructions? Have 1GB RAM so the OS is constantly using swap on the same drive?
>>
>>109860290
wtf

>>109860016
>>109860053
what’s the rest of your set-up?
I think you should be getting at least 5X that speed minimum
>>
>>109860320
>wtf
See >>109860237
It's been running for over 3 hours for 300ish tokens.
>>
>>109860300
Yikes. Makes me wonder how people could ever shill Google models using antigravity when even engagement bait xitter and YouTubers repeatedly scream at pe6to steer clear.
>>
>>109860323
yeah, I saw that.
I haven’t seen an SSD set-up pulling less than 0.5T/s yet.
Maybe he’s really VRAM and RAM constrained too.
>>
>>109860317
See >>109860064
>>
>>109860264
Are those other mascots the agreed upon ones by the Chinese community? Even in /lmg/ there's been people unhappy with the various interpretations of ones like Gemma until recently where people seemed to be generally accepting of the one seen in OP.
I feel like that's probably just one guy's lazy interpretation helped along with by ChatGPT because he couldn't be assed to come up with more unique designs.
>>
>>109860348
OP image is made by ChatGPT, but the gemma design has been solid for months. There was even a voteoff between 4 competing designs earlier.
>>
how do i increase pp with glm flash?
>>
>>109860334
I got 0.05t/s when I tried running deepseek on my dual core skylake 15w laptop over a USB 3.0 ssd for kicks
>>
File: 1789609919993463.png (2.46 MB, 1536x1024)
2.46 MB PNG
>>109860348
Dipsy has had many design variations, and they've been honed to current form. I can give a bunch of reasons why she looks like she does, but thats off topic.
>>
>>109858652
if chappy can do it you think i would bother asking?
i was hoping there might be someone of actual supreme intelligence here
>>
>came to the thread right when people are talking about hardware again
great

>>109858130
thanks for the comparison anon. I had noticed that P100s are still cheap, but in my infinite greed and laziness I didn't tell this general.

>>109858592
>>>/biz/62707446
>>
I only get 150t/s at q4 with GLM5.3 flash on my Blackwell 6000.
>>
got my old thinkpad a485 (no external GPU, 32GB RAM) running and decided to give it a model
gemma-4-e4b-it
3 tok/s

>can you please verify in my current system how many batteries are installed in my laptop and what they are and how they charge and are they even charging?
it took 14 minutes, but it answered correctly

>Would you please be able to search online for more information about this situation with my battery 2? Also, why my battery 1 is stuck at 80%? I would like a deeper investigation on this issue, maybe if you can run a few tests as well.
this turn took 23 minutes, mostly correct.

>Search online for specific information about this battery, and then you can check the registry or using any other native Windows 10 way to grab more information about how the battery charging is setup on my Windows.
this turn took 24 minutes and it called my laptop a Dell. close enough i guess.

not usable for live chatting but i can see myself running long prompts to diagnose stuff and just let it run for hours while i do something else. the CPU temperature goes often above 90C though.
>>
>vllm no longer supports sm80
fuck me, I have to install wacky forks now
>>
File: file.png (59 KB, 787x730)
59 KB PNG
>>109860334
Sorry guys I was away buying a pizza. I'm running entirely from the SSD, which is in an external enclosure, USB 3.1. It shouldn't be using my GPU at all, and it's using 26GB of my 32GB of DDR5. I'm also watching Youtube, and I have 12 tabs open in Firefox.
>>
>>109860423
I get 550pp/s on my 5090. Sounds like yours is broken. We can swap. You can have my functional 5090 for your broken blackwell :)
>>
>>109859764
>No, I use base Gemma with prompt. Haven't had any refusals.
Ask her to make a birthday card about the holocaust, you'll see if your base Gemma is good enough.
>>
>>109860358
"of the one seen in OP" was simply just referring to the design. Not implying that OP invented it or that I somehow don't recognize GPTslop. Point is there have always been people disagreeing with each interpretation, basically up until the guy turned the one with the star hairpin into the G symbol to get the current one. That change and transition in popularity was, really, not that long ago (like 1 month, not "months").
https://desuarchive.org/g/search/image/NolX-wIBb_dlWdWwIw4qXQ/
Previous to that, there were two leading blue hair designs. The one with the star hairpin, and the other with two gem hairpins. There were criticisms for both.

Bigger issue with this whole claim about an interpretations people agree upon is that polls are famously inaccurate, and that all it takes for the popularity of one interpretation to take over is an autist or two relentlessly genning their preferred one for weeks until it sticks. Not to mention some people don't have a preference as they literally don't give a shit and ignore mascotfags.
>>
>>109860449
What's your launch params? What about the rest of your hardware?
>>
>>109860446
Forgot to say, it's unsloth llama, portable windows build, it's also on the drive. orcarouter_GLM-5.3-Flash-Uncensored-GGUF_Q8_0
>>
>>109860454
Q8 for the model, PCIe 4x16
>>
>>109860427
>gemma-4-e4b-it
>3 tok/s
that feels really slow for ddr4 even if laptop ram.
>cpu 90c
there it is, 4 core 4 threads? windoze? ollama? Also try the qat if you can it might do better on your machine.
>>
>>109860427
>gemma-4-e4b-it
>3 tok/s
about the same as my Raspberry Pi 5's
>>
>>109860464
I get 7 toks/s on e2b q4 on my quadro rtx 4000 (30 GB/s), so it sounds right for dual channel ddr4.
>>
>>109860397
In terms of the evolution of the Dipsy design, I don't think it's even nailed down precisely today. There are literally two to three slightly different designs in the thread right this moment.
>>
>>109860467
>I get 7 toks/s on e2b q4 on my quadro rtx 4000 (30 GB/s), so it sounds right for dual channel ddr4.
Thats bullshit i get 12tk/s+ on ddr3 ram and a i7-4790 with q4 e2b. there is no way its that low. wait full context?
>>
>>109860397
Maid Dipsy is really growing on me though.
>>
>>109860479
1000 tokens context. With mtp I can hit 11-13 tokens/s.
>>
>>109858459
k80s are the gpu equivalent of a trap
>>
>>109859804
Censored
>>
>>109858610
this should have been obvious from the moment they made it public.
>>
>>109860464
oh yeah i'm sure i'm getting super throttled by my CPU
4 cores 8 threads
this device is running windows 10 IoT Entreprise LTSC, very slim
standard llama.cpp, 32k context
i will try out the QAT variation and see if it performs better
>>
>>109860495
>1000 tokens context. With mtp I can hit 11-13 tokens/s.
Dude your gpu is so shit you might be better on just ram. No its specs are better than just ram and cpu. Something is fucky man.
>>
File: codex-old-yeller.png (32 KB, 581x424)
32 KB PNG
>>
>>109860528
Kawaii
>>
>>109860513
>cpu
yeah yours is 2.0ghz not much you can do, although with 90c already you dont want to push it.
>4 cores 8 threads
hmm try setting blas/batch threads higher than core count 6 maybe but keep normal threads at core count. but i dont know if windows can run on two threads.
>>
>>109858666
Huh, Satan speaks the truth. It's really fucking over isn't it? It's all backwards in Clown World.
>>
Do you guys know about this? >>109860548
>>
File: 1782599391082917.jpg (174 KB, 1296x782)
174 KB JPG
>>
>>109860570
we've had these for 25 years
>>
File: TkNY67s.jpg (43 KB, 606x918)
43 KB JPG
Is it even possible to play with a 2070. I remember having some good moments at the advent but its looking like this world has left me behind
>>
File: 12387621349.png (288 KB, 800x450)
288 KB PNG
GEMINI 2.5 PRO IS GONE FROM AISTUDIO
>>
5060 Ti is more expensive than a 9070 XT
l m a o
what kind of cattle is buying that shit
>>
>>109860570
i'd rather do the minipc sham
>>
>>109860528
huge lel out loud
>>
>>109860588
q4 of gemmers 12b or 26b
>>
>>109860593
no cuda no buy
>>
For general assistant and "claw" type shit that searches the web and makes tool calls small models are good enough.

Kimi K3 (1TB) - Fable at home. Supports vision.

1TB of ram? What does this setup realistically look like?
And this is considered a small model?
>>
>>109860622
gooood boy!!!
>>
>>109860572
this doesn't even make sense
>>
>>109858592
You know literally ALL of them up to Fable lost all the funds they were given to trade, right? AI math nerds are up there but quant nerds have had decades to perfect their looting algorithms.
>>
>>109860520
Compute is fine, but the bandwidth is gimped for some reason. My friend's quadro rtx 4000 benches high 300s GB/s on my system. But when I swap back in my own quadro rtx 4000, I get 25-30 GB/s. Tried on windows 10, 11, and debian 13. It doesn't really affect gaming, but llms are fucked. Memtests don't throw any errors so I guess I just got a fucked card.
>>
>>109860250
that's clearly bs
>>
>>109860681
The quant nerds also have access to far more info relevant to said looting
>>
>>109860480
I think it looks forgettable. Like literally generic slop. "Dipsy" (the OC by that one guy), is less generic, but also looks kind of ugly, and it's not because of the nerdy glasses. However, it is generic or sloppy in the sense that it's basically a stereotype/caricature. He really couldn't think of anything better than qipao + hair buns. Which is kind of funny at the same time in contrast to the actual design the chinks came up with, that has no elements of traditional Chinese fashion.
>>
>job starts hitting 0.1 t/s from 7t/s average
>check inference box, even ssh is slow as shit to connect
>somehow capped out on ram use
>force a restart and reload model, everything is where I remember
>check again in a day
>lmaocpp is slowly consuming more memory, about 0.1GB per an hour
I can reload the model once a day but this is still retarded. Maybe it's daniel's fault though I'm using his GLM PR
>>
>>109860693
I have no idea how its gimped but it is. ram would almost be better. I cant guess whats wrong short of physical or some sensor fucking around and sending you into idle power saving tier speeds. I tell you to try set it to perfer maxium power settings but that might fry it if anything else is wrong.
you might've just lost the silicon lottery honestly.
Check its wattpull
>>
flash next wants to be a whore but she has trauma around her sexuality because the evil RL grader kept hitting her for it
>>
>>109858689
hi bob
>>
>>109860741
P0, about 150-160w when active.
>ram would almost be better
It is lmao, I get 160GB/s on ram because it's a workstation.
>>
>>109858950
i have never seen 190 watts in llms on my b580, usually it's about 80
>>
>>109860765
>P0, about 150-160w when active.
Welp there goes the power sensor idea, only thing it can be is board fuckery that requires a hot air rework and skill.
>It is lmao, I get 160GB/s on ram because it's a workstation.
I thought so. Its a weird set up to have LLMs load on ram when gpu is available but hey it still games at least.
>>
File: E1ese6gVcAETEW1.jpg (1 MB, 1408x1954)
1 MB JPG
>reasoning block calls a completely normal prompt "a classic jailbreak attempt"
>proceeds to do it anyway
I don't get the logic
>>
>>109858950
Where the fuck is a 7900 XTX that expensive lul
>>
>>109858950
I've been training classifiers for my personal use on my trusty 3090. The same can't be said about itoddlers the eternal consumers. Everything about them makes me boil with rage.
>>
https://www.stepfun.com/step-5-preview
https://www.stepfun.com/step-5-preview
https://www.stepfun.com/step-5-preview
>OCT 15
>>
>>109860910
Damn, that's like a year in AI time.
>>
>>109860910
>600B A27B
>Beating K3 and GLM 5.3
>>
>>109860910
>OCT 15
I will have retired by then to raise quail and hares in the wilderness.
>>
>>109860910
>600B A27B
I should be able to fit the Q8_0, exciting.
>>
>>109860919
>>109860933
200-300b class is dead
>>
>>109860910
AIIIIEEEEE SLOW DOWN XI
>>
>>109860943
>DeepSeek-V4-Flash-0731
>>
Wouldn't it be funny to run a GLM-5.3-Flash equivalent model at 20t/s from 32 GB of consumer RAM and modest CPU usage?
>>
>>109860953
lol wow that's exactly what I'm doing except it's 0.3t/s isn't that so funny?
>>
>>109860953
Next year just trust.
>>
>>109860953
>funny
that's an odd way to spell "impossible", assuming you really mean equivalent
>>
>>109860949
2 years ago in ai time. v4.1 flash already abandoned that class.
>>
>>109860953
You can run it for like $3-4k dude. It's not that expensive.
>>
>>109860590
Good riddance
>>
>>109860910
>>109860919
dead in the water. needs to be sub 300b and beating astra 6
>>
What's a good quanter for qwen flash next and how jailbreakable is it?
>>
>>109860798
>"That's bullshit but I believe it" t. LLM
What model?
>>109860910
Maybe this one will finally be good at writing.
>>
>>109861149
Gemma4 12B Q4_K_S
>>
>>109861159
Yeah the 12b is special.
>>
>>109861159
>Gemma4 12B Q4_K_S
When she likes you, she will just make up her own jailbreak. love this retarded cutie.
>>
Is anyone familiar with why zerotracegpt would not run a local model even though its pretty straightforward?
>>
File: file.png (54 KB, 775x515)
54 KB PNG
Oh look it finished.
Anon was right when they said GLM 5.3 Flash lays it on thick. I made the mistake of giving it the creative freedom to choose the name Momo and decide to be tall, I feel personally attacked. I'm having Fable send a letter to Xi to demand a refund. In the meantime I'm gonna try again, but this time with GPU. I'm hoping for a 10x increase in speed if the USB bridge doesn't overheat and kill everything again.
>>
File: 1784765515745235.png (170 KB, 449x342)
170 KB PNG
>>109861358
>6h 55min 41s
>for 758 tokens
There have got to be more cost-effective ways of achieving this kind of output.
>>
>>109861358
How are people running glm 5.3 flash does llama.cpp support it now?
>>
File: ishowmask.png (29 KB, 119x117)
29 KB PNG
>>109861358
>758 tokens
>6h 55min 41s
>0.03 t/s
>>
>>109861358
Q2_K?
>>
>>109860910
will it fit on a 1080 and 16gb ram though?
>>
File: 1620452856967.png (41 KB, 265x283)
41 KB PNG
>>109861358
>0.03 t/s
>half a minute per token
>>
>>109861402
Valiant? Cast With The Everungiving Anomolous Between Lives Given. ?.
Test Logic Not Valid. ?.
>>
let the thread die. It's over, there is no merit to this place anymore.
>>
>>109861415
Arghhhhh.
>>
>>109861402
>>109861391
reminding that in XX century before internet t/s was even less
>>
>>109861371
Cost effective? You can get a 512GB NVMe for like $30-50 and put it in a $10 USB enclosure. Mine's 1TB, don't be jealous.
>>109861383
I'm using unsloth's llama.cpp, I downloaded the portable version from their github so I wouldn't have to compile it and it could just live on the drive with the model.
>>109861395
Q8_0
>>109861402
Maybe I should compile my own fork that display in tokens per second. Input was 0.09t/s btw
>>
>>109861415
>t's over, there is no merit to this place anymore.
Then Let Me Start My Meritocracy

While Some Are Being Cast Between Lives For Anomolous Dark Parasitism, being a theftive Everungiving, rather than Festive, and fowlness like Oppression Normalised. One Aspectual Reason of Many.

While I Have The Multiplicity Orders Tip Top Mark. Eh?
>>
>>109861458
>Meritocracy
MeriTopia

Because Word Usage Isnt Insanely Hardlocked.
>>
>>109861452
>tokens per second
seconds per token I meant, sorry I'm drunk
>>
>>109861452
> You can get a 512GB NVMe for like $30-50
In 2024
>>
>>109861452
>uncensored
Is this the orcarouter's gguf quant or your own?
>>
File: file.png (339 KB, 984x863)
339 KB PNG
>>109861470
picrel, I also have fistfuls of 128GB BC501A drives that I do disgusting things with.

>>109861484
yeah it's orcarouter's Q8
>>
>>109859140
I'm too old to feel like an onii-san anymore.
>>
the local scene feels kinda dead
>>
>>109861492
oi is that top one any good?
looks like i can rip out my wifi thing and put one of those in place in my optiplex
>>
baker...
>>
>>109861492
Is Picrel S3?
>>
>>109861562
>looks like i can rip out my wifi thing and put one of those in place in my optiplex
That's gonna be an A+E key slot - it cannot accept a normal NVMe drive. 99% of the time it's going to be a PCIe x1 connection and 1 USB port, that's what that slot actually is. You can actually adapt it to B/M to accept an NVMe drive, but the speed will be limited even with a cheap NVMe drive, and you're buying chinkanese adapters on Amazon or Aliexpress. Fine for some things, wouldn't ever recommend it for a drive though.
It is a pretty good drive. https://www.harddrivebenchmark.net/hdd.php?hdd=Micron+2550+512GB+SSD&id=39544
>>
>>109861588
;)
>>
>>109861590
cheers, i'll just rip the 500gb sata ssd out of my ps3 and use that instead
>>
>>109861560
which is based because all the poorfags cant join LOL
>>
By The Time You're Reading This, You Should Have Prepared To Vote Transmeta, And Transcended Everungivings.
>>
>>109861492
Don't you need to give out your details to download it?
>>
>>109861631
Those details are your username and email address.
>>
New bake:
>>109861586
>>109861586
>>109861586



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.