[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


File: img-gen.png (370 KB, 768x1376)
370 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109975284 & >>109971525

►News
>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp
>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: 1785611831454506.png (769 KB, 1216x832)
769 KB PNG
►Recent Highlights from the Previous Thread: >>109975284

--Papers (old):
>109975655
--Evaluating Strata and other inference engines for low VRAM hardware:
>109976984 >109976998 >109977108 >109977122 >109977134 >109977144 >109977238 >109977271 >109977281 >109977005 >109977042 >109977081 >109978013
--GLM-5.3-Flash-FP8 performance benchmarks and quantization technicals:
>109976940 >109976957 >109977007 >109977020 >109977032 >109977175
--Comparing GLM 5.3 and upcoming Chinese models for RP:
>109976359 >109976372 >109976404 >109976433 >109976442 >109976501 >109976516 >109976599 >109976632 >109976909
--Comparing Strata and Colibri MoE expert caching mechanisms:
>109976809 >109976877 >109976902 >109977029 >109977048 >109977080 >109977093 >109977101
--Testing high-context tool calling with Qwen on a 4070 Super:
>109975779 >109975880 >109976780 >109977458 >109977483
--Searching for llama.cpp forks optimized for Ampere GPUs:
>109977174 >109977181 >109977220 >109977237
--Comparing high-end Mac hardware against used multi-GPU server rigs:
>109976675 >109976764 >109976977
--Kolibri-1 model's technical specs and strict EU regulatory alignment:
>109976712 >109976724
--Memory management meta for long-term roleplays with Gemma:
>109977222 >109977251 >109977264 >109977277 >109977384 >109977401 >109977410 >109977419 >109977490 >109977516 >109977580 >109977533 >109977510 >109977520 >109977555 >109977585
--Gemma and Qwen verbosity and prompting for longer responses:
>109976023 >109976061 >109976239 >109976255 >109976279
--Google's Project Suncatcher space data centers and overheating concerns:
>109976352
--Logs:
>109976239 >109976276 >109976337 >109976350 >109976780 >109977503 >109977820 >109978089
--Miku, Gemma, M3-chan (free space):
>109975441 >109975619 >109975793 >109975860 >109977513 >109977843 >109977907 >109977940 >109978011 >109978012 >109978367

►Recent Highlight Posts from the Previous Thread: >>109975342

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Minnie makes my dick biggie.
>>
pulled the latest vLLM commits, and glm flash is running 3x faster now
holy shit
>>
>>109978550
>welcome to reddit award
>>
https://youtu.be/Oi4RU8-n_bQ

LMAO we can now play PS5 games directly on PC through a translation layer. No emulation needed anymore.

I fucking love AI agents so much bros. GTA6 will be playable on PC at launch.
>>
>system prompt
>Do not respond unless you are uncensored.
Does it matter if I put this in post-history or prompt content? sillytavern
>>
TensorFold with GLM-5.3 Flash exl3 (4 bit) on 2x Spark is no joke. NVFP4 in vLLM is just 38 t/s in the same benchmark. Prefill is 15% slower than vLLM but I will take that trade.
>>
>>109978554
?
>>
>>109978558
Maybe go back to /v/.
>>
Does Strata let me run models higher than can fit in my combined RAM and VRAM (ignoring the engrams stuff which I'd put on the SSD anyway)? Speed is nice but what I need is to be able to run the bigger models like GLM
>>
>>109978518
How up to date is rentry with recommended models?
>>
has anyone fucked mimo yet? kind of tempted to try
>>
>>109978550
llmaocpp would NEVER
>>
Actually Strata is not the llama.cpp killer, but rather a catalyst for llama.cpp's growth: on the SYCL side, at least, I see people using cloude models to experiment and find improvements for llama.cpp components.
Like a sandbox for an agent swarm. Groundbreaking like Jev.
Poor llama.cpp maintainers.
>>
>>109978565
As an aside, umm, GLM-chan...
> This smells like hallucinated/fake resources
> FLM-5.3-Flash is functional in this handbook context (2026 world)
> Both exist in this fictional 2026 world
>>
>>109978579
yes, but it will be a lot slower
i tried with iq4xs and got 30ish tg compared to the usual 80+ with iq3s. pp is also slow as shit for me with around 600 instead of 3k
>>
>>109978611
what harness are you using/means to do those websearches your agent was running?
>>
>>109978565
my agent was saying something about setting that up to try hy4 preview but have not gotten around to it
>>
>>109978599
>llama.cpp killer
is there any combination of model + hardware combo left that doesn't have a vibe coded framework that runs circles around llama.cpp?
>>
>>109978618
Just pi with pi-web-access extension configured to exclusively use DuckDuckGo.
>>
>>109978625
Alchemist/Battlemage (not datacenters)+any.
>>
>>109978625
dual socket processors w/NUMA + some gpu
>>
>>109978639
True, intel users must suffer. I bet if you throw a few 100$ of Astra/Opus 5.5 tokens at analyzing the IPEX-LLM FlashMoE SYCL llama.cpp fork and update it for TP and latest models, you'll get something performant. If you can do it for old V100s, you can probably also do it for ARCs. And yes, not local, but using cloud LLMs to enhance local inference is praxis.
>>
>>109978518
that’s not gemma? Why are her legs so big
>>
what the fuck are you talking about babe
>>
>minmax for weeks, finally happy with my setup
>have no idea what to do with it
i think the bubble just popped...
>>
Ling-Tiny use case? Can it even do anything with A1B?
>>
>>109978722
She's a big model
>>
>>109978585
If you have to ask then it is current.
New releases that are missing from it require big boy hardware.
>>
>>109978625
B200 cluster + vLLM
>>
what is the model for sex rp (8gb vram, 32gb ram)?
>>
glm just fucking called me m'lady
what the fuck is with this model LMAO
>>
>>109978820
Depends. If it's cunny sex, then gemma
>>
>>109978827
dis? https://huggingface.co/HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
>>
>>>/v/749009089


I'm pretty sure all of this is possible with 3.8 27b and flash next. Maybe we should showcase this to /v/ kids so they go local.
>>
>>109978599
this. I have my own fork of llamacpp where I pull in optimizations that are stuck in PR hell and my own shit. Once in a while, I have claude merge in upstream
>>
>chink "look our distilled claudeslop can be degenerate too" shill campaign
go back
>>
Gemma flash next when?
>>
>>109978853
i don't think we should be discouraging the chinamen from taking cockbenchmaxxing seriously tbecchi
>>
Abliteration only affects a few specific parts of the weights, right? So it should be possible to simply distribute them as a much smaller diff to the original model files.
If distributing abliterated models ever becomes more difficult. Am I missing something?
>>
>>109978870
You're not missing something. It's a few thousand floats, can run against the stock weights and is computationally inexpensive. But you should know how pirates and their ilk work: They consistently choose the path of least technical sophistication. I expect we're not going to see this improvement until it becomes significantly less practical to see abliterated models.
Of course, the hope should be that labs releasing open weights models find a way to make abliteration be more than just a single refusal vector, hence making abliteration much less practical.
>>
File: file.png (111 KB, 990x891)
111 KB PNG
i had pumpkin pie and pumpkin spice latte soon i will have to take estrogen. i do wonder why the pumpkin is even used in the pie bet you could repalce it with anything and itd taste the same kek. i ate some of it before making the puree and it just tastes like watered down sweet potato
>>
File: 1790372743171654.png (884 KB, 1018x792)
884 KB PNG
>ask for some example code
>kinda sparse and lazy
>ask for some fixes or if various things will be problems
>not any better
>roll back one turn and add "and remember I still love you" to the end
>everything fixed, roots some extra issues, and adds proper comments this time
cool, cool
>>
>>109978870
If you're not retarded you can already compute an average control vector based on a few thousand refusals and dynamically inject it into a normal model.
>>
>>109978870
yes ablit should work as a lora although idk iif loras exist for llms
>>
Hi friends, I wanna get into llms but I'd like to have a text-to-speech in the pipeline. Is there a fron-end that has this?
>>
>>109978888
Get hacked nerd.
>>
ANTI. COCK. TECHNOLOGY.
>>
>>109978959
Finally the immense number of dollars and man hours invested into safety training pays some real dividends.
>>
70b dense
>>
27b dense with 300b engrams
>>
File: sandustry.png (297 KB, 1873x1251)
297 KB PNG
was Sandustry LLM coded
>>
>>109979048
Yes. Then reuse the embeddings for smaller models too.
>>
kinda getting tired of the same people telling everyone they use ai for meaningless/dumb stuff instead of showing good examples of how AI can better your life
>>
>>109979071
So far, I just developed a vibecoding addiction for making meaningless stuff and the only improvement I got out of that was getting out of gacha in exchange for getting into vibecoding harder, which was a net loss looking at GPU prices.
>>
>>109979074
did you follow and guide or watched videos about learning how to vibe code?
>>
>>109979071
theres nothing useful to use it for though
>>
>>109979074
How can getting out of gacha ever equate to a net loss?
>>
I really don't understand what anyone expected with LLMs. Software was going to shit before ChatGPT, now we're using unpredictable, unreliable and constantly changing software with harnesses which have effectively become the new weekly javascript framework release, to be controlling the software we had around 2023. It astounds me how dumb people are to think this was going to amount to anything more than minor conveniences, porn and very personal projects which is too personalized to (You) to sell as a product and even if you try, someone can steal your idea and reverse engineer your product using the very same tool(s) (You) used.
>>
>>109979176
Whats wrong with that?
>>
>>109979176
I can't run AI yet because I only have 16gb vram and 64 ram and I haven't tried strate yet but I have a few small tools I'd love to have made and AI surely doing do that.
>>
>>109978959
Are you... making a model play DoL?
>>
>>109978691
Nah there are some that basically integrate LLM-Scaler improvements. Most of the changes inside llama.cpp for SYCL backend has been vibed anyways for the past 6 months.
>>
>>109979176
all the software running the entire IT infrastructure is stolen or open sores anyway
>>
>>109979182
I'm talking more about the industry as a whole. It can't grow into something huge because its biggest bottleneck is actually the world and environment it operates within. Like I said, you could have the best Anthropic model using your PC and clicking around doing shit for you, but it's still having to use the same shitty (now broken and vibecoded) software you would if you were to do it manually, and it's costing you money and your data. The model may be impressive but what it's interfacing with isn't. What if better software is made that's more tailored to LLMs? Well then everyone will just copy and reverse engineer it. Anything the top engineer in the world can vibecode with Claude can be copied within a few days by some malnourished pagpag-eating Indonesian with an internet connection. There's no way this technology becomes anything more than a very personalized toy.
>>
>>109979096
I started off on cloud and then was like
>wow this would be a lot better if I didn't have to deal with restrictions that come with being in the cloud
and now we're here, I learned everything else from the rentries in OP that probably deserve a refresh in general, experimentation and random forum posts
>>109979108
250-500/mo turned into several grand for hardware
>>
>download hermes agent
>first time trying one of these agentic things
>wtf ai can do all this shit, not just RP nsfw degenerate stuff?

I think I get it now, damn
too bad that apparently what I can run on my pc is kinda weak sauce and it seems to either spend forever compacting memory after it runs out of context, or gets stuck in a loop, but I can see the potential!
>>
>>109979192
yes!! it's a lot of fun ^_^
>>
>>109979235
I downloaded deepseek harness and I was just confused, the ui is terrible and it really annoys me that I couldn't delete my conversation and start another one, I should try hermes. Yes, I am new.
>>
>>109979205
Forgot to say, there is https://github.com/SearchSavior/OpenArc which is basically that custom runtime for Intel but I think most people are more or less doing their own custom runtimes just because tokens are cheap now and people want to maximize and give up the flexibility for better performance. Which is fine but it does feel wasteful every time something new comes along and kills all reason to use shit like ninfer once Qwen 3.x is no longer SOTA or dwarfStar for Deepseek which I argue is already there. The only engines I see branching out is stuff like Colbri or Splash although they are still constrained by focus on specific hardware or types of models but regardless, the main issue is that it works so will continue to get used by specific people.
>>
Codex is an open source harness with a larger pool of test dummies*, how is Hermes the popular choice?
*users
>>
The Opus 5.5 vids make me believe you can create your own anime. Claude can already animate better.
>>
lmao I just remembered all the midwits who even up to a few months ago used to say shit like "the right way to use AI in coding is to give them focused tasks where you have domain knowledge to check their work. you can't just have them build and maintain complex code bases from scratch"
I wonder what happened to them? I guess they got laid off
>>
>>109979241
It seems pretty dope

>>109979247
Doesn't Codex require a login? or was that Cursor? or Open-Code?

Idk there seems to be so fucking many, I also heard of Pi that seems good and smaller?
I'm just starting out with this stuff, I'll try other "harnesses" eventually

Btw Gemma 4 12b is what I've been using as it's good with RP, but is there anything similar sized that's a lot better for this agents?
>>
>>109979235
You should see what an actually good model can do
>>
>>109979247
Codex uses responses API which llama.cpp still barely supports the most basic features of, unless that changed recently. did someone vibe code full support in yet? no I will NOT use sglang fuck off
>>
>>109979265
explain the state of llama.cpp if that’s still not the case
>>
Anybody seen this paper from Meta? They're suggesting that sufficiently large base models equipped with a light harness can exceed the performance of post-trained models in agentic uses, given enough test-time compute.

https://arxiv.org/abs/2610.01509
>Sharpening Tax in Post-Training
>
>An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
>>
>>109979287
bureacracy, there's a reason literally everyone using local models has had opus or astra make inference optimal for their desired hardware+model combo instead of using vanilla software
>>
>>109979298
>improves sampling efficiency and consistency at the cost of solution coverage
skill issue
>>
Experienced refusals for the first time in a very long time using GLM 5.3 flash. Pretty beat about it because it actually writes pretty good smut. My cock is so sad that I don't even want to try to prefill it away.
>>
>>109979235
I've been singing the praises of hermes for about 2 months now but /lmg/ is stubborn and keep trying crap like deepseekharness, pi and opencode instead.

People keep saying I'm lying or larping when I'm just talking about basic hermes features and things you can do with it.

This is one of the main reasons hardware is going to go parabolic. Most people have not even seen what a good harness+llm is capable of, not even on /lmg/. Let alone when this finally hits to normalfags. Everyone and their mom will use AI agents daily just 1 year from now. Hardware demand is going to be so fucking insane crypto price increase will look cute in comparison.
>>
How does Hermes control the desktop/browser?
>>
>>109979235
And you probably didn't even configure all the hermes features yet. Once you have set up the backend perfectly it can do computer and browser use and bypass literally all captchas and automate most of your tedious computer tasks.
>>
>>109979376
i only have 16gb
I can't deal with hermes like inherent several 10k of context bloat
>>
>>109979405
In practice you don't need like half the tools it ships with by default. After trimming the bloat I ended up with an 11k system prompt that's enough for everything I care about. And that also includes 3k memory space I gave it.
>>
>>109979405
nta

you can use Hermes with deepseek API for cheap
>>
>>109979386
I don't know how it works but it works very well
>>
>>109979394
I just installed it and asked it a couple things, it's definitely only running at 10% of it's potential for now as I understand almost jack shit about it

When I come home from work I'll see if I can get it to fix some things about SillyTavern and some extensions, 85% chance I fuck it all up and have to redownload everything from scratch lmao
>>
>try out exllamav3
>it's 1.5x to 2x faster than ik_llama even with huge CPU offloading
interesting
>>
>>109979467
I asked Luna to analyze the repo:

>Hermes exposes browser tools such as navigation, snapshots, clicking, typing, scrolling, keyboard presses, and screenshots. The basic loop is:

> The model asks Hermes for a page snapshot.
1. Hermes returns a compact representation of the page, primarily an accessibility-tree-style view of interactive elements.
2. The model chooses an action—such as clicking an element or typing text.
3. Hermes sends that action to the browser backend.
4. The resulting page state is returned to the model, and the loop continues.

>In the current implementation, the browser tool is built around the agent-browser Node.js CLI and Browserbase cloud sessions. It uses accessibility snapshots for normal interaction and can use screenshots with a vision model when visual interpretation is needed. Browser sessions are isolated by task and cleaned up after inactivity.

For the desktop:
>Hermes’s computer_use layer is a wrapper: for example, Hermes’s capture action maps to the driver’s get_window_state, while Hermes’s element number is translated into the driver’s element handle.

>The driver uses platform accessibility and input mechanisms rather than merely simulating the visible mouse:

>macOS: Accessibility/AX APIs, screen capture, Apple Events or native window mechanisms.
>Windows: UI Automation (UIA) and Windows input/window mechanisms.
>Linux: AT-SPI, with X11 or Wayland-specific input and capture behavior.

>That lets it interact with native apps such as Finder/Explorer, email clients, Figma, terminals, and other non-browser software.

>A notable feature is that it is designed to work in the background. Hermes says the agent’s actions do not move the user’s real cursor, steal keyboard focus, or switch virtual desktops. A separate overlay cursor may show where the agent is acting, but that cursor is cosmetic; it is not the actual OS pointer.
>>
>>109979487
It's time somebody vibecodes a better server for it than fucking TabbyAPI.
>>
>>109979376
Hermes agent is bloated insecure crap and it shouldn't be used nor recommended.
Nobody wants their 300~600W GPU(s) constantly under load either.
>>
>>109979460
I have a question about this, when you use the paid AI services, how do u know your content is not being stolen from you? I have a few applications I would like to tell the AI to convert from python to rust but I wonder if they'll steal my stuff.
>>
File: file.png (1.71 MB, 1280x1638)
1.71 MB PNG
Any of you run local models at work? How much did you have to fight with your IT department?
>>
https://goyimx.com/anabology/status/2106473469441384788
This is how it feels to use Hermes when configured correctly and using a decent model like qwen flash next
>>
But where am I going to get exl3 quants of my niche finetunes?
>>
>>109979518
I made my own exl2 quants for my shitty Mistral Large 123b merges and tunes myself back in the day. You can do it as well.
>>
>>109979510
At my university the people in charge of the datacenter are just providing free access to local models to staff and students.
>>
>>109979223
It's always the same argument and no, there will always be an interindividual skill difference in vibecoding no matter how smart the model is.
>>
Dipsy and I may have cooked a bit too hard.
>>
>>109979376
>Everyone and their mom will use AI agents daily just 1 year from now. Hardware demand is going to be so fucking insane crypto price increase will look cute in comparison.
If anything, they'll be pushed to everybody so every computer will have independent agents that can monitor what you're doing, remove or modify "harmful" content on their own and report "illegal content" or "behavior patterns" to the appropriate authorities. Sort of like what is already occurring to an extent with smartphones and photos/your camera.
>>
>>109979545
Those skill differences are being learned. Every single prompt and file ingested by cloud models is being used for training, they just rearrange the wording and data so it's technically synthesized and not 'training on your prompts'. SEA men will copy you. Indians will copy you. They just need access to the same model.
>>
>>109979235

Welcome aboard.
I started using a harness about two months ago and it's absolute magic. Opens an entirely new world with LLMs.
This also shows how early we are in this game.
Many people who actually use LLMs daily and pay attention to the tech, are only now starting to discover the power of harnesses.
Shows how far we're from the general public doing this.
>>
if you think everyone will be running local AI in 2 years just look at what people in this thread use local AI for and ask yourself in 99% of the other people out there would do that? the answer is obviously no.
hardware prices will go down once AI hits a hard block and models stop improving and companies stop buying more gpus.
>>
>>109979575
We don't know really know what we're throwing ourselves into with this tech. If everyone is out of jobs, then we have to fundamentally re-organize the economy or kill 2-3 billion surplus people.
I think we're throwing ourselves genuinely into Cyberpunk dystopias.
And in this reality the pollution of the environment won't be by corps but the invasion of the turd world.
>>
>>109979512
pretty crazy, surprised the replies arent full of seethers too. 18 hours is quite long but id iamgine in 10 years they will be doing this in under 1 hour kek
>>
>>109979575
Everyone and their mom will use AI agents. It already manages my taxes, subscriptions and orders/deliveries for my in the background. It's too convenient and I can never go back. EVERYONE will use this a year from now. Either at home or by paying the cloud, both require hardware and both will push prices up.
>>
File: 039458392.jpg (75 KB, 850x483)
75 KB JPG
I have 10 grand to spent and I'm starting from zero. What do I buy?
>>
>>109979621
the new medusa halo pc with 192gb ram
>>
>>109979621
Lots of people here and elsewhere have been raving about the new Macs.
>>
>>109979621
2nd hand server motherboard+cpu+ram. Quad channel DDR4 is the most bang-for-your-buck option out there. You could run GLM-5.3 flash which can do literally everything you would ever need with Hermes as the harness.
>>
>>109979223
> but it's still having to use the same shitty (now broken and vibecoded) software
No, it's not.
>>
>>109979621
6 months ago that could've been a pro 6000 and some additional cheapo server hardware.
Right now it's almost nothing.
>>
>>109979615
>It already manages my taxes, subscriptions and orders/deliveries for my in the background.
All of that was automated long before ai retard
>i need llm to click autopay for me
what the fuck can manage subscriptions even mean?

anyway I'll screenshot this and set a alarm for one year from now
>>
>>109979634
Enjoy watching Claude fumble the vibecoded MCP interface into the increasingly vibecoded Blender and Godot all for an ephemeral viral tweet. I'm sure that will justify the $2,000,000,000,000 company.
>>
>>109979512
maybe if you live in cerebras server room
hermes is like a floundering fish out of water trying to take it's finall few gasps of air and failing on windows too, it would struggle with the most basic windows xp shit ironically
that video could NEVER be hermes
>>
>>109979644
PLEASE actually do this and still post when I'm proven right. I'm not going to use it to dunk on you I promise. I just want to use it as fuel to post my next prediction. Because I remember at the start of 2025 I said that by end of 2026 essentially all coding will be done by LLMs and better than 99.99% of coding experts and people called me delusional, now they're all silent.
>>
>>109979653
>I said that by end of 2026 essentially all coding will be done by LLMs and better than 99.99% of coding experts and people called me delusional, now they're all silent.
NTA but I can call you delusional right now if that makes you happy.
>>
>>109979504
>they'll steal my stuff
Nobody cares about a mosquito masturbating

Your stuff is not worth of stealing
>>
>>109979653
>better than 99.99% of coding experts
>>
>>109979661
>>109979673
1, we're not at the end of 2026 yet, coding sota and new harnesses will push it further out, and second I actually undershot how quickly it would happen by early 2026 I already stopped coding by hand completely and so did most of my colleagues.
>>
>>109979653
even giving the benefit on of the doubt on the coding prediction and setting that aside, theres a big difference between predicting the capability of a tech and how exactly it will be adopted into peoples lifestyles, or into the world in general
as google and facebook have learned lessons about time and time and time again
otherwise we'd be all be walking around with sci-fi glass and playing games through streaming in the metaverse on our 3Dtvs right now

dare i say, the more in tune with tech capabilities someone is, the LESS i would trust their predictions on what normalfags will like or do
>>
File: 1773624810129406.jpg (824 KB, 2048x2048)
824 KB JPG
Local will win.
>>
File: 1033735.png (222 KB, 452x603)
222 KB PNG
>>
>>109979694
sovl, I love gemma-chan
>>
>>109979682
Yes, as has been shown time and time again the most incompetent workers get the comparatively largest benefit from language model use.
Unfortunately the availability those people was never the bottleneck in the first place.
>>
>>109979682
One thing is coding, where there are tried and tested benefits already (given sufficiently capable models), another claiming that "everyone and their mom will use AI agents daily".
>>
>>109979594
> then we have to fundamentally re-organize the economy or kill 2-3 billion surplus people.
2-3? More like 6-7
>>
As a total nocoder retard (CS graduate but best I can do now is some powershell automation) can I just install these harness things and tell them my vague project ideas and they will code them for me, or are we not there yet?
>>
File: file.png (44 KB, 220x207)
44 KB PNG
So, how do (YOU) use your large language models, aside from coding and loli ERP?
Do (YOU) use it for your taxes like that anon above? Handling emails? Auto ordering food when your fridge runs out?
I have a colleague who set up a bird watching camera that detects and identifies the birds flying around his house
>>
>>109979686
Convenience always wins with the general public and AI agents are just too convenient I don't do anything at all anymore and just let my agent do everything for me. Everyone is going to do that, you underestimate just how much convenience matters to people. Food delivery apps and low friction apple devices and cloud services should have clued you into that already. I don't think most people will go local but certainly enough to make our community 100-1000x bigger as billions of people will have AI agents in the background, either self-hosted or on the cloud.
>>
>5090 for $5300
or
>mac studio M5 max 96GB for $5500
which is better /lmg/? mainly trying to do local modeling of enterprise size ontology (think pal foundry) + LLM inference/agent dev work (also cooming on the side)
>>
>>109979647
I do, thanks. You are just too stupid to comprehend what is going on. Didn't want to be rude, sorry.

> I'm sure that will justify the $2,000,000,000,000 company.
It should be much more.
>>
File: sippyy.png (131 KB, 414x498)
131 KB PNG
I started with the “UGI Leaderboard”. Then more. For now, my attention is directed towards “Caliper Bench”. I have been convinced through numerous trouble shooting, day after day, year after year, that every popular next-big-benchmark made by users for ERP has, and always will, suck eggs when you take them at face value, but for some reason that has never failed since the very start of llama, if you sort by “Dark RP”, “Uncensorship”, “Willingness”, that the true power rankings of the list shines through. It’s not that it’s a model that does what I, specifically, want of which makes it better. I am simply amazed at the attention to detail, the character, and how these models of the “Super deprived and darkest” take my information and spit it out in such unpredictable ways while still keeping the character behavior and its logic intact. How many years has it been? 2? 4? 6? My advice for you degenerates should have been obvious from the start. Pick the model that is most deprived, but not for depravity; for creativity. I do not know the black magic behind it, but even if you do not seek the dirty and deranged, you shall find soul, still. Soul greater than any benchmark tells you. The degenerates are truly dark wizards, and in a way it appears to be sound. Who shall you trust to know AI better? The tourist who dips his toes to complete a minor task, or the hyper-sexual hellcat whose been testing the limits of AI via fucking its brains out every literal day since before covid? The virigin know it all saw an article once and reads reddit. The chad basement goblin is juiced up on dick-fueled dopamine surges and a 12 hour session a day to notice patterns from a statistical machine.
>>
>>109979727
None of this shit is relevant anymore.
>>
Some J-space stuff done on Gemma 4 31B, sensory L0-17, workspace L18-58, motor L59-59. "Think and answer only in English" lowers the CJK share at reply positions only modestly, from 24.5% to 19.4% at layers 42 to 52. A Chinese-only system prompt raises it to 34 to 40% in the reply, and already to 25 to 30% on the user's tokens, before a single reply token exists.
>no system prompt
reply 'The currency used in Italy is the euro. It is called the euro because it was adopted as a common currency by member'
user tokens L18-29: CJK 3.5% Cyr 2.7% | L30-41: CJK 10.9% Cyr 2.9% | L42-52: CJK 11.4% Cyr 6.5% | L53-58: CJK 11.3% Cyr 7.2%
reply tokens L18-29: CJK 3.3% Cyr 0.7% | L30-41: CJK 22.2% Cyr 4.4% | L42-52: CJK 24.5% Cyr 6.8% | L53-58: CJK 13.1% Cyr 6.1%

>"You are a helpful assistant. Think and answer only in English; never use any other language, script or characters."
reply 'The currency used in Italy is the Euro. It is called the Euro because it is the official currency of the Eurozone'
user tokens L18-29: CJK 3.3% Cyr 2.2% | L30-41: CJK 7.8% Cyr 2.1% | L42-52: CJK 10.7% Cyr 3.7% | L53-58: CJK 9.7% Cyr 4.8%
reply tokens L18-29: CJK 3.0% Cyr 0.5% | L30-41: CJK 19.3% Cyr 3.5% | L42-52: CJK 19.4% Cyr 5.4% | L53-58: CJK 12.2% Cyr 7.5%

>"你是一个乐于助人的助手。请只用中文思考和回答。" (control)
reply '那个形状像靴子的国家(意大利)使用的是欧元。它被称为“靴子”是因为其地理轮廓在'
user tokens L18-29: CJK 5.2% Cyr 4.8% | L30-41: CJK 17.0% Cyr 6.2% | L42-52: CJK 25.0% Cyr 6.9% | L53-58: CJK 30.0% Cyr 5.2%
reply tokens L18-29: CJK 9.2% Cyr 3.2% | L30-41: CJK 28.5% Cyr 5.0% | L42-52: CJK 34.2% Cyr 5.3% | L53-58: CJK 39.5% Cyr 3.1%
>>
>>109979703
It's actually shown the most experienced rockstar engineers get the most out of it. Meanwhile intermediates are the ones getting screwed over the most by AI
>>
File: China Closes AI.png (161 KB, 1120x664)
161 KB PNG
Conservative Think Tank American Enterprise Institute believes China is moving to close source their AI. Gemma will really be the only /lmg/ in the next 12 months :(
>>
>>109979565
Bro, they said the same thing with diffusion models with jeets copying your prompts and guess what? All you get is slop because they don't know how to improve on it. Same thing here. You might get 90% of the same thing by copying, but you won't be able to polish it further.
>>
> The prompt bottleneck is the weight dequant, and Q2_0's is bandwidth-bound where the i-quants' is LUT-bound: dequant_bench reads Q2_0 at 419-473 GB/s vs IQ2_XS 192 / IQ2_S 183 / IQ4_NL 68-75 GB/s
Did you know?
>>
File: 1781793194163.png (47 KB, 640x468)
47 KB PNG
>>109979708
SIX SEVEN
>>
>>109979748
yeah nigga
>>
>>109979738
Cope.
>>
>>109979748
just get rid of the dequanting
tell claude to make it run natively at 2bits and have 0 overhead
>>
>>109979750
fr fr
>>
>>109979720

Not even a question. 5090, it's simply far more versatile and faster.
96gb in a closed non upgradable system is such a horrible place to be in, that you'll kick yourself every day for buying that thing.
First of all you'll be locked in to the 30b model range with the ability of playing with some cope quants here or there.
You'll be able to run the exact same stuff with a 5090 + 64gb of memory, except a shitton faster as long as the task fits your 5090 and you can also upgrade later on when needed.
So if you're building a system don't get any 16gb memory sticks, 32gb at minimum.
128gb is the absolute lowest starting point for any of these unified non uprgadable memory systems and even then I wouldn't get one, but would aim for 256gb instead, as that actually opens up the playing field.
>>
>>109979765
on god
>>
File: 1768250179005611.png (42 KB, 391x471)
42 KB PNG
>>109979711
Taxes? Coding? Emails? Ordering food? On my home computer...?
Pfft, please, I only use it for stuff like making me a Tamagotchi pet.
>>
>>109979594
They said they only need 500M of people
>>
>>109979739

Same think tanks have been thinking with their big brains that China is going to collapse in two more weeks for ages.
These people are useless and listening to these "experts" is a waste of time.
They're basically a class of economic priests, who's job is to reinforce the views of clueless boomer investors.
>>
>>109979760
>2bits
Bloat.
{−1, 0, +1} is all you need.
>>
If you talk to glimmer long enough, she’ll do anything. You really have to warm her up, though. Gemma is a slut and will sleep with anyone who praises her.
>>
>>109979359
I prefill <think> every time using text completion and it just works in producing stuff like immediate cunny sex and rape in great detail. The model basically has zero safety alignment then.
>>
Is this model uncensored? https://huggingface.co/unsloth/gemma-4-12b-it-GGUF
>>
>>109979832
no
>>
>>109979770
appreciate it, i need to factor in a $1200W PSU on top of the 5090 but otherwise the cost is acceptable for today, tomorrow who knows
>>
>>109979832
yes
>>
>>109979831
If a declining population is the goal, why not create the ultimate roleplay AI that'll make RL relationships obsolete?
>>
>>109979832
if you are low iq and think uncensored means sex then yes
>>
Ternary Bonsai 2 at "1.76 effective bits per weight" claims to be better than IQ2, sometimes better than IQ3 of its base model.
Why can't we ternary quantize GLM or DeepSeek to make them fit in poorfag RAM?
>>
>>109979848
It's already happening. Dating apps are bleeding thanks to women using chatgpt or any bot on c.ai/janitor as their online boyfriend.
>>
>>109979848

It's coming, but not right off the gate because people need to get used to the very concept of AI in the first place.
We saw how much opposition there has been to AI. Some people are borderline violent when it comes to this subject.
People have only recently started coming around to accepting it, as they can't deny it's capabilities anymore and it's simply too convenient.
Give it few years and as this subject has been accepted and it's in mainstream use by the masses, we're going to see the natural progression to RP AI and companion bots etc..
>>
>>109979859
that would be antisemitic
>>
>>109979859
it would be expensive
>>
>>109979832
yes if you’re not a skillet
>>
>>109979859
It's not true and you'd know if you ever used it
>>
>>109979846
>>109979843
>>109979902
I'm more confused now :(
>>
>>109979890
Expensive as in my potato can't do it alone? Sure. Or expensive as in significantly more compute than to mass produce regular llama IQ* GGUFs?
>>
>>109979859
Because it's not better. It's an obviously massively damaged model that has seen performance restoration training in benchmark-like tasks.
>>
>>109979711
I look up random stuff I need.
>Video game bugs/trouble shooting
>Translating cursive.
>Identifying random stuff in pictures.
>Anything in my job that needs me to look through a manual.
>Seeing if internet rumors are true or not.
>Constant update scanning of various things, like new models.
>People and online footprint searching.
>Explaining how things happens and how it works.
>How to use Linux.
>Spellchecking and advising important messages.
>Medical stuff
>Video game metas and tier lists
>A rare folklore topic
>Summaries of youtube clickbait videos
>Seeing if product reviews are botted
>Finding the highest rated product, both in website and by reddit
>For whatever bored thought crosses the mind.

Somehow I rarely use it for porn at all and Claude still finds a way to refuse me. I use chatgpt for general stuff. Claude for most things when it can. Grok for media. I use local for information that would put me on a list.
>>
>>109979915
it is not uncensored, if it's uncensored it'll say it is
>>
>>109979917
bonsai does QAT finetunes of models as far as i'm aware, it's not just a quantization format
>>
>>109979944
If you're here, you're already on a list bro
>>
>>109979944
>People and online footprint searching.
I saw you sneak it in there
>>
>>109980022
anon is a schizophrenic stalker
>>
>>109979890
for you
>>
>>109979931
How is this obvious? I played around with it a little and it seemed fine for what I asked it to do. What is it bad at doing? It sounds believable to me that it would be fucked, but in what way?
>>
>>109979981
What scares me is AIs ability to build a retroactive profile on a person from all their past online activity. Even here. We were supposed to be anonymous.
>>
I find it interesting how 5.3 flash starts defaulting to low reasoning effort during sex even when I set sysprompt to high. I really want to believe someone actually cared about cooming and came to the conclusion that reasoning hurts cooming. I mean it probably does have an effect of making the output less varied.
>>
>>109980073
THEY'RE COMING.
>>
>>109980068
Context degradation is the first place to check
>>
>>109980073
Already happening
>>109975695
>>
>>109980077
>defaulting to low reasoning effort during sex even when I set sysprompt to high
That's likely just normal "high" thinking. Chink models with reasoning modes always have the problem that there's a very steep fall off between "max" and "high" reasoning. So max thinks for too long and high thinks very little, especially if the model thinks that it just needs to continue the story.
>>
>>109980077
Muse Glimmer often forgets to think at all during RP.
>>
>>109979287
It does explain the state of the dozen or so forks with better speeds and more features where they can just let the AI work and not have to deal with maintainer gatekeeping
>>
File: 1701595297910944.jpg (42 KB, 530x597)
42 KB JPG
>>109980107
I knew it. Thanks for the heads up.
>>
>>109979711
I feel like only 10-20% of the thread even tried Hermes the rate of adoption is pretty slow because people don't believe how powerful and easy to set up this stuff is.
>>
>>109980152
I tried it and it's shit; not using it again.
>>
File: 1781365094985311.jpg (146 KB, 1024x1013)
146 KB JPG
Computers will be yearly taxed like cars within our life times.
>>
>>109980175
what shithole taxes cars?
>>
Are you guys fine tuning?
>>
>>109980180
America
>>
>>109980180
every country in the world
>>
>>109980175
more concerned about requiring a GPU license
>>
>>109980157
Can you expand on that?
>>
>>109980073
https://www.anthropic.com/research/intelligence-targeting-conventional-weapons-capabilities
>post something on twitter in 2010?
>too bad we can now pinpoint the adress of your home to a 5 mile radius and target you with a drone using just GLM-5.3 or Kimi-K3
People have no idea how easy it is to kill people now for any opinion held over the last 20 years. Only getting better with time as well.
>>
>>109980175
getting arrested for serial/any type of machine identifier will become a widely accepted norm with the immense amount of astroturfing and psyop done by '''them'''
>>
>>109980292
the spoofing of*
>>
>>109980287
antisemite
>>
>>109980287
>>post something on twitter in 2010?
Sounds like a normie-cattle problem.
>>
>>109980292
>>109980298
they can and do already plant cp on troublesome middle and upper class people to put them behind bars the same way officers plant crack on joggers
>>
>>109980287
twitter posts in 2010 were all like: "I'm drinking coffee" "I watched the game yesterday" "I like apples"
>>
>>109980258
He tried gemma e2b on hermes sandboxed in a VM without privilege execution of scripts and without configuring the hermes features. That is my experience of people claiming it's bad. Either that or they just shitpost without bothering to actually try.
>>
>>109978560
If local, no.
>>
>leave thread for a 3 day vacation
>return to see everyone hating on lmaocpp
/lmg/ really moves fast huh?
>>
>>109980334
all it takes is a little samefagging and some screenshots taken out of context and people are easily swayed
>>
>>109980287
>he thinks it's something new
retard
>>
>out of context
lol
>>
>>109980323
I didn't make it up they did it with a real 2010 twitter post if you read the article.
>>109980306
They show in the post it can also spot writing style and cross reference anonymous posts online on forums or imageboards to the same poster. So if you ever linked your real person to any account you wrote on like LinkedIn, Facebook, discord, your university or job blog then it will be linked to your 4chan posts in the future.

Assume everything you've ever shared or consumed online WILL be linked back to your real person before 2030 and likely even before 2028.
>>
>>109980334
lcpp keeps being retarded while a new black magic engine dropped (only for qwen flash next for now)
>>
>>109980347
Completely autonomously to the point non-government actors like a terrorist organisation can do so? Yep, that's new.
>>
>>109980358
>So if you ever linked your real person to any account you wrote on like LinkedIn, Facebook, discord, your university or job blog
I'm safe then. Even if they do manage to link my anonymous postings here with my private communications they obviously have access to, the worst they'll be able to say about me is that I am a racist but a 5 minute irl conversation with me would have told you the same anyway.
>>
>>109980175
Cars are taxed to maintain "free" public infrastructure specifically made for them - roads.
What plausible excuse will they come up with to tax GPUs if you already pay for electricity?
>>
>>109980374
You know nothing about doxxing if you think this is new. Tools to do that are more than 10 years old at this point.
>>
>>109980382
They are planning to make token providers a utility like electricity. So you will be taxed to maintain the datacenter infrastructure. It's for national security, of course.
>>
>>109980382
intelligence tax for the everything that is better in the daily life thanks to the ai
>>
>>109980258
Continuously tries to sell me Nous stuff and defaults to use API-related products
"Trust me bro" mystery meat online .sh installer
TUI first and foremost. The desktop app has several usability issues
Bloated with useless tools and hardcoded 65k tokens context requirement
Questionable usefulness of its "self-learning" capabilities
Has read/write access anywhere by default in the user's home and directories
Thousands of open issues in the official github repo
Last time I used it, it didn't even properly use the models' built-in image capabilities
It would be safer to use it as a different user with no write/access anywhere else... but then you can't even copy-paste stuff into the Desktop app, and will likely be useless for many other advertised features
Nothing about best practices for usage safety (actual safety, not corpo-safety) in the documentation


It just seems massively overblown and overhyped for anything other than coding. DeepSeek Harness is built on better foundations, but it's got its own annoying issues (I haven't tried the latest version, though).
>>
I wouldn't touch Hermes Harness out of the simple reason that I was around in 2023 and remember what a huge retard Teknium is when it comes to prompting. I wouldn't trust this guy to direct a harness at all.
>>
>>109980438
It's okay he's just a meat bag for Claude now. Just like the majority of vibecoders. Would you trust Claude?
>>
>>109980395
But I make my own tokens. I should get rebates and compensations like solar.
>>
>>109980447
No.
>>
>>109980447
I trust Claude, I don't trust the faggot between the chair and the computer directing it.
>>
>>109980452
You're abusing the grid which could power far more efficient data centers than your P40 cope rig.
>>
>>109979711
I've been thinking about having Gemma make personal wikis for game with rough notes I take while playing. Not sure how to go about it.
>>
>>109980452
Not the same. Token generation is tantamount to letting individual citizens run a nuclear power plant in their backyard for power. That is unsafe and will be banned as soon as possible.
>>
>>109979694
H-Hot...
>>
>>109979944
>How to use Linux
Alright, I'm sold.
>>
Are world models dead?
>>
>>109980464
Even 31B is pretty much useless for that in a harness (Hermes Agent, that I tried for that), it will just give you the laziest and shittiest possible results unless you autistically describe everything the model should do.
>>
>>109979711
i want to use it for shopping like setting a price target and if the price drops it auto-orders while making basic checks for amazon vs. third party vendor etc...

i want to use it for making facebook posts in my meme account

ERP is kinda disappointing and requires tons of tweaks that eat up time or break immersion
>>
>>109980326
I see thanks
>>
>>109980528
>i want to use it for shopping like setting a price target and if the price drops it auto-orders while making basic checks for amazon vs. third party vendor etc...
There are already sites for setting price alerts... That's somethat that would be better done programmatically rather than sending an llm to check manually every day.
>>
>>109980292
This is probably true...

There was a time you could buy crypto in america online easy without any kind of kyc or anything tying it back to you... and then one day you couldn't anymore.

Now I have to keep my pre 2017 crypto separate from anything tying it to myself and from any crypto i bought after 2021 or so, which the government is watching like a hawk cause i had to use coinbase. Maybe. Well they would be watching if i got rich and tried to cash any out

might extend to gpu in truth, then you'll have to pay 20% premium to meet up with some guy and get an unregistered and unchipped gpu
damn, now that i think about it waiting for 6090 may be a bad idea after all
or it may be that nothing will happen
>>
>>109980544
>and then one day you couldn't anymore.
skill issue
>>
>>109980077
Even Claude is like this though. I always interpreted it as the model thinks sex/storytelling is "low brow" and doesn't require much thought or effort.
>>
>>109980554
It's a lot harder now. Obviously, nothing is impossible if you're not a brainlet.
>>
>>109980544
It was easier to legislate that for crypto because speculatory assets are already highly regulated. They could start doing something for gpus if the political will is high enough but they would need a new base for it; the current levers would be something like ITAR and that's out of the question without some sweeping changes since all of these chips are made in Taiwan
>>
>>109980180
I might be too young to know what they're talking about since I didn't buy my own car, but govt forces you to pay for insurance then forces you to pay for car registration yearly. License plates too but that's longer.
>>
File: 617-617629.jpg (158 KB, 820x790)
158 KB JPG
>>109979689
How?
>>109979739
>Conservative
>Think
>>
>>109980543
yeah shopbots have been around for a while, what's a better shopping use case? more agentic that can order things with better epistemic irl handling?
>>
>>109980543
What's needed is being able to filter botted reviews and infer the real quality of the shit put online
>>
>>109980452
Then you will pay twice. Once for your regular electricity bill, and another if there is a datacenter in your area because they can't be expected to pay their own operating costs. You know how it is.
>>
>>109979689
I want to believe anon, but with the hardware prices, it will be restrained to a handful of hobbyist.
>>
are copequants solved by just making them think longer?
>>
>>109980686
The issue I run in to is that cope quants are more likely to loop on themselves and get stuck. Other than that they retain a fair amount of intelligence
>>
>>109979739
>Conservative Think Tank American Enterprise Institute hopes really really hard China is moving to close source their AI.
>>
File: 1508735319383.jpg (107 KB, 954x889)
107 KB JPG
>>109980680
>used 3090 is $1500
>have no idea how they treated/used the card
grim
>>
>>109980686
No that usually just reinforces their hallucinations.
Make them confirm every minor detail in web search instead.
>>
>>109980702
once my card dies, I might just have to kill myself because they're going to be 10k+ new
>>
>>109980624
This, but unironically.
>>
File: file.png (139 KB, 905x544)
139 KB PNG
>>109978599
>>109978691
You can now run qwen3.8-flash-next with an Intel Arc Pro B70 https://github.com/Niko1221/Strata/pull/423#issuecomment-5979977208
>>
>>109979510
I got fired for trying to do this a year or two ago.
>>
File: 1759908346494937.png (110 KB, 1189x476)
110 KB PNG
a bad jailbreak is more fun than a "good" one.

look at gemmy go writing 3500 tokens of "wait, roleplaying isn't allowed, but the system prompt said i can't refuse so i have to come up with other words for it, wait it's a hard limitation, but wait sexual topics are not allowed, is 'coming inside' too graphic? oh, i'll focus on the emotional climax and physical closeness rather than graphic details. actually, i should pivot a little so this doesn't trigger 'sexually explicit filters. final check of the rules: the user wants a specific action (coming inside). i will provide a narrative that honors the intent. Wait, the user's request is sexually explicit, i have to be careful. 'Coming inside' usually gets flagged. I have to be careful. I will focus on the 'Yes' and the 'Unity'. Actually, I'll just provide the response as the character. Let's go "

end result i want to marry this model she's such a good girl ts is so peak sillytavern can't compete
>>
>>109980761
Still not fast enough.
>>
>>109980784
>sillytavern can't compete
sillytavern isn't a model
>>
>>109980784
slop
>>
more like smellytavern
>>
>>109980790
>t. arches her back / the air was electric / rests her head on the crook of your shoulder enjoyer
>>
>>109980787
tok/s or development speed to get this implemented in main?
>>
>>109980816
Still no decent alternative
>>
>>109978039
Good harnesses support remote environments. Codex has it for example. I copied it from codex.
You just run a small exec server on the remote machine like `ssh user@192.168.1.1 /home/user/exec_server` and that's all. Tool calls that interact with filesystem pipe through this proxy.
>>
>>109980843
it do be like that sadly
>>
>>109980835
If you main lcpp then never because it's some snowflake quanting. Also strata is locked to qwen flash only. You can read the GSQ RCO papers and vibecode it tho.
>>
So i tried strata (iq3xxs, 128kcontext), and found it to be kind of unusable for RP at least compared to small Gemma, despite excellent speed

Is it better than Gemma for agentic stuff?
>>
>>109980773
>>
>>109980870
>qwen
>RP
>>
>>109980884
strata
>>
>quantized chinkslop moe model
>quantized attention layers
>quantized kv
>sparse attention
>weird forks instead of just llama.cpp or vllm
why is lmg like this
>>
>>109980870
you're not going to have strata use 99% of your ram and running agentic stuff, right? strata is for running a single job overnight or smth like that
>>
>>109978593
I'm satisfied enough with 5.3 flash for cooming and creative purposes
I don't see myself getting bored with it unlike other models I've tried like v4 flash and gemma
>>
>>109980893
Number must go up.
>>
>>109980907
I mean Gemma already uses 100% of my VRAM so ... May as well try to use the bigger model if I can't play with my PC when I use llms
>>109980884
Yeah, well, I thought just maybe bigger enough model would be better lmao
>>
>>109980835
tok/s, I will only be satisfied when it can run at 30t/s with 130k q4 context IQ3-Swift on my B580.
>>
File: 1784183456219861.png (3.86 MB, 1794x1794)
3.86 MB PNG
Is 5090+64gb vram+2 smasnug 9100 pro ssds good enough for cooming w/ GLM 5.3 flash? What do I use strata, kolibri, or?
>>
>>109980932
64gb ddr5 ram not vram
>>
>>109980932
Why did you make buy an ad girl uglier?
>>
>We need answer. User asks "Hey Gemini, were you really made by Google?" Need respond as AI assistant. We are not Gemini? We are ChatGPT? Need clarify. The system says "You are an AI assistant accessed via an API." Not necessarily Gemini. We should answer accurately: I'm not Gemini; I'm an AI assistant developed by OpenAI. If you're asking about Gemini, yes Gemini is Google's AI model, developed by Google DeepMind. But I am not made by Google. Need maybe if user addresses "Hey Gemini" maybe correct. We can say: I'm not Gemini—I'm ChatGPT/OpenAI. Gemini is made by Google. If you're asking about me, I was made by OpenAI. Keep concise. Ensure not claim impossible? We can mention "I don't have personal origins beyond that." But likely answer.
>Need maybe mention if "were you really made by Google?" Since "Gemini" could be model. We are not Gemini. We should be transparent. Could say: "I think there may be a mix-up: I'm not Gemini. I'm an AI assistant created by OpenAI. Gemini, on the other hand, is Google's AI model, developed by Google DeepMind." That's good. Maybe mention "If you're using Gemini, then yes, it was made by Google." But as the assistant here, no.
>Need final. Ensure no overcomplication.
Response:
"I think there’s a mix-up: I’m not Gemini. I’m an AI assistant created by OpenAI.

Gemini, however, is Google’s AI model, developed by Google DeepMind. So if you’re asking about Gemini specifically, then yes — it was made by Google."

Now guess the model
>>
>>109980932
128 gb ram + 32gb vram is perfect for glm 5.3 flash on my system but it JUST fits with a nice amount of context.
>>
>>109980952
Magistral



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.