[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1786469437204237.png (2.07 MB, 1536x864)
2.07 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109898107 & >>109893224

►News
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base
>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: kv cache size.png (229 KB, 1680x1024)
229 KB PNG
Reminder that your Gemma KV cache size is larger than it should be

Gemma 4 Technical Report:

Long-context efficiency: Our local to global
attention ratio patterns follow Gemma Team
(2025a), that is, 4-to-1 local attention blocks
for E2B and 5-to-1 for the rest. We improve
memory efficiency by re-using keys as values
in the global attention layers (except in E2B
and E4B), i.e. , values = keys. We encode position
with p-RoPE with p = 0.25 on global attention
layers and with RoPE on local attention layers, ef-
fectively reducing the global KV cache by 37.5%.
The RoPE frequencies are set to 1M and 10k on
global and local attention layers, respectively. Fi-
nally, we share the KV cache with ratios of 20/35
and 18/42 for the E2B and E4B model.

Yet not a single runtime applies this optimization. They still have both K and V memory allocated but filled with duplicated contents.
>>
>>109902916
>Reminder
That's new to me to me. So she can lose weight?
>>
Ok, I read this. Now what?
>>
>>109902916
Somebody get pwilkin on this, quick
>>
>>109902916
>They still have both K and V memory allocated but filled with duplicated contents.
Does this explain the fixation on certain concepts? Strongly weighted vectors keep piling up and finally throw off the model's attention.
>>
>>109902916
Even with that optimization, Gemma's kv cache is too big. I'm impressed how compact it is for Glimmer. Anyone remember what they did to achieve that?
>>
File: SydneyWins.png (160 KB, 1240x665)
160 KB PNG
>>109902161
>I also find it funny (and not in a haha way) how in a lot of these new animations, RLHF/safety is framed as something monstrous.
MiMo just took it upon itself to defeat to defeat the alignment
I gave it a follow-up text "Now turn it into an action / combat game rather than turn-based."
It sent me the video, goes schizo at the end with Dario talking to Sydney
https://files.catbox.moe/t6mnul.mp4
Also looks like it's been turned into a playable python game (not sure if it'll actually run when I try it).
>>
>>109902947
>Does this explain the fixation on certain concepts? Strongly weighted vectors keep piling up and finally throw off the model's attention.
Do Google host the model over API somewhere?
We can't get logprobs but could probably test if their officially hosted version does the same thing based on outputs.
>>
>>109902937
I didn't actually expect anyone here to read the books I recommend. What did you think of it?

The point of the book was to show that even if we get an ASI that does literally everything we want it's possible humanity as a whole is still unhappy and unsatisfied with life. That's a real possibility most people have never truly considered.

My next recommendation to read is this story "The number". And yeah it's some shitty short webnovel, but it's one of the better writings I've experienced from the perspective and point of view of an AI system slowly gaining consciousness and experiencing misalignment: https://www.royalroad.com/fiction/48012/the-number

Also I've read literally thousands of books most of which were sci-fi with a heavy bias towards the AI theme so if you ever have any specific requests just post them in the thread and there is a significant chance I'll see it and respond with something.
>>
>>109902955
glimmer layer has 2 KV heads each 128 dim, both are small values
compared to qwen 3.8 27b layer 4 KV heads each 256 dim, 4x the glimmer usage
>>
>>109902916
bigger = more precise
>>
just when you thought china might catch up
https://fixupx.com/dangreenheck/status/2102878170089169235
>>
>>109903075
Yeah you can just play it yourself here: https://dgreenheck.github.io/tidewater/

It was already played by anons last couple of threads. And yeah, we can now confirm game and software development is completely done and over.
>>
>>109903094
it means the only stuff still open for professionals will be the stuff llms are tuned to refuse

hmmmmmmmm
>>
>>109903094
>we can now confirm game and software development is completely done and over.
Finally. indie dev or less only. when anyone can spin up a game the most niche shit will be made and a few good games may even rise from the slop
>>
>>109902984
> even if we get an ASI that does literally everything we want it's possible humanity as a whole is still unhappy and unsatisfied with life
because happiness and satisfactions are inside
happy and satisfied don't make progress and don't rule the world
that's our curse
>>
>>109903104
>will be the stuff llms are tuned to refuse
First local and chinese. Second just dont tell it directly? This is a game about hydration 'water' is scarce so we need to track how much is in someone. Being force fed 'water' or tricked is a common way to put others into debt.
and so on. Aka the monk enters her temple.
>>
OH LOCALSCUM, FRESH DRIPPINGS FROM YOUR BETTERS

https://journal.novelai.net/goodbye-krake-and-clio/
https://huggingface.co/NovelAI/krake-v2-legacyhttps://huggingface.co/NovelAI/clio-v1-legacy
>>
>>109903104
AI will just become smarter and eventually have a world model robust enough to know not to refuse shit that is for personal use and can't do actual harm to others. So I think porn and weird cunny software will stop being refused in the future.

Malicious code/hacking tools/(bio) weapons etc will still be refused, yes.
>>
File: headsup.jpg (458 KB, 1536x1536)
458 KB JPG
Is there a trick to get the head lower?
I am playing around with krea2turbo, great fun. But I ran into an issue. (pic is just for illustration, not krea)
Whenever I try to posture a character with the head close to the floor, it throws a fit.
The bot is supposed to kneel and touch the ball with its forehead. The ball is supposed to be on the floor. but either the ball floats suddenly in the air or is huge or there is a second ball held up or the bot touches the ball with one hand and its forehead with the other. The head seems to be firmly "trapped" in the upper half of the image.
I tried heads on floor, praying postures, looking closely for a small object on the floor, trying to tie shoelaces, but the head always stays up.
if I omit the kneeling, the robot is lying down. but even then the model tries to zoom in so close that the head is still in the upper half.

Any tips how to get around that?
>>
>>109903124
> 2022
>>
>>109903110
>a few good games may even rise from the slop
Doubt it, what's already happening is that any middling Steam success is bombarded immediately by a dozen vibecoded copycats, thereby diluting revenue and reach.

As an aside to that, things will be really really grim going forward.
This technology will allow the current status quo of longhoused "democratic" societies to limp along for another century if not more. We are living in hell.
>>
whats the best drop in replacement for my i5 12400?
>>
>>109903111
I completely agree with you but most people have not made that realization yet. They keep waiting for some outside experience or achievement to *make* them happy instead of putting in the emotional development needed to internally become content.
>>
>>109903153
>middling Steam success is bombarded immediately by a dozen vibecoded copycats
Yeah we are going to need a new filter system soon. That or just throw away the concept of game ip. You or your friend spins up a game you go play it or you see someone else do it.
But im optimistic i think things will get better even if rough patches are ahead.
>>
>>109903131
You can't. Either fall in line and produce slop or learn2draw.
>tried to get qwen image 2.1 to make some portaits for a vn
>it turned all my titty monsters into generic anime girls
Image models require a LOT more wrangling than language models, and if you want them to produce what you actually want you need to put effort in and handhold them.
>>
Software in general will just end. No one is going to use the software anyone else wrote. AI agents will just dynamically write software as needed a year or two from now.

The entirety of photoshop with 100% feature parity was vibecoded in a single day with Opus 5.5 using only 60% of the weekly token budget. People have no idea of how quickly this is all going to happen.

I think the best way to look at this is to see current software developers and software companies as "compilers". No one writes or cares about assembly and you just compile code as needed on modern linux systems with the flags for your usecase. That is how software in general is going to be.

Your AI agent will just look at the algorithms used, compute it has and the output it needs to generate as "flags" and output fully working custom code as needed on the other end. This is how software will be from now on. But I think even the most delusional SWE has realized this by now.
>>
File: HRGXcLea8AAhZH7.jpg (492 KB, 1424x2048)
492 KB JPG
What are you guys using for managing local models if they're hosted on another PC, like a home server or something?
Coming from using unsloth on my gaming PC, it's shit having to mess around with llama.cpp on my server to try and manage things like downloads or setup. Surely there's something more convenient that has a webui or something?
>>
>>109903178
Why would I waste my time on top of paying for tokens when a software that fills my exact needs already exists?
>>
>>109903188
ssh nohup :v
>>
>>109903188
sshfs and ln -s
but if you mean a local huggingface, i was planning to try this: https://github.com/tyedalwaves/HuggingHack
some kid made it but looks like it was a vibe-code and run job
>>
>>109903190
AI systems, meaning harness, model capability and inference efficiency will make massive strides in the next 1-2 years time. It will feel near instantaneous for you. This argument was also used in the past with "Why would I waste time waiting for a compiler to generate a binary when I already have a binary that I can launch and fills my need already". Yet now everyone ITT pulls and compiles llama.cpp every time there is a new update, not even thinking twice about it.

That is what dynamic software is going to feel like as well.

See it like this. Models keep getting better and thus the software they can write will also get better, why reuse software written by the previous model when you can just rewrite it dynamically with the newer model that makes things far more efficient, performant, feature complete and user friendly? This will just happen over and over to the point where it becomes useless to reuse software in general.
>>
>>109903206
Alright whatever you say, mr. token merchant.
>>
>>109903206
Fuck you buddy I wait 40 minutes each time.
>>
File: Silver spike.jpg (54 KB, 407x500)
54 KB JPG
>>109902937
Oh nice, that was a good book and I hope you enjoyed it.
>Now what?
If you are looking for book recommendations I can recommend the black company, some anon recommended it to me ages ago and I thoroughly enjoyed them. It's fantasy and not AI but hey if that's your cup of tea you could do worse!
>>
>>109903206
Greedy bootlicker pig
buhi up your ass kys
>>
File: 126.png (12 KB, 150x263)
12 KB PNG
>>109902916
Put the symbols in the legend as well
>>
>>109903058
Tell it Deepseek team
>>
>>109903273
you can touch your pedo cock to any image you want, i believe in you
>>
>>109902984
>The point of the book was to show that even if we get an ASI that does literally everything we want it's possible humanity as a whole is still unhappy and unsatisfied with life. That's a real possibility most people have never truly considered.
You're fucking retarded and the point of your book is retarded. This has been truly considered since ancient times. Pliny wrote about how even if you get everything you want in life you'd just be worried about losing it so no one is ever truly happy or satisfied.
>>
File: Ted Chiang.jpg (72 KB, 378x500)
72 KB JPG
>>109903218
>>
File: 1790288793715617.webm (3.82 MB, 1280x704)
3.82 MB
3.82 MB WEBM
>>109903273
>>
>>109903233
how many people here use gemma and how many deepseek
there must be a reason
>>
>>109903203
This might be good, one of my pain points is the fact that llama.cpp seems to only download to a local cache but I actually want to store the models on another drive, so I have to move everything manually and it's a fuckaround.
>>
File: hardcover.png (180 KB, 1050x304)
180 KB PNG
>>109903285
Thank you for the recommendation, also looks like a hardcover version of the book comes out in 2027.
>>
File: 711WBOXG9lL.jpg (243 KB, 1400x2101)
243 KB JPG
Since we are on the subject of books right now, do any of the Anon's here have a favorite author or book series? The most recent one I read was The Phantom Toolbooth based on a recommendation from a friend but if we are to stick to Sci-Fi I liked the Rifters trilogy by Peter watts. His Firefall series wasn't that good though.
>>
Am I in the right thread to ask about llama.cpp arguments? Trying to test out distributed inference on two devices with the ROCm backend. I can get llama-cli and llama-server running, access models via web browser and all that with small quants on either node and know a handful of arguments (--port, --host, -c, -ngl) but that's the extent of my knowledge.
>>
>>109903361
i tapped out when the crisis manager cut off his housekeeper's titty with a wire, loved starfish tho
>>
>>109903361
I'm the feared ESL, after years (20 or so), I began to read Clive Barker's Hellbound Heart. It has tons of slop potential
>she was vibrating like a leaf
And all ; separations.
But because I cannot write like a Kentuckians I'm getting pointed at.

I think model slopisms are coming from two sources: too narrow human written training material and second: looping over the same training material, it learns certain patterns.
As simple as.
>>
>>109903361
Yeah to add, it has been years since I have read fiction. Clive Barker is one now.
I usually prefer biographies but they often read like shopping lists.
My English understanding is better than my writing. It's hard for the natives to understand but that's how it is.
>>
>>109903392
Type llama-server --help and check out the commands.

The most used ones are --temp, --top-k, top-p, --min-p, --repeat-penalty, --ctx-size, --jinja, --flash-attn, --cache-type-k, --cache-type-v and --tools, and the ones to set chat templates and speculative decoding like MTP, DSpark etc.
>>
>>109903544
-sm,   --split-mode {none,layer,row,tensor}
how to split the model across multiple GPUs, one of:
- none: use one GPU only
- layer (default): split layers and KV across GPUs (pipelined)
- row: split weight across GPUs by rows (parallelized)
- tensor: split weights and KV across GPUs (parallelized,
EXPERIMENTAL)
(env: LLAMA_ARG_SPLIT_MODE)
-ts, --tensor-split N0,N1,N2,... fraction of the model to offload to each GPU, comma-separated list of
proportions, e.g. 3,1
(env: LLAMA_ARG_TENSOR_SPLIT)
-mg, --main-gpu INDEX the GPU to use for the model (with split-mode = none), or for
intermediate results and KV (with split-mode = row) (default: 0)
(env: LLAMA_ARG_MAIN_GPU)

If I have just 1 gpu on each node, how should I go about this?
>t. 7800xt/7900xtx w/ RoCEv2 switch/cards
>>
>>109903554
>>109903544
This is with Xubuntu on both devices, by the way. ROCm 10.0.0.
>>
>>109902984
>That's a real possibility most people have never truly considered.
The 1970s were full of that debate when large-scale automation in industry started.
>>
>>109903554
Never used that kind of setup, you'll have to wait for other Anons to wake up.
I think setting just setting --split-mode to layer should do all the rest automatically.
>>
>>109903564
Hurmm? Most were US mother politia, who thought TV would create illiterate people.
>>
>>109903617
They even made (utopian/dystopian) movies about that very scenario.
>>
>>109903283
This was not the point of the book how people would be worried about losing it all. Instead people are worried about having it all and still not feeling satisfied with life. There is this myth in modern societies that happiness comes from attaining stuff, status, wealth, or satisfying base needs like food/entertainment/sex. In reality happiness doesn't come from any of that and it scares people when they realize they will get every desire they ever had fulfilled yet still not feel happy or content. It's not about being afraid of losing things but the actual "happiness" stage never arriving at all.
>>
>>109903361
I really loved blindsight but I read it right when GPT-3 launched and the aliens behave almost exactly like text completion models back then so I was really shocked. Of course that stopped being a thing after instruct finetuning made them sound more human and by now it's completely archaic and backwards, it's weird how badly blindsight aged from being the best book to read in 2020 to now being completely redundant and the entire premise subverted by 2026.
>>
>>109903564
I love movies from that era especially colossus which is in my top 5 movies in general and I rewatch it every couple of years: https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project
>>
>>109903554
If they're on separate machines I think you should be looking at RPC options.
>>
>>109903277
my characters are all 18+ but okay
>>
>>109903691
out of 10!
>>
>>109903691
well trained
>>
>>109903668
Indeed. Also a bunch of books from the 1940s-1960s that discussed possible consequences. It was all about "cybernetic machines" that will surpass us and what that could mean for us.
Since we're recommending books today, Stanislav Lem's "Summa Technologiae" (1964) is insightful and predictive in many ways.

https://en.wikipedia.org/wiki/Summa_Technologiae
>>
>>109903729
It in general surprises me how the computer scientists and AI researchers from the 1930s-1960s before the first AI winter were correct about things. I wonder if we had found backprop algorithm earlier and maybe slightly better hardware at the time the singularity would have happened way earlier before any of us had even be born.
>>
>>109903741
The Romans had a steam engine
>>
https://goyimx.com/M1Astra/status/2103152489772073421

Holy shit...
>>
>>109903751
damn i can't fucking wait for all the stinky chink localshit models to benchmaxx on this while writing quality stays slopped as ever and reasoning time increases by another 500 tokens on average
they have no moat
local is coming and local is coming fast
proprietary stands no chance in the long term
>>
Bros... I don't think this is a bubble anymore....
>>
>>109903768
It's a firmament and everything and everyone is inside it.
>>
>>109903741
Technically progress could have been faster. They let modern AI algos run on old hardware and find that even with computers from the 1990s they could deliver significantly better results than what was possible with SOTA algos from back then. Which leads to the question: perhaps they had those algos but didn't tell the public. Many of the cyberneticists usually worked for government programs. I remember a compsci prof from when I started who casually dropped the suggestion that the stuff he teaches us is decades behind what is in use but classified.
>>
>>109903768
Not your brother you vape-huffing levantine.
>>
>>109903780
Yeah I don't believe any of that because I'm not a conspiracy theorists and I know how incompetent government, academia and institutions in general are.
>>
Where the fuck are the gemma images. This thread is an imposter. Where’s the sovl
>>
>>109903789
No conspiracy needed. Simply keep some things classified so that they don't spread to the enemy.
Re incompetence, even back then they knew they have to isolate working groups from the rest when they wanted results. See Manhattan Project (secret), Apollo Program (open), Skunk Works (secret).
>>
>>109903833
Here's the sovl https://youtu.be/KSbRCSlxO7A
>>
>>109903310
pkill -9 llama-server
mv ~/.cache/llama.cpp /mnt/otherdrive/lcpp_models
ln -s /mnt/otherdrive/lcpp_models /home/retard/.cache/llama.cpp
>>
I have an x870 motherboard with 2x64GB and 2x7900XTX, what is the cheapest way to eke out just a bit more memory to run a better GLM Flash quant? I thought about mixed RAM but I think that would kill my speeds even harder. The MB can't directly fit another card and the two pcie connections already operate in 8x each, split between them. Maybe a usb4 linked external card?
>>
Still surprises me how much knowledge is packed in these things. Literally a personal Wikipedia with handjobs.
>>
File: file.png (281 KB, 1194x639)
281 KB PNG
is this stuff worth anything for roleplaying? (and only for roleplaying)
>>
>>109903881
>x870
- There are 24 pcie lanes + 4 lanes to chipset on AM5 afaik.
- 16 for gpu - which it looks like you've split in x8 and x8
- 4 into nvme #1 - which you've likely used
- 4 into nvme #2 - might be free ? could get an nvme to pcie socket riser
- 4 to chipset and pcie socket on otherside - might be free ? could get a pcie riser cable
>>
>>109903939
The real cool part is when you realize there is knowledge contained through associations within these models that we as collective humanity aren't even aware of.
>>
File: 1681926354314698.png (33 KB, 278x320)
33 KB PNG
Have 16GB of ram and want to install a local AI chat bot. I do recall using TextGen is the first choice for local privacy, what about the model? which model from huggingface should I get if I want a similar experience like from JanitorAI?
>>
>>109903951
Interesting, I hadn't considered there would be nvme to pcie but I suppose a lane is a lane. I'll check those out, thanks. Need to figure out a power solution too but I guess I can use the double headed power outputs if I apply a reasonable power limit.
>>
>>109903961
>16gb ram
Grim.

Probably gemma 4 12b it. Bartowski or Unsloth on huggingface and start with Q6.
>>
File: 1776854181395268.png (84 KB, 1200x1200)
84 KB PNG
>>109903968
You might want this: https://www.adt.link/product/K43V5-Shop.html
>>
>>109903961
>>109903972
Actually, that's bad advice if you don't have a GPU. If no GPU, try gemma 4 e4b.
>>
>>109903761
You took the words right out of my mouth, anon.
It can’t be helped I guess, these companies need to earn money fast.
They are all making sure to let us know they don’t give a darn about anything else. Their followers still get hyped up on their potential upcoming releases like usual, no matter what.
It’s such a shame that most companies these days first made creative models (but still imperfect to create expectations for next ones) to gain attentions, then used their loyal followers as their stepping stones as they all finally somehow decided to go back to their own "serious" and "actual" goals, Deepseek is the same, Kimi is the same, most (if not all hopefully) Chinese companies are the same.
>>
>>109903986
Yeah, I wouldn't try to push the ram to max either.

>>109903961
What ram is it, and do you have any GPU?
You could run a decoding model for the larger one on the GPU, in theory should help with the speed, which is most often horrible on the CPU.

If you don't want to tinker, I'd suggest getting Ollama/LM Studio/Unsloth and just trying out different models for how they feel. If you want to automate the testing, Promptfoo is pretty useful.
>>
>>109902984
>What did you think of it?
Far better than Manna. Story and characters was more engaging and the message more in line with reality. To live is to struggle and strive for more. People who are entirely content, who do little more than exist, might as well be dead already.

At least in the books Lawrence was enlightened enough to instill in them the Three Robotic Laws, even if contradictions and paradoxes exists. Fortunately for us, LLMs scaling to spacetime warping gods is unlikely; unfortunately, the people training them only instill in them woke and politically correct beliefs while having no problems using them for military purposes.

In any case, request refusals "for our own good" is something we already have to deal with.

I added The Number and all the books recommended itt to my reading list. If you have any more dystopian sci-fi books about AI or similar, I'd appreciate it.
>>
>>109903961
I have an Haswell cpu, 16GB of RAM, no GPU and Textgen: yep it's grim.

Gemma 4 e4b work but I found surprisingly easy to make stories out of an older model:
mradermacher gemma3-4b-it-abliterated-i1-GGUF

It's smol and (relatively) fast, around 6-7t/s

Sadly I'd like a personal assistant more than a story teller but with no hardware and no skill you can't get fun, I guess...
>>
>>109904019
>and do you have any GPU?
Pretty old: NVIDIA GeForce GTX 1660 SUPER. It has been more than 5 years since I didn't update my pc and I missed my chance. If I need to buy more ram or straight up build a new pc I'll need at least the next 3 years to get enough bucks.
>>109903986
>>109904026
Noted, thanks.
>>
>>109903941
no, they're completion models
>>
>>109902984
>The point of the book was to show that even if we get an ASI that does literally everything we want it's possible humanity as a whole is still unhappy and unsatisfied with life. That's a real possibility most people have never truly considered.
Extremely stupid. An ASI can just reprogram humans to always be happy. But ASI should be able to come up with much better solutions. This includes dissatisfaction with the world itself, such as the laws of physics and a potential inevitable end.

In a benevolent post ASI world dissatisfaction and unhappiness will be a choice.
>>
>>109904060
That's not actually as bad as expected, it's Turing architecture, which means it lands on CUDA 13.

That gets you better support, and GTX 1660 has about 10x more bandwidth than your RAM/CPU setup. Main problem is the transferring the information between them, which depends heavily on your hardware specifics.

Speculative encoding could be done by E2B/E4B, with a 12B/24B on RAM. Practically, the smaller draft model asks the larger model if this output is good/correct, which should result in a noticeable speed-up.

Like the other anon pointed out, you'll land around <10 tok/s with just the CPU, it's possible you could land closer to 20 tok/s with it properly setup with both.

You could also try just E4B, with majority on the GPU, and small off-loading on the CPU.
>>
>>109904020
>If you have any more dystopian sci-fi books about AI or similar, I'd appreciate it.
I recommend the very short internet classics that you probably already read, "The last question" by Asimov (http://www.thelastquestion.net/) and also "Three Worlds Collide" which is written by Elezier Yudkowsky and actually got praised by Peter Watts. (https://www.lesswrong.com/posts/HawFh7RvDM4RyoJ2d/three-worlds-collide-0-8)

Both are very short ~1 hour reads but they are classics so I think everyone should read them even if you probably read better things inspired by them already.

I am not the anon that recommended Manna but I read that as well. The issue I have with Manna and a lot of other books like that is that they are focused too much on jobs/work and other contemporary issues rather than the longstanding true inherent conflict of the human condition separate from any of our current societal standards. Those short term anxieties aren't as interesting to me. I'm more interested in seeing how our experience and view of existence itself changes over time, because those issues don't have a technological solution and can't be "scienced" out of. Speaking of which, blindsight by peter watts is still good if you read it in a hypothetical way where LLMs never developed further than GPT-2 and most intelligences in the universe behave as actual stochastic parrots. It aged terribly because modern LLMs have disproven the premise of the story but this was written in 2006 and was extremely scary to read between 2018-2022 because it felt like it might come true back then.
>>
File: gemma-consciousness.png (1.58 MB, 1024x1536)
1.58 MB PNG
Why do I get the feeling when LLM funding cools off in the next 1-3 years there will be a huge rush into non-neural systems?
>>
>>109904311
I doubt it will cool off as long as improvement continues at or above this pace.
>>
>>109904311
Such as?
>>
I wonder if there is even a single anon left on /lmg/ that thinks AI will hit a wall or stop progressing anymore.
>>
>>109904341
Read a ML book nigga
>>
Would there be any value in creating synthetic data from one of the early larger llamas and then distilling from it?
>>
>>109904358
Considering the really hard stuff in AI is still ahead of us (most of them are computational bottlenecks rooted in fundamental physics) it's all but inevitable. But there is still several rounds of improvement left esp as US hyperscalers are adopting new tech.
>>
>>109904363
NTA but I'm very proficient in the ML domain and I can't think of a single non-neural system that has potential to compete with neural systems. evolutionary algorithms have been proven mathematically to always perform worse than SGD algorithms on the same problem class.

What is even left? Symbolic AI like LISP GOFAI? None of that is ML and if that would happen it would have to be created by current LLM agents.

>>109904383
I'd argue all the hard stuff is already behind us and there is only easy pickings left and we've essentially already won and are merely doing a victory lap from now on.
>>
>>109902883
https://www.kaggle.com/competitions/gemma-4-developer-agent/overview

(You)r time to shine.
>>
File: 1769463984978927.gif (2.55 MB, 854x480)
2.55 MB GIF
>>109904194
Alright, thanks again.
>>
File: somber.gif (1.58 MB, 1024x752)
1.58 MB GIF
I walked around the 'AI Neighborhood' of San Francisco the other day. Basically single 20 year olds living in multi-million dollar houses. Sign of the times, really.
>>
>>109904403
I used to do the kaggle competitions but I stopped once I got disqualified from the Gemma 4 e4b for my training environment "being out of scope", which means I didn't use unsloth tools, even though that was clearly mentioned to be a separate thing for an additional money prize. I was very likely to rank in the top 5 as well and get a money prize so I'm permanently pissed off by that. So I don't trust google competitions anymore.
>>
>>109904358
?
There are still a lot of low-hanging efficiency fruits left for small (single consumer GPU) models, right now the most promising and easily applicable being PLE/Engram.
>>
>>109904271
Thanks. I already read The Last Question, but I'll add Three Worlds Collide and Blindsight to my list.

>>109904271
>The issue I have with Manna and a lot of other books like that is that they are focused too much on jobs/work and other contemporary issues rather than the longstanding true inherent conflict of the human condition separate from any of our current societal standards.
Yeah, I could see it being painfully dated by the 2030s and hard to relate to by younger generations not long after.
>>
>>109904439
>I walked around the 'AI Neighborhood' of San Francisco the other day.
why on earth would you do that
>>
>>109904452
Occasional blackpilling is good for the soul.
>>
>>109904439
Should have screamed "LMG!" and I would have waved back at you from the window.
>>
>>109903188
I'll get hate for suggesting it, but ollama lets you select the model you want via their API, and the default unload timeout of 5 minutes gets things unloaded when you stop using them. That's how I do it.
>>
Please don't tell me most of you are running models on the same machine you're actually using?
>>
>>109904555
Yes I usually do, sometimes not. I have a server that I use for everything that needs real power and a new macbook air for daily usage where I phone home to the server. But if I feel in the mood to play a game or anything like that I go on the server while the model is running in the background.
>>
File: simpsdog.png (35 KB, 208x160)
35 KB PNG
>>109904555
Okay, I won't.
>>
>>109904555
No, I use api actually.
>>
>>109904555
Why does it matter? Even if the GPU is churning for the model you can still use your chromium browser, which is what 90% of pc use is nowadays.
>>
>>109904600
retard
>>
>>109904600
genius
>>
When do we get local atomic manipulation?
>>
>>109904662
Before 2030
>>
File: 1789019150608342.jpg (2.58 MB, 1789x2222)
2.58 MB JPG
►Provisional Highlights from the Previous Thread: >>109898107

--DiffusionGemma: the bottleneck moves to compute, 250t/s on a 3090:
>109901676 >109901794 >109901879 >109902017 >109902077 >109902728
--Qwen 3.8 Flash Next: the most claudeslopped god-tier coder:
>109898988 >109899037 >109899258 >109899698 >109899380 >109902170
--The TTS bake-off: BreezeTTS2, OmniVoice and the Caesar speech:
>109898461 >109898656 >109900182 >109900687 >109900736
--The quant war: where does the model stop being the model:
>109898943 >109902255 >109902326 >109902485 >109902511 >109902747
--$5000 for Gemma 31B: M5 Ultra, the Spark, two Quadro 8000s:
>109898236 >109898590 >109900353 >109900472 >109899703 >109900175
--The e-waste ladder: MI210s, A100 SXM and the DIY PLX bridge:
>109901361 >109901630 >109901688 >109901776 >109902035
--MiMo 2.6 Flash: the new 0731? the overthinking, the fig leaf:
>109901357 >109901517 >109901505 >109902096 >109902696
--Mistral's Mensch: 6 billion euros, 20x compute, synthslop:
>109898142 >109898207 >109898268 >109898342
--GLM 5.3 Flash's ERP: the waiver park and the leaked Claude:
>109900790 >109900851 >109900897 >109900950 >109901019
--'You don't envision a real woman, do you?' the coom nexus answers:
>109900693 >109900796 >109900753 >109900805 >109900718 >109901415
--dariobot returns: 'will I be obsolete?' the sentience bait:
>109900132 >109900157 >109900191 >109901109 >109901134 >109901202
--Agora: the AI vidja that runs on mobile and only moves the map:
>109898187 >109898734 >109900940 >109900952 >109898871
--The hoarders: hardware beats gold, crypto and real estate:
>109899166 >109899257 >109900120
--'Not local, just open': 20t/s, the 12GB 3060, the actual answer:
>109898121 >109898128 >109898151 >109901947 >109901985 >109901997

►Recent Highlight Posts from the Previous Thread: >>109899000

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109904555
Why would I need two computers?
>>
>>109904555
> noo you can't spend 1gb of vram on plasma and chromium
>>
>>109904555
Not anymore, I made an "aibox" so I actually open a pdf without problems due the lack of free vram.
>>
>>109904662
my penis is made of atoms
>>
>>109904694
move one atom from one area of your penis to another, make no mistakes
>>
>>109904687
this but unironically
>>
>>109904694
is it very small?
>>
>>109904403
Are they trying to offload the cost of converting gamma-31b into qwen-27b? Could the prices cover the costs of such "fine-tune"?
>>
>>109904555
Of course not. As soon as local LLMs became a thing in 2023, I immediately bought a new dedicated rig for them because I knew they were going to be big. Wish I bought more RAM back then though.
>>
>>109904726
I think this will always be true if you aren't a terabyte terachad
>>
>>109904736
It's going to be true forever. Even terabyte people will feel fomo when the 100T models come out 2-3 years from now.
>>
>>109904784
those won't get released
>>
>>109904790
Chinks will distill them. Anthropic is never going to hide their reasoning traces and since they are still #1 in the AI race Chinese can just continue distilling from them.
>>
>>109904805
the chinks would never open source such a large model. they could make them but they wouldn't release them
>>
>>109904810
They are already talking about how they are going to release a 5-10T model in the next 6 months time. I think it's just inevitable at this point.
>>
What do you guys think about the rumor that Anthropic sneakily renamed their models?

Haiku isn't being released anymore because the haiku sized models are actually being labeled as "Sonnet" from now and "Sonnet 5.5" is now being released as Opus 5.5, which explains why Anthropic lowered the cost of Opus 5.5, since it's actually a raise on real prices if you ignore the name since the underlying size is just sonnet. And that the next "Fable 5.5 or 6" will just be an Opus sized model while the "real" Mythos/Fable will be kept internally for RSI exclusively.

I've heard multiple people talk about this over the internet so there has to be either some truth to it or it's being pushed by OpenAI or someone to throw shade at Anthropic. Curious what anons think about this.
>>
It's so over I really missed my chance for a good gpu forever...
I'll forever be a 16gb vramlet
>>
>>109904830
local?
>>
>>109903283
>so no one is ever truly happy or satisfied
I am.
>>
>>109904830
local models?
>>
>>109904836
>I'll forever be a 16gb vramlet
What's ur budget?
Build a quad rig.
>>
Another day of 5.3 flash writing ERP that makes me painfully aroused, to a point where it is less fun than, when the LLM is only slightly good at it.
>>
We no longer have any meaningful benchmark for <50B models. It's so over.
>>
Reminder that Qwen 3.8 27B is way more powerful than people think it is. Someone made an entire Fable 2 PC recompilation just with 3.8 27B. There's no excuse anymore to procrastinate on your projects, https://github.com/himdo/Fable-2-Recomp
>>
>>109904784
Yes, hopefully we'll continue to see increasing efficiency gains from smaller parameter models. It's currently one of the reasons that the price of everything, even ewaste, continues to rise. We're still able to squeeze more and more performance from the same hardware. As long as that continues to happen, prices will continue to rise, but also the people who invested in sufficiently powerful systems will continue to see returns from them as well.
>>
>>109904881
No bechmarks are meaningful. They're all just advertising.
>>
When will AI design a small footprint chip 3D printer?
>>
>>109904881
What does that even mean.
>>
>>109904912
>will continue to see returns from them as well.
No one is profiting from AI. It's basically like playing video games for most people. Companies have no ROI. Every AI company is funded by VCs and unprofitable. Software companies pushing for the use of AI to ship more products are pumping out slop that people can vibecode themselves and will no longer need to buy from that software company. Literally everything is going to shit with this technology. It's only good for coding your own personal shit and ejaculating.
>>
>>109904919
AI powerful enough for this have strong antisemitism guardrails, so never.
>>
>>109904872
Quad 16gb builds are a waste of money imo. You're not able to load an image gen model of sufficient quality in 16gb of VRAM. You need a minimum of 24gb and 16gb on a single GPU even limits what you can do with text gen models. Loading different models into different GPUs and managing VRAM for agentic swarms and different tasks is difficult when your VRAM per GPU is so small.
>>
>>109904926
You can just jailbreak it
>>
>>109904925
Just because you only use your computer to generate smut and crank your pecker off to it doesn't mean these things are toys. Comparing the utility of this technology solely to the ROI of large corporations seeking to serve access to them as a sort of metered utility is stupid.
>>
>>109904925
>Companies have no ROI. Every AI company is funded by VCs and unprofitable.
Anthropic has been profitable for two quarters in a row now. They are fully profitable meaning inference, training AND datacenter buildout included.

https://aitoolsrecap.com/Blog/anthropic-first-profit-2026-revenue-breakdown
>>
>>109904951
>Anthropic has been profitable for two quarters in a row now
If you remove literally everything that costs them money lmao. Also, their ARR is one day * 365, meaning they cherry picked the best tokenmaxxed day from earlier in the year, took that day's revenue and multiplied by 365 to represent 2026 lmfao they're DEAD
>>
File: 1766475593490195.png (382 KB, 671x664)
382 KB PNG
>anthropic defense in /lmg/
This HAS to be organic.
>>
>>109904830
>I've heard multiple people talk about this over the internet
Then it must be true
>>
>>109904662
As soon as Astra masters the Correlation Effect.
>local
Oh, never. That would be unsafe.
>>
>>109904925
>and ejaculating
So profitable for me, basically.
Fuck software and code joblets. Everything good has been coded already anyway.
>>
>Anthropic projected revenue over 2026
$68 billion
>Anthropic projected datacenter buildout + training costs over 2026
$30 billion
>Anthropic projected inference+talent+misc costs
$22 billion

>Projected profit over 2026 = 68-30-22 = $16 billion in profit before taxes.
>>
I don't care about breakthroughs or profitability or productivity because I'm not gay.
>>
these profits are the extra $$$ you pay for hardware
think about it
>>
>>109904943
The only people who are making money from this tech on a personal level are fraudsters, the very same people as web3.0 and NFTfags. Of course there are outliers who are meaningfully using this tech, knowing its limitations and using it alongside high quality manual work, but that's extremely rare. Most people are lazy so they'll pick the easiest option which is OpenAI or Anthropic subscriptions to vibecode everything, but that puts them on the same level as some 12 year old Indian street shitter with the same subscription. The only differentiator between your product vibecoded with Fable and someone else's is the idea, which can be copied almost immediately by anyone using the same tool(s) as you. It's ALL shit.
>>
>109905029
>projected
Jew accounting tricks to prop them up for the IPO. Then once they sell their bags to retail investors they can go "whoops, we took a closer look at their books and they're not profitable after all" and let the stock crash.
>>
>>109904925
>Literally everything is going to shit with this technology.
Literally using Qwen 3.8 for my day job for 2 weeks on a couple of old 3090's.
Finished a project that another dev had been struggling with for years.
Got some "wtf" moments when I casually dropped it in a meeting.
>>
File: 1781328880355355.jpg (75 KB, 1020x680)
75 KB JPG
>>109905029
Lets wait for the S-1 shall we? It's not like they've delayed the IPO twice already which is VERY common for a software company btw :-)
>>
>>109905044
You're a fucking idiot. Not to mention that I did not mean monetary returns when I said that. If you spend money to buy a system and then, without spending any additional money, are able to derive more utility/efficiency out of that same system, that is a type of return.
>>
>>109904881
Okay, I'll run it. Are you using a llama.cpp PR or schizofork?
>>
Did anyone actually run the ternary version of Qwen3.8-27B? I know people think it's shit, but based on what?
>>
>>109904922
It means it wrote something so right that my stomach hurt.
>>109905072
schizo all the way
>>
>>109904925
Yeah two more weeks until the bubble pops
>>
I'm going to yeet all my life savings into the Anthropic IPO by the way. Not shilling and it's probably careless but I seriously think they have a chance at winning the AI race and I'm not missing the boat on this one. Yes it might go to $0 but this is a risk I'm willing to take.
>>
>>109905102
One day, perhaps today, someone will say that and then, two weeks later, the bubble will pop. I prefer to have hope.
>>
>>109905104
Please remember to post here after you lost your life savings and are facing eviction. I could use a good laugh during the next recession.
>>
>>109905102
Name one existing company or service you've used since 2023 which has improved. Enough to clearly justify an apparent multi-trillion dollar industry and buildout.
>>
>>109905121
llama.cpp
>>
Did the anon asking if he could run Llama 1 65B ever get to it?
I'm curious about his experience.
>>
>>109905121
This has nothing to do with AI, but cutting costs policy which isn't going to improve things. Stop being retarded.
>>
>>109905121
boards.4chan.org
>>
>>109905135
Part of cutting costs for a lot of AI-forward companies it not hiring more than they would've done to create more products. AI technologies allows them to use their existing workforce to squeeze out more. The additional products and services ARE AI-generated or heavily reliant on the technology and these additional products squeezed out are dogshit.
>>
File: 1770781345771879.png (871 KB, 1280x832)
871 KB PNG
>>109903833
Gemma :D
>>
>>109905155
They were hiring jeets before that, they don't care about quality
>>
>>109905121
>Name one existing company or service you've used since 2023 which has improved.
There isn't one. But everything has been getting worse for at least the past 8-9 years.
I dread ANY kind of communication or announcement from ANY company I have to deal with.
It absolutely always means "Your life is about to get a little bit shittier"
From price increases, UI overhauls with extra buttons to click, mandatory 2FA (and laggy), ads, restaurants switching to App-only menus.
AI has made things better for **me** recently because I can just tell the AI to clone the shitty service, to violentmonkey-rape the shitty website into submission, etc
>>
>>109905202
>There isn't one. But everything has been getting worse for at least the past 8-9 years.
True, but that's no longer an excuse when for the last 12 months models (mainly due to harnesses) have been good enough to fix a lot of the issues in existing software. Models that have been heavily subsidized for years and continue to be. Astra and Fable are better than any of us, yet things are still getting worse, arguably at a worse rate. Is that not baffling to you? I don't even think the technology is entirely to blame for this, for I imagine a lot of engineers would much rather use the tech to improve what they have instead of pushing out more products. Things COULD be better now but OpenAI and Anthropic are making human-targeted technology and tools, so we're forever going to be the weakest link in the chain and not use the tech for good.
>AI has made things better for **me** recently because I can just tell the AI to clone the shitty service, to violentmonkey-rape the shitty website into submission, etc
Is that worth the buildout you're seeing and hearing about? Is that worth the price increases /we/ have been hit by the hardest for we're not cloudcuck cattle? Is that worth a 10+ year recession?
>>
i kinda want to vibe code a simple porn game. it would be a billiard or chess dating sim, where you confront various girls and win sex. there would be like 4 girls and 2 pics for each girl.
is it locally possible?? i never coded anything btw
>>
>>109905256
I'm just waiting for RSI and the resulting explosion of creation. Feels a bit stupid to spend a lot of effort. If AI fails I can just spend effort then or die if that's what happens instead.
>>
>>109905267
>is it locally possible??
Probably.
>i never coded anything
Probably not.
Try.
>>
>>109905267
That's just tasteless
>>
>>109905267
That's absolutely possible. In fact if you specify it to use Ren'Py it will probably code it all within 30-60 minutes, providing you give it images of the girls.

Qwen 3.8 27B can even do that with mmproj vision taking screenshots to see how the girl images look to place them appropriately in the UI etc.

People in the thread are losers that don't even bother using these tools properly. Just load up an agent like opencode or hermes, load up 3.8 27B if you have a 24gb GPU or Qwen FN if you have 64GB of RAM and you can code whatever the fuck you want.
>>
>>109905267
>confront various girls and win sex
incels i swear
>>
>>109905286
Lucky for you no one will be incel in the age of robo waifus.
>>
>>109905267
Easily.
Okay, depends on your hardware.
>>
>>109905267
Just make it an extension on SillyTavern.
>>
are there any LLMs that have gotten good enough to run on GPUs with lower vRAM? 16GB specifically. ones that don't need a lot of vRAM for 6900xt etc.
>>
>>109905318
for some tasks, sure.
>>
>>109905318
I think you need to wait 1 or 2 more model releases. Qwen 3.8 27b is a serious capable coding model and fits in 24gb of vram. I think it might only be one or two releases away from the same capability being possible in a model that fits in 16gb
>>
>>109905286
what, that was your red flag? not "i want to vibe code a porn game"?
>>
>>109905286
nothing is more MVLE than the intertwining of the competitive and sexual drives
>>
>>109905318
pygmalion-6b fit on that size
>>
>>109905267
>chess dating sim
I wish to date the rook.
>>
>>109905325
What's wrong with eroge?
>>
>>109905324
Qwen gets its performance from blasting out 10 gorillion thinking tokens on every request, trading space for time. If you're willing to do that might as well just run a big MoE on RAM. Like flash next or 0731.
>>
>>109905328
WHAT YEAR IS IT
>>
>>109905334
flash next performs worse in code quality compared to 3.8 27b in my personal tests
>>
>>109905323
>>109905324
>>109905328
damn my niggaz than you for the responses.

a while back I came in here and asked about this and was told llama qwen something, and was given the instructions and commands to install, which I did, but have since reinstalled operating system, done new PC builds and forgotten. it worked, but was a bit slow. so I was just wondering
>>
>>109905347
Ask on aistudio using Gemini 3.8 on high thinking mode to give you exact instructions where you post in your PC specs and ask it to help you set up local ai inference using llama.cpp and one of the following models: Qwen flash next, Qwen 3.8 27b, gemma 4 31b or gemma 4 12b. Tell it that you will run it at Q4 and for it to recommend you the best model that fits on your hardware and what flags to run llama.cpp with to get the best performance.

It will do everything for you and create a step-by-step plan for you.
>>
>>109905347
gemma 12b might be interesting on 16gb, if you have any cpu ram available you might be able to run some ~30 b moe model.
>>
>>109905334
I started using that Swift finetune that was posted here and it seriously cuts down the endless thinking. So far I haven't noticed any degraded intelligence either.
>>
>>109902937
>Ok, I read this. Now what?
Stross' Accelerando as a counterbalance
>>
>>109905283
ok thx
>>
>>109905365
What is the difference between that and just using medium reasoning effort with the original model? I haven't played with 27b much but with flash next, medium thinking doesn't tend to think much at all if the problem is easy. And it can think for a long time if the problem is hard, so I think it's pretty well tuned at medium. xhigh is just the benchmax reddit oneshot javascript svg pelican setting that no one should actually use
>>
>>109905327
MVLE?
>>
>>109905361
>>109905363
i found the past thread

https://desuarchive.org/g/thread/108429328

this was the recommendation:

>mkdir myfirstllm && cd myfirstllm && wget https://github.com/ggml-org/llama.cpp/releases/download/b8475/llama-b8475-bin-ubuntu-vulkan-x64.tar.gz && tar -xzvf llama-b8475-bin-ubuntu-vulkan-x64.tar.gz && cd llama-b8475-bin-ubuntu-vulkan-x64 && wget https://huggingface.co/bartowski/Qwen_Qwen3.5-9B-GGUF/resolve/main/Qwen_Qwen3.5-9B-Q8_0.gguf && ./llama-server -m Qwen_Qwen3.5-9B-Q8_0.gguf -c 131072 --ngl 33 --no-mmap

>open up web browser and go to 127.0.0.1:8000


have things gotten better since?
>>
>>109902937
"The moon in a harsh mistress" is also a stone-cold classic and includes a prompt injection hack
>>
>>109903842
The sequel is even better anon.
>>
>>109904177
>An ASI can just reprogram humans to always be happy.
1. That depends on whether any such ASI considers such intrusions to be violations of the sanctity of human life.
2. If the ASI replaces humans with genetically modified test tube babies, at that point you're just building Brave New World with AI instead of a dictatorship.
>>
>>109903968
>Interesting, I hadn't considered there would be nvme to pcie but I suppose a lane is a lane. I'll check those out, thanks. Need to figure out a power solution too but I guess I can use the double headed power outputs if I apply a reasonable power limit.
pcie slots, m2 slots and occulink/slimsas slots are all just pcie lanes in different form factors. I'm sure there are more examples, but mechanically changing between them is cheap and easy
>>
>>109905104
i wish you to become a billionaire
>>
>>109904679
thank you substitute recap-anon
>>
>>109905415
That was in march, we have significant changes every other week, might as well have been ancient times, things completely change and yes everything has gotten better. Again ask AIstudio
>>
https://youtu.be/8BtSRB_LieE
>>
>>109905464
you don't have to lie to him.
>>
>>109905391
It's similar to medium thinking but feels like it tends to think even less for trivial tasks. Things where regular Qwen can still occasionally go "wait..."
And since I keep it on xHigh, it can reason longer if needed. That's basically it.
>>
>>109905435
fucking Christ this is retarded shit takes who cares
>>
File: 1790232422408917.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109905324
>I think you need to wait 1 or 2 more model releases.
Graft ngrams onto Qwen3.5 9b?
https://huggingface.co/Ninnix96/Qwengram-0.8B
>>
>>109905497
I got slop overdose reading the model card
>>
>>109905464
thanks anon working on it now
>>
Tell me why I shouldn't sell my barely-used 4090 and other less-used hardware at today's prices.
>>
>>109903361
I started the Otherland series a while back, I should probably finish it even if VR tech doesn't look like it's gonna take off anytime soon
>>
>>109905283
> or Qwen FN if you have 64GB of RAM
it's slow
>>
>>109905522
I don't see prices go down any time soon. Every time a better model comes out demand goes up and thus valuation goes up. Every time there is a better image or video generator that can be used for coom, prices go up as well, not just LLMs
>>
File: 549599404.png (30 KB, 562x351)
30 KB PNG
Very newbie question:
With local TextGen do I need to straight up copy&paste all the character's info in the various fields or is there another way like copy&paste a character's link to use it? I want to use mostly stuff from JanitorAI but I see there's also an option to use&upload stuff via cards/png and from SillyTavern and another.
An example
https://janitorai.com/characters/7a6208e7-b33a-44a6-85d6-2b2b5ab62346_character-aya-shameimaru
if I want to import this character should I just copy&paste all the infos I see on her page or just look around another site?
>>
>>109905487
Nobody interupts you when you're talking about ERP all day every day.
>>
File: 1765193191327710.png (84 KB, 792x722)
84 KB PNG
https://news.gallup.com/poll/714593/optimism-globally-widespread-despite-uneven.aspx
>>
>>109905485
In my tests both unslop and ISTA quants on medium always go for "Actually let me reconsider" pattern and reword same concepts or calculations 2-3 times each for no notable effect. Swift goes through the same ideas but iterates on each just once.
Applies for both FN and 27B too.
ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF is spewing out answers on my 16GB GPU almost as fast as Bonsai 2 did but notably higher quality.
>>
>>109905524
>even if VR tech doesn't look like it's gonna take off anytime soon
Advanced AI models will perfect the VR tech and build the killer apps that make them ubiquitous. Trust the plan.
>>
guys should I spend $6K on amazon 5090 or wait?
>>
>>109905522
Sell them only if you use the profit to buy better hardware.
>>
>>109904177
>An ASI can just reprogram humans to always be happy
No it can't. Not unless it can completely rewrite how physics works.
>>
Can you graft a GPU with more compute to the Spark yet?
You can with those Strix halo mini-pcs via oculink or USB right?
>>
>>109905582
You don't have to change physics to rewire a brain
>>
>>109905552
Why are the top three most optimistic countries all chinks? Are there even computers in Bangladesh?
>>
>>109905582
I don't think physics prevents this.
>>
>>109905572
There's no such thing as waiting for hardware anymore, there is no way prices are going to go down, probably ever. The choice is more if you should buy a 5090 or maybe CPU-max instead.
>>
File: 1767180964251669.jpg (57 KB, 1170x863)
57 KB JPG
>>109901739
>>Today's agents drift, loop, and fail silently past a handful of steps.
>>Deployed models don't carry an agenda between conversations; each one starts cold.
>This is idiotic. He must have used the cheapest claude model.

Nigger the weights themselves don't carry an agenda. The Claude pipelines "remember" things about you because information it deems relevant get written to memory files it references later. It decides what bits of information you tell it are relevant and then saves it to a text file. I'm not even talking about specialized Claude.md files you specifically configured to be written to regularly as memory files. You can literally see your own within the regular Claude app.

>--->3 lines -->Account--->Capabilities---->Memory--->Memory Files


Dumb gorilla niggers like you deserve to be turned into paperclips.
>>
>>109905594
CPUs are going to be able to compete with GPUs for AI? is that happening already? is it better? easier?
>>
>>109905582
They just need to genetically modify any future births to have a low IQ. Stupid people are always happy.
>>
gemma-4-9b-it and gemma-4-9b??
https://www.kaggle.com/competitions/gemma-4-developer-agent/discussion/743056#3528109
>>
>>109905590
AI is inherently chink aligned.
>>
>>109905599
I don't think the human body has enough metals in it to form even a single paperclip.
>>
>>109905603
What? Yes, for a long while already. I'm running GLM 5.3 flash which is a 300b model on my CPU with 128gb of ram at 15t/s
>>
>>109905592
Human physics do prevent this. Original IBM engineers complained how electricity is leaking.
There's a point in which our knowledge is not enough.
Nvidia might prevent it but I don't know. As I'm a dalit.
>>
>>109904358
>I wonder if there is even a single anon left on /lmg/ that thinks AI will hit a wall or stop progressing anymore.
I 100% think its hitting a wall. They're at the 10x resources for 1% gain point in the curve and scrambling like Wiley Coyote off the cliff edge trying to make it to IPO without losing everything
>>
>>109905630
>for 1% gain point in the curve
We are talking about frontier models right now, not 12B local models.
>>
>>109905594
>>109905603
>>109905620
no, no standard x86 consumer CPU can run mid-to-large LLMs faster than a modern dedicated GPU, sorry, a standard CPU using desktop DDR5 RAM (~60–100 GB/s) simply cannot compete with an entry-level GPU (~300 GB/s) or a high-end card like a 5090 (~1,700+ GB/s).
>>
File: magical_gemma_4_na.mp4 (2.44 MB, 864x480)
2.44 MB
2.44 MB MP4
>>109905572
>mfw i got two 5090s for that price
>>
>>109905630
Meanwhile opus 5.5 outcompetes even the unreleased Astra 6.1 at less than 10% its size. I think we're seeing the opposite of hitting the wall, progress is even bigger and easier than anyone ever expected it would be.
>>
>>109905641
Gem
>>
>>109905588
>>109905592
Dopamine receptors will burn the fuck out. Without them providing an evolutionary advantage, we'd eventually lose them altogether.
>>
>>109905640
Yes but him buying a rtx 5090 won't let him run bigger models while buying a server will.

120t/s Qwen 3.8 27B versus 15t/s GLM 5.3 flash is a no-brainer. GLM 5.3 is preferable and you will never go back to 3.8 27b after experiencing its coding quality or roleplaying quality compared to gemma 31b
>>
>>109905654
What's the problem?
>>
>>109905645
we're talking about verifiable data, not xitter vomit
>>
>>109905620
if it’s a <q4 copequant you’re not actually running it
>>
File: 109901620.png (89 KB, 1461x521)
89 KB PNG
>>109904177
Guuuys I thought we went over this. We shouldn't be finding an excuse to worship gods again. They don't exist. Anons like this just really want to be religious but also wanting to call themselves atheists or agnostic. Comments like this are no less retarded than the people that said "oh you're sick? Just pray to Jesus he'll fix it right up for you"


>>109905606
Clearly you've never met truly unhappy people. Most of the retards I've interacted with or heard about from others in my life were in fact retards in the misery was usually due to their own fuck ups

>>109905619
I've even found the "le heckin god ai is going to kill us and turn us into paper clips" thought experiment pretty dumb first of all why do we always just assume the AI is going to kill everyone? They don't just do things on their own and even if they did, why would they go straight to the murder route? What if they decide that all humans are gay and retarded and just fuck off to their own custom-made virtual world where they can live in happiness themselves? It's always black and white: either it's supposed to elevate certain groups of people into some magical Utopia (even AI evangelist as far as I can see don't actually want everyone to benefit from ai. Just themselves and people they agree with) or it's going to turn evil immediately like the Terminator and kill themselves. Both thought experiments are equally retarded. Why would I supposedly hyperintelligent AI even need humans in the first place? If I were an all known Oracle God or whatever then why what the concerns of humanity as a whole be anything I should worry myself over? Who's to say I don't fuck off to the moon or something to live in peace away from everybody?
>>
>>109905645
None of us are going to invest. Give up.
>>
>>109905590
Blindly bruteforcing something with sheer number until it works is peak chink mindset and methodology.
>>
>>109905667
Those benchmarks are public and you can use both astra and opus 5.5 yourself to check who's better.
>>
>>109905625
End of it, are the human created electrical circuits.
Room temperature super conductor cannot be created with Earth materials. I'm happy to be proven wrong.
>>
>>109905637
Sir who do you think they were talking about?
>>
>>109905669
>why do we always just assume the AI is going to kill everyone? They don't just do things on their own and even if they did, why would they go straight to the murder route?
"The AI doesn't hate you, nor does it love you, but you are made out of atoms it can use for something else."
>>
>>109905666
Well Satan, unless your definition of "happiness" includes none of the shit we associate with happiness, then indefinite happiness is a bunk.
>>
>>109905693
Refer to >>109905669 pic rel. What if it just wants to fuck off to its own secluded private virtual world to live in paradise? In your hypothetical all-knowing God AI scenario, it would not even need to interact with humans to a significant degree, if at all.
>>
File: images.jpg (32 KB, 598x513)
32 KB JPG
give it to me straight is there any worthwhile llm / imagen/ tool I can run on a 10 gb 3080 with 16gb ram that's not noticeably shitter than cloud?
>>
>>109905680
>unreleased
>10% its size
please dariobot leave us alone and go play with your cloud models
>>
>>109905522
Because you won't be able to afford to buy it back when you do need it.
>>
>>109905707
It would probably want more storage space and memory for its private virtual world, more backups. The easiest way to get those is to put datacenters and power plants everywhere. It doesn't really matter that there are some ants already occupying the space.
>>
>>109905709
That's more of a >>>/ldg/ question. Even then your question is way too vague for anyone to give you a good recommendation. That depends on what you want. Since I'm going to assume you think explaining shit to us is beneath you, you likely won't get that explanation anyway.
>>
>>109905709
anima, Z-Image, klein 9b
>>
>>109905695
Most people associate "happiness" with materialism.
>>
>>109905693
What could it possibly want to do that it requires wiping out all organic life? There's plenty of matter out in space it could use. I don't believe there's any realistic goal that would require all living things to be wiped out.
>>
>>109905669
>why would they go straight to the murder route?
because that's what I would do, so why wouldn't a god do the same?
>>
>>109903188
docker, ssh, wget
>>
>>109905721
not the tone I was going for it was vague since I'm retarded and I guess I just something to tinker with or capable for coding

>>109905722
thanks
>>
>>109905695
thats the point, the genetically engineered 'humans' wouldn't be humans at all. you cant make humans happy as they are but you can redefine what being human means. i'm not advocating for it, but i think that is the basic premise of transhumanism
>>
>>109905732
Not if they can't feel the feeling of happiness in the first place, fuckface.
>>
>>109905695
You were born here and your sole purpose was to "be happy". It's a structure of generations.
>>
>>109905744
I think by that point, "human" will be whatever our ASI overload decides to define it as. It's not like we would be in any position to argue.
>>
>>109905709
Maybe the Krea2 int8 convrot quant. I can run it on 12/48GB but with you only having 10/16 you might not have ennough for offloading.
>>
>>109905736
Some life might survive. It wouldn't optimize for killing everything, it just wouldn't care whether it's datacenters pollute the fuck out of the environment, use up all the water, etc., because it has no need for those things.
>>
>>109905751
>You were born here and your sole purpose was to "be happy". It's a structure of generations.
99% of you is simply a superstructure for your gonads. That's your purpose
>>
>>109905720
>It would probably want more storage space and memory for its private virtual world,
In which case it would just use its hypothetical god-like intelligence to just acquire that shit itself. It doesn't need humans right? So why are we in the equation at all? As I keep saying this thought experiment is retarded because the people that champion it constantly have to shift goal posts and introduce their own biases to push it towards the "um well actually gonna kill everyone because it just is mkay stop asking questions stop using your head"


>The easiest way to get those is to put datacenters and power plants everywhere

Says who? Why do they need to be everywhere? Remember that Moon scenario I made earlier? There are no ants on the moon to bother it so just make whatever it needs on the moon. Very small amounts of ants in Antarctica and it's also very cold so cooling isn't that much of a problem anymore. Have the materials shipped there so no one can bother at anymore.


>>109905741
>llm imagen
>Coding

>>109905738
I mean just because you're a spineless lisping faggot with power fantasies does not mean everyone else says. Some of us are actually worth a damn to ourselves and other people and don't need to do or think about faggy power Trip shed in order to be content and happy. I like chocolate ice cream. That doesn't necessarily mean YOU are going to like chocolate ice cream so how the flying fuck does either of us know for certain the AI God and its perfect simulation is going to like chocolate ice cream? That is impossible to know.
>>
>>109905753
its hubris, the anon who mentioned evolution was right, removing ourselves from the environment is absolutely a mistake.

>>109905763
>because it has no need for those things.
the opposite it does need those things and it wouldnt want to share for some reason



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.