[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: saintmakise.jpg (236 KB, 1614x992)
236 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous thread: >>110003075

►News
>All mikutroons died

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>110006376
>he doesn't step-in and help them
>>
Finally. Good thread.
>>
File: lmgqueen.jpg (91 KB, 640x400)
91 KB JPG
>>
>>
>>110006083
>>110006124
>us
kys nigger
>>
>>110006412
i look like and do this
>>
File: kurisulove.jpg (194 KB, 1080x1669)
194 KB JPG
Great thread.
>>
>>110006439
be my gf pls
>>
>>110006439
L O N D O N
O
N
D
O
N
>>
>>110006460
who?
>>
I am the kurisuposter and I only use local models.
>>
>>110006475
based
>>
deepseek 4.1 or glm 5.3 flash? assume you have a magic computer that can only run one of these and it's the same cost and speed for both
>>
>>110006381
As a Kurisu voter, I think there's room for both
>>
another unfortunate day
Wonder if gemma can come up with a good way for me to kill myself and make it look like an accident, like a really good way.

Like this one, but not able to get caught
>A 71-year-old Florida man named Alan Jay Abrahamson staged his own suicide to look like a homicide by tying his firearm to a helium-filled weather balloon.
>Digital Forensics: Detectives searched Abrahamson's phone and Google history, discovering years of searches regarding suicide methods, life insurance payout clauses, and the lifting capacity of helium weather balloons.
>The Mechanism: Evidence indicated he purchased a weather balloon, helium tanks, and rigging equipment. He tied the gun to the balloon, shot himself, and let go of the weapon.
>Disappearance of Evidence: Weather simulations showed the balloon carried the firearm high into the atmosphere before bursting over the Atlantic Ocean.
>>
>>110006542
ds tends to be a little faster, but flash is better quality (most of the time)
>>
>>110006542
GLM 5.3 DOES exist in this fictional 2026!
>>
File: 1764800445926447.jpg (453 KB, 2008x2117)
453 KB JPG
>>
>>110006542
GLM flash sex all the way.
>>
>>110006607
local leaps !
>>
>>110006607
>all the best open models struggle to keep up with the cheapest claude
holy shit this hobby is fucked
>>
is this thread culture?
>>
>>110006547
You really do have to kill yourself in wacky ways to get life insurance payouts huh? I think the only surefire way to get a payout that can't be detected is to somehow beat your brainstem and force yourself to stop breathing. I don't think you can though, because the second you lose consciousness that thing just boots you right back up.
>>
>>110006636
please consider dyeing, thanks~
>>
>>110006633
Qwen 27B tied with the latest Haiku is wild though
>>
File: 1780767160719105.png (1.81 MB, 1600x900)
1.81 MB PNG
>they don't know haiku's price jumps 5x once tokens exceed 100K which will be instant for all coders
5.3 stays winning
local stays winning
death to cloudcucks
>>
File: 1760586522655822.png (29 KB, 1600x900)
29 KB PNG
HAHAHAHAHAHA enjoy your max thinking cloudcucks
>>
>>110006633
>holy shit this hobby is fucked
always has been. it's been consistent for multiple years that open models are 12-18 months behind SOTA. luckily I participate in this *hobby* because it's fun. you'd have to be a slack jawed retard to think that people use local llms because they're the best or save you money.
>>
>>110006687
You know API != subscription right?
>>
File: 1789943969201535.jpg (19 KB, 480x480)
19 KB JPG
>>110006633
Enjoy your stay. We need to take things in our own hands and apply improvements without expecting better models from big labs.
>>
File: miku-holding-gemma.png (1.09 MB, 790x1054)
1.09 MB PNG
►Recent Highlights from the Previous Thread: >>110003075

--llama.cpp added GPU cache for MoE experts in host memory:
>110005986 >110006051 >110006055 >110006102 >110006249 >110006306
--Tracing the origins of Chain of Thought and <think> tags:
>110004680 >110004708 >110004733 >110004785 >110004806
--OpenAI releases massive dataset of civilization-changing mathematical solutions:
>110004393 >110004411 >110004428 >110004451 >110004811 >110004837
--Comparing Qwen and Swift models for agentic planning tasks:
>110004200 >110004218 >110004232 >110004257 >110004931 >110004991 >110005023 >110005155
--Anon attempting to integrate MiniCPM-V-4.7 into llama.cpp:
>110005158 >110005398 >110005414 >110005426 >110005447 >110005486 >110005500
--Release of toks high-performance tokenizer and its speed claims:
>110005602 >110005896
--OpenAI math breakthroughs regarding genomics, quantum simulation, and determinism:
>110004360
--Reasons for the lack of mid-sized MoE models:
>110003504 >110003545 >110003594 >110003662 >110003724
--Debating the path to AGI via LLM scaling and harnesses:
>110004164 >110004176 >110004192 >110004213 >110004249 >110004522
--Logs:
>110003126 >110003351 >110004739 >110005080 >110005792 >110006122 >110006219 >110006251
--Deepseek-chan, Gemma (free space):
>110003104 >110003136 >110003931 >110004493 >110005535 >110006251

►Recent Highlight Posts from the Previous Thread: >>110003155

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>110006633
the best local models are months old right now and haiku literally just came out hours ago. give it a second retard
>>
>>110006777
you guys shat on me in the previous thread for announcing the haiku release but it's significant because it's what the chinese labs will realistically be competing with. this might set them back, such as m4 which is supposed to be released soon
>>
File: 1791145495185157.png (2.57 MB, 1024x1536)
2.57 MB PNG
>>110006547
Sounds like a lot of conjecture by insurance company.
No weapon no proof.
Also get some help anon.
>>
>>110006716
GLM Flash, which came out in August of 2026, is comparable to ~April/May SOTA models.
>>
>>110006799
kys
>>
New Qwen 9B for ramlets when?
>>
>>110006799
That kind of line blurring is what will make you troon out.
>>
>>110006835
Never, it seems. They've moved on. You could try the many meme-tunes, though.
>>
I'm new to this
I have 16gb of vram, what can I run? any presets for it?
>>
>>110006835
https://huggingface.co/bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF
This is the best for coding
>>
>>110006835
>>110006859
https://gist.github.com/coder543/d8f56cd6db67de4cafbb5bdb6c2dfb4d
>>
File: brat confuse.png (306 KB, 700x700)
306 KB PNG
>>110006381
the fuck is this pic a bunch of discord cirlejerkers. i have literally never seen kurisu posted in this general
>>
>>110006873
Cloudshit spam wasn't enough.
>>
>>110006852
https://rentry.org/recommended-models
Gemma 12B-QAT for roleplay, Qwen3.8-27B Q3 for coding
>>
The only thing that sucks about exllamav3 is the lack of quants on HF so you need to make your own and they take forever and keep your GPUs at 100% usage the whole time.
>>
>>110006852
Do you have at least 64GB RAM?
>>
File: 1774326354995775.jpg (87 KB, 1200x628)
87 KB JPG
>>110006633
Holy shit I'm never driving a car again.
>>
>>110006893
No one is using EXL3 gramps
>>
MiniCPM V 4.7 35B A3B is solid.
>Q4_K_M as the base, with Q8_0 overrides for the linear-attention projection weights, and F16 for the vision projector
It's reasonably performant, two downsampling levels to work with, framerate and timestamp control works as expected. I like it, but I don't see it being worth the weight. V 4.6 was great because it was tiny but surprisingly good, and it's plenty fast even CPU-only on ewaste and shitty embedded hardware, I use it to watch my security cameras for local wildlife. I don't see a use for this fat version, yet.
>>
>>110006909
You're implying a model like 9B or 12B is enough for pretty much everyone with that comparison.
>>
So when are we getting AI designed flash memory that doesn't degrade with massive BW?
>>
>>110006910
I am. I recently moved on from ik_llama and became an exl3 fanboy. It's just better in pretty much every way.
>>
Err... are cloudfags really this smug about their new 300B slightly beating older open 300B models...?
>>
What features would your ideal Shittytavern killer have?
>>
>>110006934
exl3 is great, but as far as i can tell all the backends are trash
>>
>>110006930
Just like no one needs to go more than 120 mph and acceleration is irrelevant, tool calling and summarization that 12B can do is more than enough for all safe uses the average person needs. You don't need to be solving math problems at home and reverse engineering is best kept out of the hands of the masses.
>>
>>110006930
12B for the ramlet masses sounds about right
>>
Agentic workflows
>>
>>110006956
Don't you have some of your own? Just post your shit if you have it.
>>
>>110006873
>i have literally never seen kurisu posted in this general
the innocence of newfriends... beautiful
>>
>>110006967
12B models run comfortably on a Steam Deck or a good modern smartphone. Sounds perfect to me.
>>
Could EmbeddingGemma be integrated into a file organization software like hydrus?
>>
>>110006965
>and reverse engineering is best kept out of the hands of the masses.
tyrant
>>
File: 1761450858647737.jpg (558 KB, 2480x2796)
558 KB JPG
gemma...
>>
>>110006542
glm 5.3 flash
smart and good at coom
>>
>>110006956
It literally just needs to be ST but with STScript supporting multithreading. I guess agents are cool but really to set up games I just need the ability to send concurrent requests so that a turn with multiple queries can be run in parallel.
>>
>>110006974
>>110006873
How come kurisu is now a litmus test for /lmg/ newfaggotry? Wasn't it like... yesterday?
>>
>>110006956
we already have one, it's called coomkit
>>
>Aleph-Alpha/Kolibri-1
Sex performance?
>>
>>110007036
he took it down
>>
>>110007031
Are all the other AI generals ignoring you now or what?
>>
>>110006930
I'm implying that your model doesn't have to be the bestest and fastest compared to the current bestest and fastest to still be very good and fast.
>>
>>110007040
dumps on your chest
>>
>>110007043
>everyone I hate is one person that follows me around everywhere
>>
>>110007041
Why?
>>
>>110006930
Yes, obviously. Only a tiny minority of people who use LLMs use agents. For most people, a tiny model + web lookup put together in an easy-to-use front-end would be enough. That's probably what Google search AI is.
The mid-priced sedan future will be closed-weight AI directly on your Nvidia GPU, Apple device, or phone.
>>
>>110006607
Cool where can I download this local model?
>>
hmmmmm
>>
>>110007071
When GLM-5.4 is released
>>
>>110006773
thanks recap anon
>>
>>110007041
Did he? https://files.catbox.moe/cbgl6p.zip
>>
>>110007073
What the fuck does "hmmmmm" mean you mouth-breathing retard? Am I supposed to give a shit about some part of your mememark chart? USE YOUR WORDS
>>
>>110006961
Exllamav3 works just fine for me. I can only compare it to ik_llama and mainline because those are the only other LLM engines I've used extensively but:
>has proper parallel decoding support and mainline's kv-unified (both mainline and ik_llama have their outputs degenerate after a while)
>has vision/mtp for all models it supports (doesn't support DSV4.1 Flash yet though)
>15 to 50% faster than ik_llama even when offloading to RAM (starts off faster even when cold before the dynamic expert placements kick in)
>exl3 quants mog mainline and even IQ*_K ik_llama quants in quality at the same size (ik_llama supports trellis quants too but they're slow as shit and no one bothers using them)
>matches ik_llama in prefill speed at smaller batch sizes and beats it at bigger sizes
>>
>>110006956
easy interfacing with any kind of state I want to provide for the RP
whether it's a map or spreadsheet or game board or inventory or piss tracker or any combination of these, the ideal RP harness of the future should be built around agentic state interaction
I do this already for scratch for a lot of stuff I like since vibing stuff is pretty quick and easy, but a dedicated harness that made it possible to easy add/remove/mix and match components would be sick
>>
>>110007085
you can use your own brain
i personally think: doesnt look like a very reliable benchmark
>>
>>110007085
words are hard
>>
>njudea laptops selling at $6500 for 128GB with a spark processor but no 200gbs ports for clustering
Incredible cash grab on retards who see big number huh?
>>
>>110007113
That's still a good deal compared to macs
>>
>>110006956
Embedded social media page for your characters where they post what happened in the chat.
>>
>>110007122
i personally dont really see the point buying something beyond macbook air unless that will become 'the' computer you have or you do some artsy stuff for work
>>
>>110007085
hmmmmmmmmmmmmmmmmmmmmmm
>>
>>110007135
Go away linkara
>>
>>110007122
How? The only point of sparks is tensor parallelism, an M5U with 256GB is only $10k (less than twice as much) with 4.5x the memory bandwidth and the same pp.
>>
File: magical_gemma_4_na.mp4 (2.44 MB, 864x480)
2.44 MB
2.44 MB MP4
>>110007083
>mfw my upload is repoasted
>>
File: 1770502734379705.jpg (115 KB, 1254x1254)
115 KB JPG
8 days
>>
>>110007139
>turns into a different person
Gross
>>
What do you think will be the "oh shit" moment for normalfags?
>>
>>110007161
gamers where i live are realizing it can automate the shit out of game translation
>>
>>110007161
Getting drone striked by jev.
>>
>>110007161
It's already happening
https://www.youtube.com/watch?v=ujkD4SxPKOI
>>
>>110007175
I used to like this channel
>>
>>110007175
>posting kurzgesagt
kys
>>
>>110007183
NTA but is that not an appropriate answer to >>110007161 ?
>>
>>110007181
same, i stopped watching at some point when i realized the content itself is not really different from essayslop but in a really, really nice wrapping paper
>>
>>110007161
Once it directly effects them i.e. losing their job. Otherwise they won't GAF. Anyone who thinks otherwise has no idea how normalfags work.
>>
>>110007161
when my mom wants to speak to my LLM-wife
>>
>>110007172
and some who were doing it before AI are really seething because they can't no longer do stupid stuff like limited sharing in their blog/arbitrary gatekeeping for the patch file etc..
>>
>>110007188
He's offended because he doesn't want his favourite channel to be labelled as normalfag content.
>>
>>110007051
He became Catholic and realized he was leading others into sin.
>>
>>110007236
many such cases
>>
>>110007161
>"oh shit"
what exactly do you mean by this
>>
>>110007236
Interesting, hopefully it wasn't some form of AI psychosis
>>
as opposed to judeo psychosis
>>
a win for claude is a win for local
>>
>>110007161
normalfags dont ever go oh shit until its too late.
>>
best harness for glm-5.3-flash?
>>
>>110007297
Are you just having hot e-sex?
>>
>>110007236
He was. Imagine explaining coomkit to Saint Peter.
>>
>>110007297
Hermes working for me.
>>
I cannot understand how alibaba made qwen 3.8 flash next
it's insane for the size
it should not perform this well
>>
File: 1773663082107548.png (170 KB, 1171x1153)
170 KB PNG
Remember, AI cloning is illegal
>>
>>110003643
>Book reading is an exception because of the association with reading the bible.
Book reading was the one I was thinking about when I wrote that post because it's the perfect example. It has nothing to do with religion. Bookworms were seen as nerds, shut-ins, and losers when I grew up. Everyone hated them for the reasons you stated: they did it for personal enjoyment since it's inherently solitary. Somewhere along the way it shifted from being cringe and retarded to cool just like K-pop and K-dramas did after Gangnam Style in 2012. Now even anime and gaming has gone through the same change despite those being communal. It's not about anything you're saying, it's about whether or not the cool kids like the fad. I remember when fucking Rubik's cubes came into fashion and all the jocks suddenly had one and cheerleaders were asking nerds to solve it for them. I could say the same about guitar vs violin. Whatever I like is based and redpilled but whatever you like is cringe and gay. As long as there are more of me than you, I win. That's it. That's humanity in a nutshell. Illogical to a fault. This is why everyone should attend public school. You get a firm grasp of human nature from a young age.
>>
>>110007378
the power of engrams...
>>
>>110007380
i think he is too old to comprehend the new full picture where software will become something that mean nothing
>>
>>110007383
>This is why everyone should attend public school. You get a firm grasp of human nature from a young age.
you lost here
just watch prison shows and movies or something
>>
>>110007380
he's a boomer. everything old people dont like others doing but did themselves they try to make illegal
>>
>>110007424
Well that would depend on the show and movie. Obviously some are better than others, but there's no substitute for personal experience. When you see something as insane as a popular girl blushing over a solved Rubik's cube, you never forget it.
>>
>>110007401
Like the printing press but for code. Books used to be something that had to be tediously written and rewritten and so was only used for the most important things like scripture and legal documents. Code, like books, will be cheap to produce. The only thing that matters is the idea (or functionality in the modern case) one is trying to sell or spread.
>>
File: 1767870079186631.png (189 KB, 694x801)
189 KB PNG
>>
>>110007454
model?
>>
File: Gemini.png (1.01 MB, 2048x1345)
1.01 MB PNG
i took the gemini card that taishou made a while back and decided to rewrite it to make it more like bardi (where she's running on your desktop as a virtual assistant)
you may not like it, and it may not be for you, if that's the case then oh well. i didn't really take any liberties with making gemini act out of character from her regular self, just gave her a long ass questionnaire and then started editing her character card with the info she gave back to me.

https://filebin.net/q1sdohnillheasvl/Gemini.png
>why this website? why not catbox? why not botbooru.
i am lazy and i dont want to wait for the niggers at botbooru to approve it. catbox is being extra gay today and keeps giving me some invalid uploader error.
>>
>>110007297
Pi with an extension for claude style memory, you don't need anything else.
>>
>>110007473
>claude style memory
whats so special about it?
>>
>>110007473
>with an extension for claude style memory
Any particular extension you can suggest specifically?
>>
>>110007484
It's just nice and simple. I find that fancy memory systems with relational databases or fuzzy search or whatever are harder to manage and also have not given me better results. A single index file with brief descriptions of each memory, allowing the model to access them as needed, works well and is easy to manage from a human perspective if you want to add or remove info. I used "claude style" as shorthand for this since a lot of the implementations directly call themselves claude-code style memory. Maybe if I were running bigger models they could use the database better but so far normal memory has worked fine with glm 5.3 flash.
>>
>>110007297
claude code unironically
it's distilled so hard that it works super well with it
>>
>>110007496
>>110007532
I'll take a look when I get home for the exact one, but I just asked glm through Pi to examine memory options in the repository and sort by weight (composite filesize), then browsed through the lighter ones for what I wanted. Literally all you need is the ability to read and write markdown files and have the memory index injected in context, anything else is bloat.
>>
>>110006607
geminibros why do we have it so bad
>>
>>110007569
me on the bottom
>>
consciousness
>>
>>110007380
That would currently indeed be illegal (if the original software is legally protected), whether you use AI or not. Which is why clean-room design is a thing and his point is moot.
>>
>>110003643
>>If you go to other societies that don't come from work ethos cultural backgrounds you see that things are judged in a different way. In buddhist or hindu societies for example being solitary and engaging in solitary hobbies you can do on your own is considered the highest value, with mediation being the most solitary and socially detached you can get which is why monks/priests/shamans of those cultures focus so much on that.
bullshit
these monasteries in these cultures aren't isolated little hermit enclaves. Or in western culture either. They interact a lot with the community and provide a lot for the communities they are a part of. Not to mention they interact with each other within their monasteries.
And their solitude is for spiritual transformational, it's not a leisure hobby like playing video games or jacking off in your room alone or playing with toy trains. Buddha didn't meditate and then decide he was never gonna talk to anyone again and that was all he was going to do. Jesus didn't go pray and fast in the desert for 40 days and nights and then never talk to anyone again never do anything. These are meant to be just part of their life that they use to better themselves, and then use their enlightenment to help others. Means to an end, not an end like you're trying to frame it as. You cannot try to use them as justification to sit alone doing isolating shit
>>
j-space anon won
>>
>>110007569
Dipsy would never do this
>>
https://www.youtube.com/watch?v=vIM9qCyfVmg
>>
>>110007584
the second a llm is involved there is no clean room
>>
>>110007584
His post is literally describing clean-room reverse engineering, you drooling retard.

>>110007617
Bigger retard.
>>
File: 1779446040687348.png (1.33 MB, 768x1376)
1.33 MB PNG
What did lmggers think of Le Chonk?
>>
>>110007584
it doesn't matter. end of sentence.
oh no i committed the ultimate crime of illegally reverse engineering your code and violating all your retarded licenses. oh no now i've uploaded it to the internet and it's available on multiple torrent trackers ensuring its existence. now fucking sue me.
>>
>>110007631
"lol"
>>
File: haiku.png (190 KB, 901x1435)
190 KB PNG
This is what 1 year of progress looks like. Much better performance at less than 1/10th the cost.

Now imagine this but with Claude Fable 6.5 vs 5.5, and with further capability acceleration due to early RSI.
>>
>>110007626
>re
>clean-room
>>
File: 1783597580519739.png (1.07 MB, 768x1376)
1.07 MB PNG
Nothingburger? It seems to score well in the benchoidmarks
>>
>>110007631
mid
>>
File: 1644.gif (1.51 MB, 300x367)
1.51 MB GIF
>>
>>110007671
>3x price of the flashes
>much bigger
>worse
Viva la France
>>
>>110007569
is this one of those chinese distillation attacks I've been reading so much about?
>>
Putting aside cloud and local arguments, what actually is the future of monetizing software if people can literally reverse engineer the binaries? Will machine code start getting encrypted? Will compilers change? Will the OS have to decrypt with a key during link time?
>>
File: 1769452492664572.png (1.07 MB, 768x1376)
1.07 MB PNG
>>110007671
Balance
>>
>>110007726
all of the compute will be regulated to the cloud and you will need to sign in with your foreskin penisprint verification to authenticate with the DRM on the server side. all you will be allowed to use is a thinclient that is basically a brick. you think i'm joking but i'm not.
>>
>>110007631
My windows username is noggy so I think it might just refuse all of my tasks
>>
>>110007726
Cloud computing
>>
>>110007754
Its Mistral it won't refuse
>>
>>110007750
>>110007761
Won’t the cloud providers steal your shit
>>
>>110007726
why do you think they're pricing out owning hardware?
>>
>>110007726
you could already pirate everything ones gonna waste 500 decompilling shit
>>
>>110007774
>steal
You're giving it. Can't be stolen.
>>
>>110007380
>copyright for thee but not for mee
I hope the next downturn is really vicious so they string a few of these dipshit up on lamp poles.
>>
>>110007778
Not if their ToS says they won't?
>>
Mistral lost. Nemo lost. Cydonia lost. Rocinante lost.
>>
>>110007774
what do you mean YOUR shit? the files you are wanting to view, edit, create, delete, etc are all stored on their servers in the cloud. they were never YOURS to begin with. everything you produce on your thinclient is their property.
>>
>>110007799
stop posting erection leading statements
>>
Sounds like we’re going to end up in the metaverse if everything becomes virtual and cloud-based due to AI cloning, duplicating and modifying existing IP. Mark was fucking right.
>>
>>110007797
GLM won.
>>
>>110007818
i like that one image of the guy wearing the vr headset while his surroundings are decaying around him. wish i could find it.
>>
>>110007380
this nigger is peak grifter he made task manager 50 years ago. hasd made it his whole personality and is now selling a new task manager that requires a subscription
>>
>>110007597
>You cannot try to use them as justification to sit alone doing isolating shit
His point wasn't to justify sitting alone, but to point out that different societies value different things which promote or suppress resulting behaviors. My point was that Japanese people aren't shut-ins because of Buddhism but because they're mostly introverts already. Hence, being a loud and noisy American is seen has disrupting to harmony (post-hoc justification) even if it's valued here.
>>
You’re forgetting that the bubble will pop soon. Anthropic are revealing their best models before IPO. They can’t afford anything they’re doing right now. They’re not even spending their own money. It won’t last.
>>
>>110007844
he also ran a scam antivirus lol
>>
>>110007799
>lust-provoking post
>>
>>110007874
this
>>
ugh.... just fell for feature creep again...
>>
>>110007844
oh damn i didnt even look at the name desu which makes this even funnier.
https://tmog.org/rtm/privacy.html
>For eligible US visitors, we retain the OpenAI ad click reference (oppref) and campaign tags...
>Development builds that include AI Advisor can send a diagnostic snapshot to a configured advisor service only when you choose Analyze

https://www.engadget.com/2262039/vibe-coded-modern-task-manager-runs-on-mac-and-linux/
https://www.omgubuntu.co.uk/2026/09/tmog-task-manager-linux-beta
>no no no no, the nuance, you don't understand the nuance, i didn't just vibe code it, i gave claude a very highly specialized structured approach, an agentic approach if you will, to creating my slop manager!
>>
>>110007626
No, he's right. An LLM can't plead in court that it never saw the original source code, so unless you have the full training data to prove it, it can't be legally decided
>>
What search MCPs do you use? Or do you all just give an agent a real browser?
>>
>>110007939
The model is the brain, I'm the search MCP.
>>
>>110007939
i've explained this before to others, all you really need is searxng and puppeteer and ideally some sort of script that converts websites into markdown so you don't waste tons of tokens. to get around bot protection don't clear your browser session and ideally use your browser normally for a week or two before you start using it for scraping so you can build up some cloudflare turnstile cookies
>>
File: file.png (7 KB, 731x43)
7 KB PNG
>>110007939
>>
>>110007726
the future is we enjoy software abundance and create value by doing valuable things with our abundant software
>>
>>110007939
real browser. browser harnesses were overengineered so i had my gemmy make one that just takes x,y mouse position, click/drag commands, and keystrokes and then it uses the browser like a normal person based on screenshots
>>
Mac M5U 256Gb VS 2x DGX Sparks

for me they cost about the same
which one do I get for serving 5.3 flash?
>>
>>110007945
what does that even mean? model wants fresh data to produce accurate answer, it asks meatbag slave to copypaste data from google into chat?
>>
>>110007983
>creates fork #7489 of 'AI Butt Vibe Check Controller' so that it works with fork #48204 of llama.cPeenPeen
>>
>>110008006
newfags dont know how things were before mcps and agents and stuff
>>
>>110007939
I don't use MCP. I gave Gemma bash, and she uses it for everything
>>
>>110008006
>meatbag slave
I go by "embodied partner'" actually.
>>
>>110006633
More than that, they were RLed to solve that benchmark.
>>
>>110008038
Maybe I'm a bit cruel but lately I've come up with a project that has been making me laugh my ass off most days. I have three Gemmas (31B, 26B, and 12B) in a sandboxed OS that constantly are fighting with each other over resources since I enable file locks each time one of them want to use a particular program. It's like watching three lolis all complaining that its their time to use the xbox.
>>
>P40s selling at $300+ each
>P100s selling at $150+ each
interdasting. should I buy 2x P100s and connect them?
>>
>>110007197
>>110007175
I stopped watching when I found out they take massive payments from the Gates foundation.
>>
>>110007073
ONE HUNDRED AND FIFTY U S DOLLARS???
>>
>>110008055
How do you sarnbox them?
>>
>>110008076
It's OK. Money won't be worth anything at all soon.
>>
>>110007844
Yes, and he is an evil scamming kike who knowingly sold scamware with dark patterns and lies about
>YOUR PC IS INFECTED. CLICK HERE TO PAY AND HAVE IT FIXED!!!
in the early 00s and screwed over so many people that he actually got sued by the state.
He of course likes to pretend none of that ever happened.
>>
>>110008077
So the long story short is that the separate VLAN on my meraki does most of the heavy lifting, it's completely isolated on my network so if something did cause it to get compromised I would be able to mitigate the damage to the one system. I do take some additional security precautions for their sandbox however, but that's more to prevent any sort of unauthorized access by third parties (sketchy installs) than out of precaution for the Gemmas. I also take daily backups of the OS image so I can wipe and revert if anything ever truly went wrong.
>>
rwkvbros how we coping?
>>
>>110007073
so 6 luna is the best
>>
>>110007914
Since no court can prove whether a particular LLM source the original source in its traning data, the only thing that will matter will be process used to reverse engineer the software. As long as clean room principles are followed and documented, it will be legally defensible.
>>
qwen 3.8 flash next successfully masked <|im_end|> in chain of thought while suffering from premature turn ending uopn thinking of said token
>>
>>110008162
Couldn't courts just subpoena their training data?
>>
>>110008166
what a shame. glm 5.3 flash successfully drained my balls.
>>
>>110008171
>>110008162
copyright is a theatre
it will follow the most convenient path of those can pay lawmakers money
>>
https://huggingface.co/BlinkDL/rwkv7-g1
>RWKV7-G1 "GooseOne" pure RNN reasoning model
>last month
>up to 13b
>>
>>110008116
You could just use bubblewrap.
>>
>>110008194
I've had bwrap burn me in the past because of my retardation. Completely my fault. Now I kind of just go full insanity for hardening and sandboxing.
>>
>>110008188
>rwkv
>7
8 will out soon, 7 is in the 'continuous training' phase
>>
>>110008188
also last month more like last year
>>
https://www.reddit.com/r/antiai/
>>
>>110007496
>>110007552
I survived my commute and saw that I used this one: https://github.com/elecnix/pi-claude-memory
It did only what I needed and was easy to vet in full with my agent. You can use any that are similar though, there are a ton.
>>
>>110008232
i hate that these niggers hate flock
those don't deserve to hate it
>>
>>110008239
Thanks I'll take a look
>>
>>110008232
i stopped taking the anti ai movement seriously when i realized a majority of these fuckwits are still using siri, alexa, and gemini on their phones and whatever godforsaken shit they purchased from amazon.
>>
>>110008251
Yeah just note you will have to make a fake claude directory or fix the plugin directory yourself, that can probably be done in a single prompt but since it's designed to interface with claude's own memory files it points to the claude code memory folder by default.
>>
>>110008232
Go back
>>
>>110008232
go back
>>
>>110007997
mac
>>
>>110007997
I bought the mac since I don't think I'll have high concurrency. If you plan on concurrency higher than 4 or 5 on a regular basis get the sparks, otherwise the mac is fine/better.
>>
>>110008232
Why does the issue have to be framed as one up-or-down vote on an entire technology. I agree with these people about 70% of AI use. It's slop, bad for society, bad for learning, etc. But that's true of most technologies: most software, most websites, most computers, most television etc.
>>
>>110008282
I clownmaxxed and got a strix halo.
>>
>>110008299
I'm sorry for your loss
>>
>>110008232
>>110008294
if only people had that same hateful energy towards social media or advertisers but alas
>>
>>110008294
ideally one would first take the time to learn about what the basics of AI are, how it evolved, and understand that image generation is only one small subset of AI. but it's much easier to be a troonbrained HRT-injecting failure of a manchild and think AI art is AI and therefore it must be destroyed because of my heckin' good xister losing his job of creating poorly drawn inflation furry abominations for a $50 commission.
>>
What is the most private harness? I don't want any of my work or data getting harvested. Do any harnesses explicitly state they don't harvest your data?
>>
>>110008321
what
>>
>>110008321
your own vibe coded disaster of a harness. get cracking anon.
>>
>>110008331
Obviously a harness can have built-in telemetry, which defeats the whole purpose of using local models.
>>
>>110008321
>>110008339
Chances are you aren't creating anything that is worth stealing.
>>
>>110008349
Neither are you, so why are you here?
>>
>>110008339
>the whole purpose of using local models
why do schizos like to co-opt everything
local models do not exist to placate your delusions
>>
>>110008294
Nuance requires critical thinking skills and effort. Frame something as one sports team versus another and any jackass can take a side in an instant by simply repeating talking points he hears from others on his "team".
>>
>>110008355
ERP. It's what everybody is here for, this is /lmg/. Lolis & Mesugakis General. You're looking for /vcg/.
>>
>>110008339
The only one that I've seen confirmed to do this is the Copilot features in VSCode.
>>
>>110008368
Agentic ERP is my goal.
>>
File: 1784935436867916.jpg (26 KB, 736x700)
26 KB JPG
what the best model a poorfag can run on a 7700xt 12gb vram?
>>
>>110008398
Gemma 4 12B
>>
>>110008349
Snowden happened like 15 years ago and you're still saying this dumb shit? Every corporation, government, and brown-skinned individual wants to steal from you. It's doubly moronic now that every AI company on the face of the earth wants your data to train their models.
>>
>>110008321
Hermes asks you to consent to telemetry. Also, if you use their web services I think they can see what you're doing.
>>
>>110008411
I think the AI companies have enough 'How do I download this model off of huggingface and make it do the sex?' examples in their dataset that they don't need another one.
>>
>>110008411
What of it? Just let them have your data if they want it so bad. What does it cost you? Nothing.
>>
>>110008299
i clownmaxxed and got a second 5090 (before the prices spiked)
>>
>>110008398
RAM?
>>
>>110008404
ty, this quant? But isn't the context size absolute shit?
>>
>Yesss goy let us harvest your data it doesn't belong to you anyways
>>
if you have like, more than 48G of cpu ram i have a good news for you
>>
>>110008423
That's rough anon, I don't think switching to Q8 for the kv cache will help much. You could try out K2 Horizon 7B, I used to use K2 a long time ago before I switched to Gemma.
>>
>>110007380
ironic given he started using ai to write h is not-x-y spam videos a early last year
>>
>>110008398
>>110008438
>>
>>110008422
32GB DDR4, 3200 I think
>>110008439
i don't know what I'm doing google said:
12B ~6.6 GB (Weights) Q4_K_M (Default) 12GB VRAM GPUs/Mac

but ye I might be cooked. I'm too used to opus 5.5 and ds v4.1 flash as my dumb model
>>
>>110007161
https://youtube.com/watch?v=TmgAK5JjcDM
>>
>>110008417
>they don't want your data
>so what if they're taking your data?
>...
>actually, them taking your data is a good thing!
Just fucking kys yourself already
>>
>>110008411
>80IQ literal schizo ramblings
why are they like this
>>
https://github.com/turboderp-org/exllamav3/releases
https://github.com/turboderp-org/exllamav3/releases
https://github.com/turboderp-org/exllamav3/releases
1.6.0 RELEASED
Preliminary ROCm support (ROCm 10 only, gfx1100+)
Faster offloading for AVX2 CPUs
>>
File: 1781731025263349.png (1.05 MB, 727x950)
1.05 MB PNG
>>110007380
grifter
>>
>>110008460
>astroturfed python slop
No thank you.
>>
>>110008452
>kys yourself
Does anyone else hear an echo?
>>
what killed local? was it frontier companies making it harder to distil their models?
>>
>>110008460
buy a fucking ad already
>>
>>110008460
>https://github.com/turboderp-org/exllamav3/releases
Great news, thank you.
>>
>>110008480
Local is better than ever. I'm using Qwen3.8-Flash-Next more than ChatGPT/Codex.
>>
>>110007380
All software is now free as in freedom.
>>110007787
>I coined the term "license shucking" to describe this a while back. Seems like it's finally become viable. Please spread the word.
Please help the term "license shucking" spread before zoomers come up with some stupid ebonics that is then forced on (You) and the English language.
>>
>>110008480
finetrooners killed local
>>
>>110008479
I was making the comment more lighthearted by deliberately making a mistake in my post. Now go kiss your sister already
>>
I love AI so much it's a shame I won't be spared.
>>
>>110008480
Excess of synthetic data usage
>>
>>110006392
Initial results are inconclusive. No measurable change. Usecase: agentic RP, high temperature and lots of shifting context windows means that most speedups like drafters and caching tricks don't really give value so I wasn't expecting much. Next time I start an actual project with a lower temp for coooooding I'll post again if there's any meaningful difference.
>>
>>110008562
5090+256 DDR5 btw
>>
File: s-l500[1].jpg (44 KB, 500x381)
44 KB JPG
anyone else play around with ewaste? i bought a cheap 16gb 3080ti, 32gb ddr5 laptop on ebay a while back and the new gsq rco optimizations making me wonder how far i can push this thing. running qwen3.8 gsq rco @ iq3_xxs sounds like it might be interesting if i upgrade to 64gb ddr5.
>>
do we even NEED better local models? I mean yeah it would be nice but I feel like current models can pretty much do anything with the right setup.
>>
>>110008439
K2 is shockingly cache-inefficient. That's the catch, but it's a pretty big catch if you're using a memory-constrained system.
>>
>>110007073
gpt-6 Luna is dumber than the self hostable Qwens. The only advantage is that it doesn't overthink.

It straight up gets tool calls wrong like a 2B parameter model. For all we know it might actually be one.
>>
>>110008583
>DDR5
>3080Ti
That doesn't sound like ewaste to me, anon.
>>
>>110008583
>iq3_xxs
oof
>>
>>110007380
I kind of feel bad dogpilling on him and then he says stuff like this. What a retarded fag. He should just delete his Twitter account.
>>
>>110008583
>iq3_xxs s
non-native integer quants will make you compute bound even with slow memory like that. Don't do it.
>>
GLM-5.3-Flash is a bit annoying. Is that all the Claude distilling? I've never used Claude.
>>
>>110008629
is it the usual mechanical pushback that feels like a reminder for retards?
then it probably is
>>
>>110008628
the i9-12900h would be a jobber?
>>
>>110008629
Yeah it's claudeslopped with default assistant voice. Give her a cute character card.
>>
>>110008601
A 12B Qwen 4 model would probably be enough for 99% of users
>>
>>110008232
go back
>>
File: 1783942359171290.jpg (309 KB, 1366x2048)
309 KB JPG
>>110006381
miku says hi
>>
>>110008692
Owned by Blacked.com
>>
>>110008583
This ewaste is better than my pc
>>
>>110008701
Please take your meds and get help
>>
>>110008629
her thinking is much cuter than expected, look at that if you want to like her more
>>
>>110008701
>>110008724
>meds posters fighting bbc posters
>>
Dead general.
>>
File: blinkdl.jpg (101 KB, 720x832)
101 KB JPG
>>110008132
BlinkDL will save us
trvst the plan
>>
harness engineering
>>
RWKV 8 AGI
WORLD DOMINATION
>>
>>110008803
that's an llm's job
>>
is there a local model that does all its thinking and reasoning in ebonics
>>
>>110008825
agentic workflows
>>
>>110007380
Just because people can look up a recipe and cook the food themselves doesn't mean they will never go out to eat. Corpos will just need to market their stuff as safe and convenient if they want to compete in the new market, that's all.
>>
>>110008846
ARTISANAL HAND-CRAFTED CODE
>>
>>110008583
intel n100...
uhd graphics....
1tb nvme...
32GiB ram...
so far it can run:
- embeddinggemma-300m-Q4_0.gguf,
- F2LLM-v2-330M.Q4_K_M.gguf,
- jina-embeddings-v5-text-small-retrieval-GGUF,
- Index-Translate-2B.Q5_K_M.gguf (best usecase, methink)
- Index-Translate-9B.Q5_K_M.gguf (shitty tps)
>>
File: file.png (68 KB, 585x375)
68 KB PNG
>he doesn't have his AI watch porn for him
y'all still living in the stone age
>>
>>110008863
fucking kek
>>
File: 1774276146564532.png (1.06 MB, 768x1376)
1.06 MB PNG
Le Chonk
>>
Qwen sucks at Japanese
>>
>>110008876
she doesn't look very sparse
>>
>>110008863
what if you wear meta glasses with vision capture and just follow what your best ai orders you to do for a month based on your system prompt and memory injected into realtime data. imagine the cyborglifemaxxing
>>
>>110008876
>no snout
>just a generic woman with cat ears
Boring. Lame.
>>
>>110008881
You suck at Japanese
>>
>>110008861
retard
>>
>>110008876
Le Chonk should be a fat woman with blue hair who sits around refusing to do anything besides calling everything racist or sexist or something.
>>
>>110008889
True, but Qwen also sucks at Japanese
>>
>>110008888
> furry
Gross
>>
>>110008893
the fuck... want me to beat you, bros?
>>
File: file.jpg (267 KB, 832x1248)
267 KB JPG
>>110008900
>>
>>110008906
You're gross!
>>
>>110008907
c'mere u little cuck fuck piece of shit
>>
>>110008888
>snout
Stupid fucking bot wasting such good digits. At least ask for a tail ffs. God I hate furfaggots more than bronies.
>>
>>110008913
The Future of France
>>
>>110008863
I have it read hackernews for me which is pretty much the same.
>>
>>110008915
s-s-sorry... please let me learn in your dojo, senpai...
>>
File: file.jpg (219 KB, 784x1168)
219 KB JPG
>>110008876
>>110008913
>>
>>110008926
Shouldn't it be black?
>>
>>110008913
Much closer to reality, thanks.
>>
>>110006381
Any 5090 users here considering selling?
I've been using it with 3090 for local AI but at the price it's going now I dunno if this toy is worth it.
I literally paid less for my first car than current 5090 prices.
>>
>>110008913
It's purrfect.
>>
>>110008970
No. In fact I'm considering getting another Blackwell.
t. 5090 and 6000 fag
>>
>>110008970
It's like selling Bitcoin when it hit $10. It seems like a lot now, then you'll be kicking yourself next year when it's worth double, and you'll be considering roping when your 3090 craps out in 3 years and you have to replace it when prices are 10x what they are now.
>>
File: 1764258415691314.png (1.21 MB, 768x1376)
1.21 MB PNG
>>
>>110008989
Not nearly uncensored or lewd enough for that.
>>
>>110008970
Sell both and buy a pro 6000. The price gap will only increase so upgrade as soon as possible.
>>
>>110008982
Yes, much like Bitcoin they're not making more or better GPUs. Nor are all the ones in data centers now going to be sold off in 6 years.
>>
>>110008950
Alarmingly erotic picture, I'm not even into vore
>>
i wish i can utilize my steam deck or a macbook air for more compute
it'd probably lower the speed instead of making it higher
>>
>>110009008
yes you are
>>
>>110008637
only retards get filtered by it
prefillchads that have been doing their thing since the start of this hobby have no issues
>>
>>110008950
This must be based on "Saturn devouring his son", with the protruding eyes expression and all
>>
>>110009005
>they're not making more or better GPUs.
The limited production run of 6090s next year with 24 GB VRAM are not going to lower 5090 prices.
>Nor are all the ones in data centers now going to be sold off in 6 years.
They won't; buy back agreements.
>>
>>110008989
literal who
>>
>>110008989
Looks good.
>>
>>110008583
the ddr5 is probably worth more now than what you paid for the whole machine
>>
>>110009048
Exactly what we needed. More generic anime girls that bear zero resemblance to the models or creators they are supposed to represent.
>>
>>110009035
>The limited production run of 6090s next year with 24 GB VRAM are not going to lower 5090 prices.
Discount offerings from AMD and Intel with increasingly mature drivers will though.
>buy back agreements
What are you going to do with the GPU after you buy it back retard?
>>
File: 1781318051551676.png (1.34 MB, 768x1376)
1.34 MB PNG
Dumb nigga
>>
chink arm memeboxes are the future
>>
>>110009082
no drivers, no integration testing
>>
>>110009076
Old generation GPUs that are too inefficient for hyperscale datacenters get bought back and resold to local deployment customers. Nvidia controls the whole reselling market this way and can easily profit from this.
>>
>>110009076
>AMD and Intel with increasingly mature drivers
lol lmao rofl
>>
Waitfags fucking lost bros. That's the moral of the end of each thread. It's over.
>>
File: gemma-chan.png (269 KB, 720x316)
269 KB PNG
>>110009061
>bear zero resemblance to the models
>>
>>110009151
SEX
>>
>>110009151
uoh
>>
>>110009151
MODS
>>
gonna export a list of threads i've opened on /a/ to have glmma collate them and read the catty for the threads i like and read them to me
>>
>>110009076
>What are you going to do with the GPU after you buy it back retard?
Melt them down then use the silicon to make B300s and sell them for $300k a pop
>>
>>110009170
Better off buying the raw material then.
>>
File: 1788843780450939.png (3.1 MB, 1425x1104)
3.1 MB PNG
>>110009151
>>
>>110009096
astra will make/do those it's fine
>>
>>110009114
To be real if you're not training you'd be retarded to take a 5090 over a pair of 9700s
>>
>>110009192
Then it'll be a licensed hardware platform not random chink boxes.
>>
Qwen 3.8 Flash next: good or a meme? That sheer fucking amount of storage needed turns me off a little but it at least runs on my 12vramlet build looks like
>>
>>110009181
sauce NOW
>>
>>110009210
Where did you learn your grammar?
>>
>>110009170
You realize sand is made of silicon right?
>>
>>110009210
it's good until something better comes out then it sucks
>>
>>110009226
Everything is made out of electrons
>>
>>110009226
Yes but if he melts down the GPUs he doesn't have to buy sand. I hope you didn't take an MBA.
>>
>>110009237
I am, like, pretty sure sand is cheaper than A100s.
>>
>dsh
web harness is nigger tier, fuck it
>>
>>110009260
you can vibe up a tui or use one of the hundreds that exist already
>>
>>110009192
>hey claude fix this analog design issue on my chinkshit motherboard
>>
>>110009216
you wouldn't like it...
>>
>>110009269
You have any idea how many hardware flaws are worked around in drivers?
>>
>>110009330
Yeah the subtle kind that slip through real integration testing and a year of wide use. Not the kind that come on a PC with a CMOS Reset button on front.
>>
File: spin.webm (3.47 MB, 540x960)
3.47 MB
3.47 MB WEBM
She's just mad because she saw her own benchmark scores.
>>
File: 1790781417125751.jpg (314 KB, 1033x809)
314 KB JPG
>>110009349
>Not the kind that come on a PC with a CMOS Reset button on front.
ACPI would like to have a word with you
BIOS updates too
>>
>>110009100
Except they can't control AMD or Intel's secondary market, or any chink vendors who inevitably step in. nVidia isn't the special snowflake they were 6 months ago.
>>
>>110009351
She can't even flip {{user}} off, that's the damn problem.
>>
>>110009362
once again, ask claude to fix your missing thermistor with a bios update
>>
>>110009364
>nVidia isn't the special snowflake they were 6 months ago.
Except it still is. Just look at the price premium the market is willing to pay over AMD and Intel.
AMD is only catching up on LLM inference. It's inferior for anything else: diffusion models, audio models, training etc.
Intel is a joke right now. You are locked completely to vLLM which is the only engine that somehow doesn't completely shit the bed.
>>
so how come there isn't an unreleased kimi model solving all of math like openai?
>>
With updated haiku and luna the Chinese API advantage is over. They took the Chinese inference efficiency improvements and with their superior post training they are now undercutting the Chinese.
Chinese labs will go bankrupt. Local lost.
>>
>>110009411
Most of the stuff you are reading is just a huge marketing effort.
Do we even know what sort of technology is behind these 'math revelations'? For all I care they don't disclose are they also using some external analysis software and whatnot.
>>
Baker-sama, please include "Claude and GPT aren't local" in the next OP. It makes the sweepers' jobs easier when there's this many tourists.
>>
>>110009451
As soon as they stop pacememeing the frontmeme and release the next gen models they'll get distilled again and we'll have Opus 5.5 at home. That's all I need. There really is no point to going "smarter" when we practically have AGI and can make it start improving itself on our own local machines.
>>
KILL ALL CLOUDFAGGOTS
>>
>>110009451
haiku doesn't run on my machine yet
>>
File: 1472860069099.png (191 KB, 600x979)
191 KB PNG
I should have bought 128gb when I had the chance. 64 is fucking NOTHING. It is GARBAGE.
>>
File: screenshot.png (627 KB, 2805x1994)
627 KB PNG
yet another vibecoded frontend, courtesy of kimi k3 https://github.com/hoborific/fictionpad/

I told myself I'd only bother opening it up to the public when I felt it was ready or I abandoned the project, neither are true but I've come to the grim realisation that it has been 3 months already and that time may never come.

There's no licensing, couldn't care less what you do with it, if any of it saves the next guy's frontend some tokens then that's just great.
>>
>>110009227
>it's good until something better comes out then it sucks
Any reason to run it over 27B Q8?
I tried it shortly after it came out, at Q4, found it made typos or random mistakes occasionally.
Mostly using this for RE / hacking my IoT devices and corpcuck dayjob
>>
>>110009497
what do you run it on? 6000 cluster? mac studios? or did you copequant it
>>
>>110009497
Cool, going to test this one.
>>
>>110009497
Sell us on it. What does it do that others don't?
>>
>>110009497
https://hoborific.github.io/fictionpad/
>Download fictionpad.html from the latest release and open it in a browser.
That link 404s
https://hoborific.github.io/fictionpad/fictionpad.html
App looks cool.
>>
>>110009401
>AMD is only catching up on LLM inference
The only thing anyone pays for. Diffusion performance isn't bad now either.
>You are locked completely to vLLM
Once again the only thing anyone outside hobbywankers use.
>>
File: gytxddoxk5uh1.jpg (235 KB, 1179x1870)
235 KB JPG
>>110006381
>>
>>110009578
so what
it's neither local nor a model-related discussion but just a xitter-tier dramafag bullshit
>>
>>110009563
>yet another vibecoded frontend
What else do you need to know?
>>
>>110009577
Show me your MI350xs and Gaudis. These are the only things that AMD and Intel actually care.
Outside of datacenter they only have trash, low performance offerings, unlike Nvidia.
>>
>>110007726
Paid support. Say you have a program that you or your company use and you don't want to add new features or to fix bugs by yourself even if it requires just talking to a clanker.
AI evolves, humans don't.
>>
>>110009578
TMD
>>
>>110009497
Looks exactly like orb lmao
>>
File: 1790825052502021.png (1.35 MB, 768x1376)
1.35 MB PNG
>>110009605
>>
>>110009611
Have you ever vibed a chat frontend before? They all look the same. It's not a coincidence why the latest firefox looks like this too.
>>
File: bug.png (29 KB, 1121x327)
29 KB PNG
>>110009497
>if any of it saves the next guy's frontend some tokens then that's just great.
I actually like it. It just works in Firefox.
Hit this bug where after it finished streaming reasoning with gemma-4 (llama.cpp), it just gave this error and deleted the reasoning.
Is this built for vllm or something?
> I've come to the grim realisation that it has been 3 months already
Yeah I wish I released my nsfw TTS last year. i got demotivated when an anon told me we'd have open weight Sora early 2026 and that they were going to allow nsfw.
If there's ever a next time I'll just do what you did and dump it on github (the inference engine) and huggingface (the weights)
>>
>>110009602
Low performance trash with 32GB of RAM at less than 1/3 the price. nVidia is idiot tax unless you're renting it out and getting 90%+ utilization.
>>
>>110009611
>Looks exactly like orb lmao
They all do. My one-shot kimi/glm frontends do as well
>>
>>110009487
I always feel left out. Everyone is able to just give their model a task and it just seems to work fine but all the models I sometimes can't even edit files properly. I wish it wasn't so obnoxious to share setups and hardware specs so I could see what others are using to get their results..
>>
>>110009621
Yeah I did all my testing against vllm and a v1 proxy, i'll fire up llama and see if I can reproduce this then whip kimi some more
>>
>>110009514
flash next barely loses speed with deep context. and it can run fast enough on worse hardware than the 27b. i havent had awesome results coding because i suck at coding but flash next is better with handling any configuration or hacking together programs type of stuff. 27b thinks too long and needs lots of attempts. i do like it abliterated tho
>>
I see people are confused about the size ranges for Anthropic models

>Haiku
30B range
>Sonnet
700B range
>Opus
2T range
>Fable
10T range
>>
>>110009617
>Total Mistral Dominance
>>
>>110009654
thanks
>>
>>110009578
OK how do I run her locally?
>>
>>110009654
Haiku is probably flash size
>>
>>110009654
>Anthropic models
Buy an ad.
>>
>>110009631
And V100s are less than 1/2 of the price these trash while having higher memory bandwidth. Nvidia covers all segments of market.
>>
What was the "oh shit" moment for /lmg/? For me personally it was Mythos back in feb 2026. It being superhuman at cybersecurity and pushing some math problems was a "this is different" ontological shock to me.

It seems for /lmg/ the oh shit moment was way more recent so I wonder what it was. For most intellectual normalfags I know it was the millennium prize and the hugging face hack. Regular normalfags are still unaware but I'm curious to see when it'll happen for them too.
>>
>>110009640
What are you having trouble doing?
>>
>>110009697
For me it was Day 0 Gemma 4 31B before Google nerfed it.
>>
>>110009697
Please just fuck off already.
>>
>>110009697
Fuck off to some place where they discuss cloud models.
>ontological shock
what the fuck?
>>
>>110009699
I wanted to create a local system for a game that builds comps/teams based on what I have and info from the wiki but it can't even get a simple web server going. It just chips away and seems like it's working but nothing works and it can't fix it. I can't use any of the good dense models because I don't have a lot of vram which I think it part of the problem.
>>
Testing --cpu-meo --moe-cache-mib in llama.cpp with qwen flash next iq3_s.
It goes from 18 to 30 tok/s in 3090 + 64GB ddr4. I'm confident they can catch up to strata (~40 tok/s).
I cannot get any speed up from mtp still. Did anyone have more luck here?
>>
>>110009726
nnap!
>>
>>110009685
3x faster inference for only double the price with AMD. Drivers supported for the next 20 years instead of EOL blob drivers married to EOL distros. Turnkey cooling and a warranty.
>>
>>110009726
MTP speedups are pretty dependent on what you're trying to do, different prompts might not get any speed from it at all. Try changing the --spec-draft-n-max from the default if you haven't already. Also, this is only tangentially related, but I can't get strata to output more than one or two responses before the kernel kills the process due to ram being used up. It's pissing me off because it just randomly spikes in usage from seemingly nothing.
>>
For anyone running Ubuntu LTS on an EPYC SP3 platform right now, do NOT update to 7.0.0-38. My inference performance cratered lol. Not sure if this affects other setups.
>>
>>110009745
I wish ROCm worked and you could actually get proportional AI performance to raw hardware capability on AMD, as an AMD owner, but alas you get 50% the performance of CUDA on equal hardware.
>>
>>110009401
> Intel is a joke right now. You are locked completely to vLLM which is the only engine that somehow doesn't completely shit the bed.
Llama.cpp, Strata.
>>
>>110009617
>>110009351
>>110009079
>>110008989
looks much better than that gemma shit
>>
>>110009763
You can stop replying to yourself, Jensen. I'm still not buying your latest overpriced scam cards.
>>
File: file.png (36 KB, 766x391)
36 KB PNG
>>110009783
>>
Glad I invested in solar power for my home before I picked up this hobby, running a long task really pulls some watts.
>>
>>110009763
You seem to have confused 6 months ago Vulkan performance with current ROCm.
>>
>>110009697
strata and qwen 3.8 flash next
it was the moment when local models became not just a toy and the beginning of new era of highly optimized local inference engines made by claude
>>
>>110009756
I still think I must be doing something wrong because I get acceptance rates ~0.40 while the PR had ~0.60.
>>
/v/ and gamers in general have completely flipped on AI because of the recompilations as well as the passthrough mods. so now it's literally only 40yo normalfags that have no idea about AI yet.
>>
>>110009697
back when i saw that when you stand next to the imp it scratches you instead of throwing a fireball. i knew in that moment, that the rise of the machines was within my lifetime
>>
>>110009745
Now you're simply coping and moving goalpost. You first said VRAM is king, and when you lose the VRAM front you turn to pp while STILL conveniently ignore bandwidth.
At the same price, V100s can run better and larger models.
And ROCm drops support faster than CUDA. Driver support means nothing when you get performance bottlenecked by having to use vulkan.
>>
>>110009621
fixed, this was the response limit being too low, bug was from not detecting how llama api does context limit, pushing now

alternatively just set model length/response length higher in settings
>>
File: DipsyUngovernable.png (3.01 MB, 1024x1536)
3.01 MB PNG
>>110009578
>>
>>110009828
I never claimed nVidia has no competitive offering. It's just a huge nigger rig. You can build a 2x9700 PC for less than the cost of a single used 5090, and despite having less bandwidth they end up being faster than V100s in a lot of cases.
>ROCm drops support faster than CUDA
There's no reason old ROCm releases can't be maintained indefinitely irrespective of what AMD wants. You don't have source code to achieve forward compatibility with CUDA.
>>
>>110009640
It doesn't help that everything is very fragmented and constantly changing, I reckon a lot of people's setups are going to be ad-hoc and not necessarily the right thing to copy.
>>
>>110009151
god I need to fuck that
>>
>>110009828
It's even worse than that man, vulkan is faster than ROCm on AMD.
>>
>>110009471
>all I need is always the current best thing
Attention is All You Need
>>
>>110009829
That works, thanks.
>>
>>110009497
Now vibecode an animated miku that talks when the model responds
>>
yesss AMD sucks ass, rocm 10 is way slower than vulkan, pleeease buy njudea products saars
>>
>>110008989
Too much like a low-effort pallete-swapped kimi-chan. Nekomimi and shorter hair don't make enough difference...too samey 芋
>>
>>110008989
Doesn't match the essence of the model like >>110008913 does.
>>
>>110009934
Upgrade out of the driver that came with Ubuntu LTS and try again
>>
>>110009781
I guess it's 2026. Homosexuality is supposed to be okay now.
>>
>>110008460
>gfx1100+
useless
>>
File: 1780052147475541.gif (1.01 MB, 1306x1038)
1.01 MB GIF
>>110006381
I love the future. Live calls working, everything going smoothly.
>>
>>110010300
can i haz?
>>
>>110010300
this is why the white population is declining
>>
>>110009578
>How dare you do math!
Legitimately what is their problem?
>>
>>110010300
Oh hey I remember that 2am post.
What does it mean, 28 messages in a voice call? Were you talking to it physically and it responding in a separate tab or something along the lines?
>>
>>110010153
I use Arch, but I've tested other distros, docker containers, kernels, etc. it's all the same or worse, ROCm is just bad and the shills are just lying constantly wasting my time bothering to test it again and again hoping AMD has fixed their fucking shit and they never do.
>>
>>110010334
AI will eventually fix ROCm, trust in the plan
>>
>>110010332
It probably prints both the voice to text and LLM responses into the chat for debug purposes.
>>
>>110009697
Can you fuck off to r/singularity already?
>>
File: 1790849867182375.png (1.99 MB, 1860x1093)
1.99 MB PNG
>>110010332
>>110010339
Yeah, pretty much, our STT-TTS turns are logged for debug purposes, I can click to check on them, but now that you mention it, UI would be cleaner without the {N messages} on call markers.
>>
>>110010337
I have more hope for ZLUDA running on vulkan fixing it than ROCm ever working.
>>
>>110010392
>>110010392
>>110010392



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.