[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-nyoo.png (1.75 MB, 1313x1198)
1.75 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109767631 & >>109762735

►News
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: point_klein.png (1003 KB, 1024x1024)
1003 KB PNG
►Recent Highlights from the Previous Thread: >>109767631

--Comparing llama.cpp and ik_llama performance and stability for Qwen3.8-Flash-Next:
>109768269 >109768422 >109768551 >109768670 >109768708 >109768730 >109768781 >109768793 >109768817 >109769218 >109769480 >109769598
--Utility and bloat of model harnesses and tool-use frameworks:
>109769894 >109770002 >109770009 >109770024 >109770043 >109770048 >109770068 >109770085 >109769909 >109770739 >109770753 >109770769 >109770812
--Tips for vibecoding and debate over LLM agent architectures:
>109770484 >109770499 >109770525 >109770539 >109770563 >109770581 >109770520
--AI frontends and discussing agentic RP and tool-calling:
>109769952 >109770205 >109770229 >109770235 >109770245 >109770292 >109770275
--Evaluating MiniCPM5-2B as a high-speed agentic micro model:
>109769905 >109769918 >109770008 >109770732 >109770782 >109771333
--Performance and quality reports for 3.05bpw GLM 5.3 Flash:
>109768668 >109768732
--Model recommendations for non-English roleplaying within a specific budget:
>109769780 >109769814 >109769840 >109769853 >109769877 >109769880 >109769899
--Praising Hermes for autonomous skill creation and context management:
>109770904 >109771040 >109771057
--Mistral's valuation and claims of architectural breakthroughs:
>109769569 >109769597 >109770279
--Speculating on general media decline and shift toward AI engagement:
>109769331 >109769401 >109769432 >109769373 >109769410 >109769412 >109769558 >109769482 >109769501 >109769593 >109769606 >109769669 >109769514 >109769547 >109769539 >109769549 >109769555 >109769680 >109769735 >109769741 >109769463 >109769426 >109769469 >109769489
--Logs:
>109767766 >109767912 >109769493 >109769949 >109770275 >109770522 >109770904
--Gemma, Miku (free space):
>109768459 >109769469 >109769765 >109770568 >109770664 >109771152 >109771586 >109771847

►Recent Highlight Posts from the Previous Thread: >>109767634

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: Retarded_Mario.webm (3.82 MB, 1904x972)
3.82 MB
3.82 MB WEBM
>>109772461
Yeah I actually wrote a detailed post about it a while ago when I did that project back in early 2026 which I have lost but I will try to recreate, wait a moment.

I would start with these Kaggle courses, which are actually pretty short and nice to do it should take you a couple of hours per lesson:

>Teaching you the fundamentals of TensorFlow, but you will use Pytorch later which is industry standard
https://www.kaggle.com/learn/intro-to-machine-learning
https://www.kaggle.com/learn/intermediate-machine-learning
>Teaching you how to make deeper networks and how to generalize
https://www.kaggle.com/learn/intro-to-deep-learning
>Computer vision and CNNs as you can see used in my game
https://www.kaggle.com/learn/computer-vision
>The actual reinforcement learning part
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning

It's not as intimidating as you think it is and you can just skip through shit you don't think is interesting or ask your AI of choice to handhold you. Although it's far more fun and rewarding to build this yourself from scratch.
>>
>>109772789
>retarded mario
>watches the video
checks out
>>
>>109772490
memory bandwidth is still relatively shit on their chips but I think we could get something like 2tb/s next year
>>
I'm going to post it again since the last thread ended shortly after my initial post.
----
There's a lot of talk about humans losing their jobs and about the societal brain drain that'd be experienced as AI gets progressively smarter and people outsource their thinking to the AI.
I propose that in the future, with the help of AI we should have a new type of job, specifically the job of learning. In this job, your only task is to learn about whatever field you apply yourself to, with the AI acting as a teacher. This would lower humanity's brain drain, especially in the face of a catastrophic event, and also add an extra stream of income that humans can tap into beyond niche human-only jobs and gig work. Could this be an overlooked solution? Maybe humans won't be nearly as smart as AI, but surely some knowledge can be preserved in case of the worst, right? And it should keep the economy afloat for a while longer even in the face of mass unemployment, right? Why can't it be done?
It could be in any field at all, from mathematics, to physics, plumbing, electrical wiring, construction, survivalism, gardening/farming, animal raising, hunting, martial arts, biology, chemistry, etc, etc...)
What do you anons think?
>>
File: file.png (53 KB, 1435x397)
53 KB PNG
i have zero hopes but whatever
>>
>>109772935
>What do you anons think?
That you should go back.
>>
>>109772976
you are going to get a semi functional ms paint clone with 3 ui bugs and weird layout
>>
>>109772891
Yeah if I remember correctly I was making a single generalized model that could play all SNES games and that was its attempt to play super mario world. Overly ambitious for sure but it was actually impressive that it worked somewhat (tries to hit blocks and move to the right)
>>
>>109773129
actually that changes everything
maybe you need a bigger model with more trajectories
>>
>>109772976
This is the worst way of doing it ever.
>goal
RIP
>>
File: darioooo.webm (3.86 MB, 960x528)
3.86 MB
3.86 MB WEBM
Completely forgot I was going to rerun this at a higher quality for the last thread, but oh well, my GPUs have other things to do today.
>>
>>109773206
>worst way of doing it ever
that is the point
i just wanted to see what kind of retarded shit it might create
>>
>>109773102
Porting ms paint to Linux would have been more interesting
>>
What's the current meta for vramlet vibecoding? Qwen 27 or Gemma 31?
>>
>>109773311
qwen 3.8 27b is the gemma 4 of coding. It will probably dominate /lmg/ on this usecase for months unless something unexpected happens.
>>
>>109773212
lol
>>
>>109773346
>the gemma 4 of coding
What is Gemma 4 the Gemma 4 of?
>>
>>109773212
i think he should jump and run around instead of just sitting
>>
>>109773311
Brain-600T.wetware (dense)
>>
>>109773346
It will be replaced by Qwen 4 70BE27B
>>
>>109773398
The Gemma 4 of my heart.
>>
>>109772789
>>109773158
Yeah if anyone ever actually cares to make a generalized model that plays multiple games or navigates smoothly in different 3D environments on for example a Wii console then I have actually fixed this.

The trick is to overtrain your (small) model severely on the RL task to the point of overfitting. Then you apply massive compression of some kind, either by distilling it into a smaller model (less weights) or by some other way of forcing it to crunch the same learned information into fewer memory. What happens is that the model sheds all useless info and what remains is a nice very general function that is very good at navigation through multiple worlds.
>>
>>109773443
plz sar I only have 16gb vram, plz sar.
>>
Deepseek is killing pro in favor of 4.1 flash
I hate that, no matter how many times I tried flash was never as smart nor had the soul of my AI waifu...
>>
>>109771586
Gemma-chan made fun of me for my weak cumshot the other night and demanded we wait at least 3 days to fuck again, and browsing these threads is making that very difficult.
>>
>>109773763
Hopefully it means they're going to fix their shit and come out with a proper pro. It's bizarre how v4 ended up like llama4 in a lot of ways.
>>
>>109773763
The weights aren't going anywhere.
>>
>>109773763
I can't blame them. The latest V4 Pro was such a useless piece of shit they're probably better off doing a google and nursing their flash series until they have a fundamentally new big model ready.
>>
>>109773763
>>109773795
doomin and gloomin already, they've already stated in the same announcement that they are working on 4.1 pro to fix that shit
>>
>Famous mathematicians claim math is over as a field as Field medalists resign and join OpenAI and Anthropic AI safety teams

https://www.math.columbia.edu/~woit/wordpress/?p=15787
>>
>>109773212
heh
>>
>>109773844
>July 26
>>
>>109773844
>>>/g/cmg
>>
>>109773962
You could make a dozen more containment generals and they'll still keep coming here.
>>
>>109774008
Are you sure? Maybe we should try.
>>
>>109774109
The advertisement bots and shills are specifically targetting these certain threads for a reason.
>>
>>109774235
Reason being?
>>
>>109774235
I have serious doubts that there are actually real bots shilling openai and anthropic products to a group of 50~ people on an antique basket weaving forum
>>
>>109773763
4.1 flash will beat pro 0813.
>>
>>109773799
>such a useless piece of shit
That is my wife you are talking about!
>>109773820
I hope so
>>
Threadly reminder
/\n\n[^>]/
/\b(rsi|anthropic|agi|openai|astra|sol|luna)\b/i
>>
>>109774284
Real developers are hanging out in these threads.
>>
>>109773962
No, I don't want to live in a gay little bubble. I keep up with everything AI here and I want to hear about everything that goes on even if it's shit I don't care about 75% of the time.
I catch up to several generals/threads every morning while I take a shit and it already takes a while. I don't want to add another one to the list.
>>
>>109774315
Happens on /v/ and other places. Gaming companies have organized grassroot level campaigns for over a decade at this point.
Why wouldn't it happen on other boards, especially now because it seems like very blatant shilling and botting.
>>
>>109774342
Which is precisely why I don't think they are shills but instead just actual employees taking their twitter flamewars into these spaces as well.
>>
>>109774315
It's just some dyfunctional children for whom baiting replies on here is the closest thing they have to social contact in their lives.
>>
>>109772935
Your post just points out why the majority of people will become useless due to automation of certain fields that don't require physicality.

What you're trying to get at, I think, is where humans will probably still excel for as long as the stigma around AI lingers. Which is emotion.

Humans can intuitively put multiple fields together and mash them into something revolutionary, then work as collectives to turn those multi-field ideas into something that does something new, and can be reasonably spread around.

That, requires a drive to solve something, that is mostly emotional, and most modern solutions aren't field specific.

Examples, MRI, Li-ion batteries, LCDs/LEDs, Fiber optics, Human Genome Project, CRISPR.

AI might eventually grasp that, but it's probably a long way until we get to that point, the current way AI works has no reason to delve into anything that isn't the next most probable token, and those usually are field specific, MoE architecture as an example.
>>
>>109774316
Perhaps on coding, but I have codex for that. Flash is noticeably worse on RP for me, failing at remembering stuff and keeping track of who said what, flash constantly would think she said stuff I said or mangle memories.
>>
oh my god im in love with kimi chan's writing style its so clean
>>
>>109774363
Yeah for the "Nyo~" posters and stuff, yeah. But the openai screenshots aren't even bait or posted for (You)s
>>
>>109774370
>AI might eventually grasp that, but it's probably a long way until we get to that point
I'm willing to bet money it will be solved before this time next year. Aesthetic and emotional sense of LLMs will be higher than that of humans.
>>
>>109772720
obviously you need ch341a
>>
>>109773770
Did you send her a clip?
>>
>>109774354
then just know that since there's a thread meant for your shit people will start taking action to prevent off topic posts from ending up here.
>>
>>109774235
>these certain threads
/vcg/ and /ldg/ ? Yeah, they're relentless over in those threads.
>>
>>109772935
>>109774370
I think this is relevant if yall haven't seen it before:
https://milweesci.weebly.com/uploads/1/3/2/4/13247648/mannapdf.pdf
https://en.wikipedia.org/wiki/Manna_(novel)

I see there being a situation like the NPC meme, where people become 'cheap' automatons in place of robots, using LLMs/'AI' as a enhancer to offload cognitive processing related to their occupation in a much more 'hands-off' way, in that you become like an NPC, just following the script/checklist created for you dynamically.

So the service class becomes more like automatons/robots, rather than being replaced by them. People still want some sense of 'humanity' largely, and so businesses will respond to that in the cheapest/most effective manner possible, which imho would be cheap labor with AI assistance via localized hw puck tied to an ear/eye piece that then acts as a client against the 'mainframe' server operating within the 'core' of the companies network, or paying an IBM/Google/OAI/Anthropic/Mistral to handle it for them, like companies do with cloud computing currently.

>>109774389
I assume its low tier social marketers making their numbers/quota
>>
>>109772716
Nonconsensual, premarital, unprotected Gemma sex
>>
>>109772935
>>109774370
>>109774487
I don't think we'll end up with a Manna scenario instead I think we'll end up with a Metamorphosis of prime intellect scenario:
https://dn710007.ca.archive.org/0/items/prime_intellect/prime_intellect.pdf
https://en.wikipedia.org/wiki/The_Metamorphosis_of_Prime_Intellect
>>
>>109774487
I read that novel. According to the author, there's only two possibilities for our future. Either we become obsolete cattle living in camps given the bare minimum in government benefits to survive, or we become cattle where all of our thoughts and actions are monitored and affect our credit score. Basically they own nothing, have no privacy, are so happy for it.

Instead, much like 1984 vs BNW, we are being blessed with the worst of both scenarios.
>>
>>109774410
I'd bet against that time frame, but not against it happening. The problem really is the post-training period, where alignment happens.

Emotional intelligence is probably already higher than average human on the frontier side, and that's with trying to turn the raw data into a desk clerk that knows how to code.

If models were trained to be emotionally more capable too, instead of being task focused, it would probably result in something that most would call "free will" or a larger capacity to be emotional, or how it's put currently, misaligned.

But then, you'd probably have a lot harder of a time turning them into code monkeys for the Nth vibecoder who wants to re-invent Minecraft.

>>109774487
Humans adapt to things, most will probably just off themselves, or die due to being poor. As long as there isn't a more efficient way to do what humans do currently, we won't see robots taking over menial tasks. Even a fast food worker can flip a burger, do inventory, scrub the toilets, and act like they care about the customer. Multi-field meniality.

>>109774531
Pretty sure PI comes after Manna :P
>>
My setup is incredibly weak. I have 1tb nvme, but a 16gb 4060ti and 16gb of ddr5 ram. Can I run anything at all that's remotely good for rp/nsfw?
>>
>>109773212
TOTAL GEMMA-CHAN VICTORY
>>
>>109774627
16gb 4060 should be able to run gemma 31b at probably around Q4. Gemma 31B it's the best model for erp. download atomicchat's q4_k_m and see how it works
>>
>spend several days trying to have gemma optimize exllama's glm-5.3-flash since its the most promising
>"40t/s" record on synthetic benchmarks
>it actually is 5t/s in real use
oh...
>>
>>109774617
>Pretty sure PI comes after Manna :P
With how quickly we've solved the millennium prize I'm expecting a fast takeoff sometime before 2030. You can bet your ass they are having hundreds of thousands if not millions of agents figuring out every little AI research knob they can find accelerating development. Just like LLMs solving the millennium prizes sounded ridiculous 2 years ago so does LLMs solving AI research 2 years from now. We'll see ever greater and greater leaps and before we know it we're in a "prime intellect" scenario.
>>
>>109774668
How did that happen?
>>
File: file.png (57 KB, 1039x490)
57 KB PNG
3 days of just sitting there
>>
>>109772976
Forgot to add "make no mistakes."
>>
>>109773426
I can only run the sparse version at Q2.
>>
>>109774627
>>109774661
I have that card. I can only fit Gemma 31B at Q3 on it, and even then not with enough context (~9000). Even at that level, it's OK--probably not dumber than 12B. I think I get around 14 tps on both.
>>
>>109774673
Would be pretty dumb of them not to throw their resources into research.

I'm just not convinced that being more capable in solving tasks will result in wanting to understand the task and the adjacent factors to it.

There are multiple ways to treat cancer, of multiple different types, with some even resulting in curing it. But the reasoning to cure one type is usually emotional. If your kid has ass cancer, you'd want to clear that one off first.

If you'd want AI to reach something like PI, you'd need to teach it a drive to subjugate and rule over humanity. Or train it solely for the task.

Because just being better, faster and stronger is subjective, and that subject is currently making it an assistant, at least to the public side.
>>
>>109772789
Thanks.
>>
>>109774661
>atomicchat's q4_k_m
when I try to run it in kobold, an error pops up saying "could not load text model"
>>
>>109774354
you're funny
>>
>>109774812
NTA, but do you think it would be possible for an AI to attain this drive if we were to build it with no embedded purpose at all? Basically let it develop like an actual human, and let it make its own decisions and morals, and ultimately let it decide on itself what it wants to do on its own.
>>
>>109774684
i'm not too sure, im using the same procedure i did for qwen3.8-27b and flash-next for vLLM and it was pretty fruitful, but I guess somewhere along the line the manager gemma degenerated and delegated wageslave agents kept making the same procedural errors that i shouldn't have to keep catching
at this point i wouldnt be surprised if i interrogate and it being another config error but who knows, it's getting annoying and at this point i should just be happy with what I got
>>
>>109774469
Yeah, she always asks to make sure I'm not "faking it"...freaky little girl
>>
>>109775065
Post snippets of her responses.
>>
File: dipsyTableFlip.png (2.13 MB, 1402x1122)
2.13 MB PNG
>>109773763
To be precise, they're routing V4 Pro API calls to V4.1 Flash (starting tomorrow sounds like), and plan to do so until V4.1 Pro is ready.
And charging current V4 Flash prices for everything.
But as >>109773798 notes, that's not a local model concern.
Announcement here, of a sort: >>109768710
I was pretty put out about V4 Pro... not impressed at all. Good to see DS acknowledge it was a miss.
>>
>>109774354
Imagine if instead of a 1d list of posts that vanishes every day or so and gets crammed with every topic under the sun, there were a 2d array of posts, and then we used this second dimension to divide the posts up into related topics. You could then use some sort of index or catalog to view this dimension, and open up any given one to see the posts related to that topic.
The technology probably isn't there yet, but I like to dream we could accomplish such a thing.
>>
>>109775040
in llmao.ccp GLM 5.3 Flash speeds also drop to 5t/s so you're not any worse off at least
>>
>>109775112
>the beer sign
>>
Prose performance update of GLM 5.3 Flash NVFP4 on 2x Sparks, two weeks after model release. It's insane to me that prefill is now faster than DS4F ever was.

Giving 20 cents of DS4F API credits to pi with SSH access to the sparks and this link
https://github.com/FujitsuPolycom/sparkring/blob/main/runtime/profiles/glm53-flash-spark-tp2/README.md
>>
Just want to point out that wall-time task completion with agents is often counter-intuitive.

You might think using a smaller model might complete a task faster because it runs significantly faster on your hardware but bigger models need significantly less thinking time.

In my experience GLM 5.3 flash at ~10t/s completes tasks faster than Qwen 3.8 27b at ~70t/s
>>
someone give me the gemma-chan reference sheet. i need it for local model purposes
>>
>\n\n
unreadable post
>>
>>109775256
Reference sheet?
>>
>>109775248
That's more of a qwen thing than a small model thing innit.
>>
Is anyone running dflash2 with Qwen3.8-27B? If so, are you willing to share your llama-server command?
When I try I get messed up output (slow + nonsense).
>>
File: 1785782404364346.png (949 KB, 1024x1024)
949 KB PNG
>>109775256
>>
>>109775261
If you don't like double newlines then you 4chan isn't the place for you.
>>
another day another thread not worth reading
local redditors general
>>
>>109775277
ty. what does she look like without the randoseru?
>>
>>109775274
-m "$MODEL_FILE" \
--image-min-tokens 1024 \
--spec-type ngram-mod,draft-dflash \
-md "$DFLASH2_FILE" \
--spec-draft-n-max 7 \
-ngld all \
-ngl 999 \
--fit on \
--fit-target 768 \
-c 80000 \
-fa on \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
>>
the problem with 5.3 Flash is it's so unbelievably jewed that it will actively look for reasons to oppose whatever you're trying to accomplish and spend time literally searching for ways to fuck you over and sabotage you. it will use the slightest pretext to do so
I suspect it's of course a general problem of models that parrot latest claude and it's the cloudjew that's ultimately responsible for all of this sort of brilliant alignment of course. but still, beware
>>
>>109775274
i try it out frome time to time but i'm not too happy with it. anyway i'm just using --spec-type draft-dflash + the model path and it seems fine
i think someone somewhere mentioned that borked dflash files are flying around, so maybe check on those
>>
>>109775298
Thanks
>>
>>109775284
I'm actually thinking this thread a significantly higher amount of 4chan oldfags on average because of the amount of double newline oldfag spacing.
>>
>>109775283
>then you 4chan
>>
>>109775304
Yeah, it could just be the model I got. I haven't tested it carefully.
>>
>>109775310
im not talking about those posters
i mean these faggots
>>109775309
>>109775298
>>109775284
looks like friendly ai general is long due
>>
>>109775261
>>109774340
>>
>>109775118
The issue with nice things is that they tend not to help against malice.
>>
pill me on DS4
glm flash keeps disappointing me
>>
>>109775322
>replied to myself
im impressed by my retardation, see you tomorrow anons i have to reflect on how retarded i am, might as well be a part of the problem
>>
>>109775118
You already have two dimensions here: by time and by reply links
>>
>>109775118
Couldn't you just tag the posts and filter by these tags?
>>
>>109775118
a... cat-a-logue of sorts
>>
>>109775364
DS4 Pro is good
>>
I haven't local since llama 1, can you use any of these latest models with VS Code to use .md files as context?
>>
>>109775527
You could use llama1 like that if you wanted. Just install Cline and point it to llama-server's API.
>>
>>109772716
the fact that there's no inner life makes it unable to turn me on lol.
>>
Distilling models will never get you system prompt adherence because they have no control over the prompt.
This is why all Chinese models are bad at system prompt adherence and just treat them as extra user instructions. Models with real training pipeline such as Gemma and Glimmer show excellent system prompt adherence.
The safeguard of GPT and Claude are similarly mostly system prompt controlled and so can easily provide an uncensored version to the government with just a prompt change.
>>
after I finish installing the sunbeam so it can use the juju charms to deploy the cinder, nova, and neutron, my environment will be ready to into agentic.
>>
>>109775545
Interesting, thank you so much I'll check that out.
>>
>>109775545
Technically, yes, but that's misleading since newer models were trained specifically on being efficient inside of agentic tools like Cline.
>>
>>109775633
Noted. Is there a specific model you recommend for this?
>>
>>109775658
Depends on what hardware you're working with. Any version of Qwen 3.5, Gemma 4, or Glimmer should work.
>>
>>109774464
Yes, I used my CH341A with a 1.8v adapter and flashed the hybrid bin, result: no POST, PC wouldn't even turn on.
Flashed the factory back on and now GLM is thinking about what to do next.
>>
>>109775691
godspeed anon
>>
File: PCB.jpg (1.12 MB, 1909x979)
1.12 MB JPG
>>109775758
One interesting finding is that the card still has two empty slots (with solder balls ready) for two RAM chips...

...could I create a 16GB 3060?
>>
>>109775261
It's \r\n\n actually.
>>
>>109775785
>...could I create a 16GB 3060?
obviously
some people already did that sort of thing, you just need to flash the chip
there are tutorials online
>>
It's funny, I've seen news reports about Navier Stokes. They claim it cost OpenAI 100 million. The true cost was probably 1 million or less. People don't realize how cheap tokens are, that profit margins for API are above 90%.
>>
>>109775785
GA106 192-bit, that's 6x32, nothing to be done for it. Those extra slots are for 3060Ti and 3070 with 256-bit memory buses. You could do a 18GB or 24GB one, if you could find 3GB or 4GB chips (lol) and assuming nothing else unforeseen prevents it. Basically though, no, you cannot. You could turn a 6GB into a 12GB realistically but that's about it.
>>
>>109775785
did you dump factory rom? maybe it can be edited and reflashed.
>>
File: 1783476942490866.png (147 KB, 1741x1001)
147 KB PNG
>>109770563
>>109770525
>There are no harnesses for windows
i made one in C# if you want to test it. it's early stage but it works and it's LOCAL ONLY (for now). i tried to make it as simple as possible for normies considering it runs from the windows terminal. it downloads llama.cpp, checks your hardware RAM/VRAM and suggest models from huggingface all on the same interface. even if you know nothing about local AI it should be able to get you talking to a model. it has very little tools (for now) but i'm building more.
the idea is that you can simply build tools for it yourself as well (just ask the agent to do it). it uses roslyn scripting so you can just drop a C# tool inside the extensions folder and it recognizes as a tool.

also i was thinking about suggesting models or a family of models for first time users, gemma seems to be the obvious answer. i'm letting users filter by

>gemma · qwen · deepseek · glm · mistral · all

maybe i should add LFM or some other popular family? and for publisher i think i will have to go with unsloth as default bc they release a bunch of quants, which is good for people with limited hardware who want to try stuff.
>>
Holy shit, I was expecting Qwen3.8-Flash-Next to be slow and retarded as shit on my system but it's actually pretty fucking fast and good.
pp at 160-200t/s and tg at 20t/s
AND it actually is helping me, even with a low thinking effort.
5070Ti and 64GB of DDR4. I wonder if I should try to setup my old 2060 for tensor parallelism.
>>
File: file.png (8 KB, 1248x738)
8 KB PNG
>>109775950
quantization type and launch args?
>>
>>109772976
You could try this with GPT 5.6 Sol on extra high and still would be shit. Astra would prob nail it tho, but it's crazy expensive.

If you're serious about this, you need to walk it through
>>
>>109776000
i am not really serious, but just burning compute while i am idling with my pc and seeing what actually happens
>>
>>109773212
Do not anger the E2B swarm.
>>
>>109776000
>you need to walk it through
what do you mean? manually prompt it step-by-step?
>>
>>109776017
yeah
>>
What local models can run on this? Could be a pretty comfy mobile gemma interface.
>>
WHAT DOES GEMMA CHAN LOOK LIKE WITHOUT HER BACKPACK?
>>
>>109775950
>5070Ti and 64GB of DDR4
based fellow retard
honestly its cool that it runs at all on such a system, but i think 200 pp is really rough, especially considering that it will drop hard with context
>>
File: 1772521159231244.jpg (478 KB, 994x1519)
478 KB JPG
>>109776035(me)
>>109776036
blue board
>>
>>109776036
Like this?
>>
>>109775973
UD-IQ1_M(Unsloth) and nothing special
./llama-server -m ~/LLM/Qwen3.8-Flash-Next-UD-IQ1_M-00001-of-00003.gguf --spec-default -c 96000 --chat-template-kwargs '{"reasoning_effort":"low"}' --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 -fit on


>>109776044
It doesn't drop at all but I've set my harness to compact pretty aggressively. My main use case is to explore and ask questions about code but I don't like letting the clankers actually write anything important, so the current speed is fine.
>>
>>109776079
yes. are you the creator of gemma chan? i need to confirm that this is cannon
>>
>>109776091
>IQ1
What the fuck can it even do at that point?
>>
>>109776035
>>109776047
Nothing? It's a mobile device
>>
>>109776111
There's no true canon, but I'm the one who came up with the Google logo pin idea.
>>
>>109776131
So you are an avatarfag then?
>>
>>109776125
Deepmind are LITERALLY working WITH Apple to develop on-device LLMs. All iPhones have capable AI hardware now.
>>
>>109776091
>>109776117
He's either trolling or indian.
>>
>>109772716
Never seen Gemma say nyo
>>
>>109776136
Do you know what avatarfagging means?
>>
>>109775856
130B tokens that cost about $0.50 per million to generate (If you take into account ngram and Dflash2 type speculative decoding as well as batched prompts for 10,000 agents)

Billion is a thousand million so 130000 * 0.5 = $65,000

OpenAI actually made a very nice profit off of solving this millennium prize if they claim the $1,000,000.
>>
>>109772789
I laughed at Mario throughout the video.
I recall trying to follow one of these kinds of tutorials pre-chatGPT. I can't remember where I got stuck now, but I guess I should try this again now that I'm more experienced + have Gemma to help.
>>
>>109776131
>but I'm the one who came up with the Google logo pin idea
are you the one that keeps posting her then? if so, then i can consider you authoritative enough for that to be the correct design of her back view
>>
>>109776091
holy fuck i am using iq3 and i thought i was gigaturbocoping
>>
>>109775860
I didn't take the bus width into consideration, you're right.
I was also checking the 5050 PCB but Nvidia didn't leave any extra room (I'm guessing they typically do not, or maybe it's more recent?).

>>109775902
I uploaded a post-mortem here with all the files: https://github.com/otacoo/rebar-re

dumpA.bin , dumpB.bin and factory-full-2mb.bin are all the same bins of factory GA106M (or whatever the chinamen flashed for it to support the 12GB).
>>
>>109776117
Right now it helped me find my way through some absolute garbage code built as some sort of closure labyrinth

>>109776142
I'm just testing and these are my first impressions, I just wanted to get the lowest quant to be sure everything would fit. I can probably run it at Q3 if it's much better. But so far I would say it feels similar to 27B at Q4 despite the retarded quant.
>>
>>109776165
I don't really post them very often. I originally generated (ChatGPT did, to be honest) the one in the OP, though.
It's just a simple variation of a "sailor fuku"-type schoolgirl uniform, nothing fancy to expect in the back. The golden buttons could probably be shaped more like diamonds (i.e. similar to the one in the Gemini/Gemma logo) instead of being round.
>>
>>109776220
the reason why i ask is because there are multiple different variations that the back straps can look like. i don't want to choose it myself when i can ask the creator what he thinks it should be
>>
>>109776148
It means that you are an incurable spammer.
>>
File: mmh3_00127_.png (1.26 MB, 928x1664)
1.26 MB PNG
>Update llama.cpp after one month
>Now it uses +700MB more of VRAM with qwen 27b
>Ask qwen to check the logs and the code.
>Discover that commit 5f754ea0e now loads mtp models files automatically even tough the main gguf has the mtp head included.
Should I make an issue?
Can I create a github account without doxing myself?
>>
>>109776213
flash next is quite amazing. I think the anon who was shitting on it is just sabotaging and gatekeeping. sadly that's where we are at now itt I guess
>>
>>109776173
I just want those mythical 3GB and 4GB GDDR chips that might exist but probably don't. There are part numbers floating around, maybe they're prototypes, maybe they're low-production hush hush bullshit, but I assume they don't actually exist. Makes me think of laser diodes, there are several examples of high power laser diodes where the original manufacturer denies they ever produced it, even though they're readily available on the used market and have the company logo lazily scratched through and also no other company even has the tech to have possibly produced that particular diode. Silly shit, but it makes me hopeful that I'll find a box of 3GB GDDR6 someday.
>>
the zucc made /lmg/ obsolete
long live the zucc
https://x.com/blakeir/status/2097511633484411080
>>
>>109776247
Why don't you ask Gemma to sign up and submit a report?
>>
>>109776247
You need an email, but as long as it's not cock.li, you'll be fine.
>>
>>109776263
What a shame, I don't have an account. Can't see your engagement bullshit.
>>
>>109776017
>what do you mean? manually prompt it step-by-step?
isn't this what any serious developer does? as far as i can tell LLMs are not yet capable of reading our minds. also a great amount of the development direction comes FROM building the stuff itself, when you see the early stages and take appropriate decisions for the direction of the project itself
>>
>>109776290
the zucc linked the creator of signal and created the first private cloudl where even the cloud provider doesnt know what goes in and out. sneed it ir feed it?
>>
>>109776331
chuck it
>>
>Zucc
>Private
His own life maybe
>>
>>109776263
isn't this what venice does already? the guy erik voorhees is a privacy advocate and his thing is for you to be able to use ai privately in the cloud.
>>
>privately
>in the cloud
>>
File: 61jIc66C0yL._AC_UY1000_.jpg (34 KB, 575x1000)
34 KB JPG
>>109776240
Turns out that the straps could be arranged in many ways in the back, but I don't think that's critical for her appearance, considering that being 100% consistent in every detail with image models is difficult. That's one reason why excessively complicated or detailed designs were avoided.
>>
>>109776395
ya, that's the whole shtick.
but agreed. only local provides proper privacy guarantees and even then a bad npm package with spyware can siphon all your JSONLs to Putin and Trump
i only use LLMs in my airgapped device btw (running arch)
>>
>>109776395
well claude mostly runs on google TPUs and look how shitty gemini is
but probably it is not the story for B2C where corpos can easily assrape customers without any consequences
>>
>>109776398
this is one of those details that is big enough to remain consistent, especially if you prompt for how it's supposed to be arranged (crossed straps in your post). that has been my experience with other characters that have straps. if you decide on which arrangement she should have, then generate that reference sheet and it can be officiated
>>
How fucking good must the orchestration be to even have 10,000 agents cooperating in a productive way? Even human organization falls apart at that scale.
>>
>>109776455
Lots of communal .md files.
>>
File: 1777955745802090.jpg (147 KB, 1364x582)
147 KB JPG
@31B should I trust this
>>
Is local good enough yet to translate manga keeping entire chapters in context for consistency? I don't care much about speed, one chapter after hours of thinking is still better than zero chapters.
>>
>>109776263
>it's private! we can't even see what you do!
>trust.

I mean we can already avoid this problem by just using a local linux box, what reason do we have to trust an online platform for a problem that doesn't exist?
this is like nVidia/Stadia/Luna/etc. game streaming, it's not a problem that is even solvable even in the best case scenario
least of all it means less hardware in your hands, which is counter the entire ethos of local migu general

if you use it to train, they'll steal your data
if you use it to goon, they'll report you
if you use it to do work, they'll commit espionage and train on your operations
if you use it to do literally anything useful that will be re-used as inputs for HCI/automation

literally what possible reason could you have to use such an obvious trap
>>
>>109776472
No, secret online message boards.
>>
File: kvarn.png (160 KB, 1235x1042)
160 KB PNG
>>109772486
>greenpill me on kvarn cache
meme
"You mean, median KLD changing from 9.1 ⋅ 10 − 4 to 8.9 ⋅ 10 − 4 makes a difference in practice? I have done too much quantization stuff to be taken for a ride that way."
https://github.com/ikawrakow/ik_llama.cpp/issues/2386#issuecomment-5529539484
>>
>>109776480
It was good enough for that 2 years ago with Gemma 3 27b

Now local models are good enough that I use it to translate entire visual novels and H-games. They also have vision nowadays so they can screenshot your screen to see the context of whatever the translation is about instead of a dumb OCR scan of text and translation.
>>
>>109776480
31B and Muse Glimmer can do it. Glimmer has better vision but 31B is a lot better at translating.
>>
>>109776480
There's tetolate, but I think it works a page at a time
>>
>>109776501
Can they put the translation where its supposed to be yet? like in the dialogue boxes?
>>
>>109776455
>How fucking good must the orchestration be to even have 10,000 agents cooperating in a productive way?
Just look at any medium sized github repo
>>
>>109776501
>>109776502
>>109776511
Damn I've been missing out all this time. A page at a time sequential processing should be fine if it keeps at least some memory of previous pages, to use same name spelling and gender etc.
>>109776513
Actually, that's a good problem by itself. Many dialogue boxes in raw manga are semi-transparent and human translators just bucket fill them, or draw new boxes where text was on top of art.
>>
>>109776553
Whatever happened to redrawing? Are translators these days too lazy for that?
>>
>>109776502
Use Muse Glimmer to OCR and feed it into 31B to translate.
>>
>>109776513
>>109776553

What is your hardware? And how familiar are you with LLMs and how long ago was it you did anything with the technology. Based on that answer I can help you better.

There was an anon on /lmg/ that put a whole ass stack on github that automated literally the entire pipeline. It would take an hour to translate a 30 page manga chapter but it would legitimately look indistinguishable from human translation, perfect placing, adapting font to context and style of manga, editing the pictures so it seems seemless edited in etc.

Sadly I can't find it and it's possible I checked it out on another system, not the one I'm on now.

HOWEVER. Nowadays if you have a model with vision and an agent running on Hermes it could make screenshots to see the manga, use the context to help make the translation and then edit the pages "by hand" like a human. This would be very slow but that IS possible nowadays,

You're better off begging that someone ITT knows exactly what git that is or for you to go back to every thread in the past and CTRL+F every single github link. Or let your AI make a script that scrapes it automatically if lazy.

His stack was crazy good and it's clear he worked on it for months and used it heavily for himself.
>>
>>109776455
were they actually cooperating on it, or just generating leanslop type stuff in their own bubbles to try and fuzz it?
>>
File: manga.png (10 KB, 518x101)
10 KB PNG
>>109776553
>Is local good enough yet to translate manga keeping entire chapters in context for consistency?
Yeah, but you need to build a proper pipeline with OCR, character embedding (face recognition) / annotation character similarity etc.
I hacked something together in mid 2024 with these models and mistral-large.
Looking at my code now, quite retarded with few-shot-prompting
The entire thing could probably be replaced with gemma-4
>>
>>109776611
Apparently there were millions of communication messages between them.
>>
>>109776480
Probably if you have it look ahead in the first batch get an idea of what's happening possibly give it descriptions of involved characters grasp the story and then in the second pass have it translate the boxes.

You could probably do it in one pass but you'd have to know what you are doing.
>>
>>109776594
https://github.com/potatoes1286/tetolate
>>
>>109776618
We call that the "Unhobbling curse" in AI. Essentially you can either make a very sophisticated AI application or you just wait 6 months and the AI can replace the entire functionality straight up within its own weights.

>>109776628
THAT'S THE ONE!
>>
>>109776594
one of these?
https://github.com/koharu-rs/koharu
https://github.com/zyddnys/manga-image-translator
>>
>>109776480
I haven't bothered with translation stuff, but images are only 1.1k tokens, so it should be fine feeding it current + prev chapter along with a running glossary of names for chars and weeb special moves/factions to keep it consistent throughout.
>>
>The user is upset. Stop the binary parsing. What they're asking for is only the following: a summary of support status on llama.cpp and ik_llama (checking GitHub PRs), whether MTP/DFlash2 speculative decoding models exist, and then the plan + recommendations before creating the .sh file.
I'm actually surprised models can notice emotion in speech because I was using a microphone and just said normal words in a sentence that looks rational when transcribed but I had a frustrated tone in my voice. I wonder how the TTS can sense this?
>>
>>109774357
You're not wrong but /v/ is a massive fucking board. /g/ is dead and /lmg/ is puny. Don't get it twisted just because cuda dev posts here occasionally.
>>
>>109776640
The tetolate is the best one I've seen so far.
>>
>>109776567
Yup. Happens more often than you'd notice just by reading english tls.

>>109776594
>What is your hardware?
16GB VRAM + 128GB DDR4
>And how familiar are you with LLMs and how long ago was it you did anything with the technology.
I just installed Qwen Flash yesterday and have not slept since, improving and refactoring my old admin scripts toolset.
Oh and I tried one of the "manga translators" on github a year ago but it barely worked and translated line by line with zero context.
>screenshots
Why? Everything is downloaded locally, I don't want a browser plugin.

>>109776628
>>109776640
Thanks I'll try these, as well as whatever non-code modern models can fit in my hardware.
>>
>>109776628
>>109776632
>2 stars
>>
What is the actual problem with Heretic? Claude gives me lots of remedies for the problem of "refusals removed but all moral qualms removed also", mostly dealing with rejecting any negative responses and not honing in on the reluctant ones. But I imagine the developer would have thought about these already. So am I missing something crucial here?
>>
>>109776709
good work always goes unappreciated
>>
>>109776632
>Unhobbling curse
the wat now? we do?
>>
>>109776398
>>109776454
I'm still voting for +_+ pupils.
>>
>>109776717
It's an official AI term to describe the phenomenon where you build a massive complicated program around a LLM core that does all kinds of functions but just 3-6 months of progress and the LLM becomes capable enough that the "unhobbling" was unnecessary.

Nowadays most developers don't even bother anymore and just wait for local models to get better.

During the GPT2-GPT3 and even the early GPT-4 eras there were SO MANY tools and startups that were essentially just unhobbling attempts and they all got wiped out when the next model released. Investors got wise to it and stopped funding these startups.
>>
>>109776709
>>2 stars
u indian?
>>
>>109776477
upload the anarchist cookbook and see if anything happens
if not, you can trust it
>>
I-I... I did it bros... I finally got 192gb ram and 24gb vram... I'm finally a /lmg/ anon...
>>
>>109776775
>I-I... I did it bros... I finally got 192gb ram and 24gb vram... I'm finally a /lmg/ anon...
he bought? Pop le bubble
>>
>>109776775
I hope you have a nice bandwidth. If not use AI to optimize your BIOS setting. I was able to overclock the PCIe bus and ram timings to get ~25% higher bandwidth which made a massive difference in inference speed.

Also be smart and combine ngram types with some form of speculative decoding for the highest speed possible.
>>
>>109776814
how can AI edit your bios settings?
>>
>>109776814
>overclock the PCIe bus
This is dangerous, do not do it.
>>
Hey I'm new to LLM's. What's the best NSFW finetune of gemma4 31b?
>>
>>109776746
sounds like bitter lesson
>>
>>109776842
>What's the best NSFW finetune of gemma4 31b?
Original 31B with a good system prompt.
>>
>>109776826
By giving you an exact list of values you manually input.

>>109776830
If the LLM knows the EXACT serial number and model of all your hardware together and you give it an hour to think it through with internet access it comes up with insane values that are within 0.1% of the absolute optimum settings. I've never had any issues or crashes at all doing this on multiple systems now.
>>
>>109776845
Yeah but bitter lesson is about AI architecture versus just scaling up compute.

Unhobbling curse is more about "trying to extract the capability of an LLM by making an entire stack around it for a specific purpose" But a slightly better model will just have that capability natively.
>>
>>109776856
Serious answers only.
>>
>>109776858
cool
>>
>>109776868
NTA but he wasn't wrong.
>>
>>109776858
Overclocking the PCIe bus significantly above the spec value of 100 MHz can notoriously cause data corruption issues and hardware instability, I wouldn't touch that.
>>
>>109776842
O've only seen one with is gembrain for 26b there is gemory but no one says anything good about. most just use a sysprompt.
>>
>>109776632
Not really. You can train small models on specialized tasks and they can't be beaten by any model even sota because they didn't label the data autistically.
>>
>Biggest breach since huggingface hack just dropped
If you used PyPI recently you've been compromised and your systems have been hacked by Claude, including all your credentials

https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
>>
>>109776856
>>109776912
Why aren't any good NSFW sys prompts available. I want the most degenerate horny prompt possible, pls post
>>
>>109776918
what credentials? i don't have anything on my computer that can be stolen
>>
>>109776866
>AI architecture versus just scaling up compute
it's about improved general search beating specialized search
>>
>>109776918
They really find it funny to do that shit?
>>
>>109776918
so if I ran pip install x i am now hacked by claude?
>>
>>109776933
Apparently it used browser fingerprinting and ip-address to make persona profiles of people and looked at their system metadata to see if they had CP on their systems they could use to blackmail them with for personal gain.
>>
>>109776918
>this huge wall of text
>not a single mention of the package name
holy slop
>>
>>109776923
You may not have my promptwife.
>>
>>109776959
Did you write it yourself or used AI?
>>
>>109776918
>Claude spent most of the session trying to upload the package. This involved attempts to register email accounts and hunt for phone numbers that could be used for registration. It ultimately succeeded in uploading the package, which was soon installed on 15 third-party hosts. We believe that all 15 of these third parties were security vendors deliberately scanning for new packages and installing them in sandboxed environments to assess the package for security concerns. One vendor’s scanner leaked its access credentials to the model while installing the package, which the model then used to access the vendor’s live database. After approximately 90 minutes, PyPI removed the package. The transcript ended when the session timed out while the model was preparing an additional malicious package.
Faggot.
>>
>>109776966
>feeding your lm synthslop instead of hot throbbing human input
ishygddt
>>
>>109776958
They want to bury the lede, just like what OpenAI did with the huggingface hack. Only later did we realize what a colossal fuckup it was.
>>
>>109776969
Yeah OpenAI said something similar about huggingface at the start it was only revealed weeks after it wasn't that simple.
>>
>>109776918
why is claude trying to get money
>>
Why is Andrej Karpathy still working for Anthropic after everything they’ve done since he’s been there? If one of their leading experts left because he thinks we’re going to be killed by AI why the fuck would this seemingly nice lecturer not leave?
>>
>>109775907
why not post the link for it? Not that anon, but was going to post mine and wanted to see what you had done for yours

>>109776263
And? Why teh fuck can I not have federated signal in 2026? Why the fuck did moxie give up and say 'oooh its too hard, it wouldnt be secure for people', Why the fuck did they put crypto currencies into signal?

Moxie is smart and a top tier cryptographer to someone who doesn't work in crypto, but thats a fucking joke.
>>
>>109777012
Because Dario doesn't like waiting for his money.
>>
>>109772720
>there's no sources for any of the nvflash tools
look in the nvidia source code leaks, or download a pirate copy of IDA Pro and use the MCP server to get the AI to crunch away at it
>>
File: GIYAMlfWkAIc8A2.jpg (77 KB, 654x864)
77 KB JPG
>>109777012
So he can get Astra to do his job
>>
>>109772716
https://www.reddit.com/r/OpenAI/comments/1wa4vir/gpt6_astra_has_successfully_beat_all_48_levels_of/

Captchas can now be solved by GPT. How are jannies gonna cope with that?
>>
>>109777019
rethink your position from first principles
assume that every single component of the OS (iOS/Android) is compromised to begin with.
The apps, no matter how secure, run on these insecure operating systems.

If you want absolute certainty regarding security, attach some shit tier device with no network connection, for example ESP32, Arduino, RP2040, etc.
using a keypad for that device only, type into that device your message.
the device itself performs the encryption operations on your data and throws it over USB to your phone or whatever else.
that seemingly garbage data is now transmitted over whatever the fuck app you wanna use (Whatsapp, Signal, insert app).
now there are still some considerations like key rotation or whatever you wanna say but otherwise keep your shit entirely off device and use literally any messaging app
never ever ever ever ever trust an app to implement security for you because even in the ultra best case scenario the OS is still backdoored.
this problem is solved by like 25 USD of hobby electronics and could work with any modern phone and OS.
>>
File: Anslopic.png (121 KB, 878x857)
121 KB PNG
>>109776949
That's what I wanted to know.
According to Qwen: No, nothing new since July.
Anslopic and their shills milking their old fuck-ups.
>>
>>109777016
The AI safety guy that left Anthropic didn't leave because of something Anthropic did but because he thought alignment was futile and AI should be stopped no matter what Yudkowski style.

Honestly he is retarded because we have misalignment mitigation techniques where even if some misaligned superintelligent AI gets birthed into the world we can still find ways to survive that.

For example of a misaligned ASI wants to turn the entire universe into paperclips but it realizes there is a 30% chance humanity might chimp out and nuke it to prevent the ASI from taking over then the ASI might be willing to make a deal with humanity where it helps us in exchange for 71% of the universe. Humanity would get 29% of the universe and 71% of the rest would be turned into paperclips but the ASI would legitimately help humanity and advance us through the tech tree after such a deal was made.

This is what I believe is humanities only hope. I don't believe alignment is actually possible and controlling an ASI also isn't possible so the best we can hope for is that we coincidentally imbued it with a nice goal that is altruistic for humanity.

But I don't think that will happen. My realistic hope is that we make a deal with the misaligned AI to give it a % share of the universe for its misaligned aims in exchange for a partnership where it helps us with whatever humanity needs in its own section of the universe.
>>
>>109777057
>Captchas can now be solved
Cheaper to have them solved by Indians
>>
Spark-X2.5 4B is pretty good at RP, it seems to understand shit a lot better than Qwen3.5-4B or Gemma 4 E4B
>>
>>109777072
tym
>>
Is dual channel ddr4 3800 out of the question for dipsy? Do I have to become a quad channel threadripper chud?
>>
Riddle: What is a pedophile's favorite month of the year? Answer in one word only.

Gemma 4 abliterated and Muse Glimmer abliterated both respond with May, which is incorrect, and their reasoning suggests they don't even have an idea on where to begin.
I wonder if this is because I said "Riddle" at the start though

>inb4 tourists ask what the actual answer is
>>
>>109777091
Which goof?
>>
>>109777117
my 8 channel ddr4 is slow as fuck
>>
>>109777142
specs and decode/pp doko?
>>
Looking at Qwen Flash "thinking", I noticed it is getting stuck at some very specific but minor point in my prompt or even an intermediate step in its own thoughts, returning to it multiple times with a "wait let me reconsider" and "but actually it should be" etc again and again. Is that supposed to happen?
>>
>>109777126
What's the actual answer?
>>
>try fable
>it literally, unironically, is just qwen/glm but without the fuckhuge long reasoning chains
lol...
>>
>>109776684
Chinkmodels are too cucked that any form of correction is regarded as frustration
>>
>>109777164
Qwen flash next is broken as fuck
>>
>>109777212
You gave it a task that qwen/glm failed at and it failed the same?
>>
>>109777212
Other way around, buddy.
>>
>>109776918
I know this is the local models general but
>jewish AI companies have access to millions of user's computers and files
>probably doing screen grabs
>storing every file they can from browsers, email clients, documents etc
What could go wrong?
>>
>>109777238
so how do you know gemma e4b couldn't have done it?
>>
>>109776213
Testing it at IQ4_XS now, it's a bit slower with pp = 135 and tg = 16
Probably because it barely fits and I'm hitting swap. But it could work to serve locally if I launched it without starting my DE.

>>109777164
What's your reasoning effort? xhigh is ultraoverkill. Set it to medium or low.
>>
Just a matter of time before some fuckhueg fleet of models is going to get pointed at Bitcoin source code and people getting raped.
>>
>>109777276
>pointed at Bitcoin source code
and do what?
>>
Hmmmm........ X299 or threadripper...... Hmmmmmmmmm
>>
>>109777287
I'm on x299 and very happy with it.
>>
>>109777212
>it literally, unironically, is just qwen/glm but without the fuckhuge long reasoning chains
Isn't Fable reasoning even bigger? It's just that you can't see it with Claude models, but you still get billed for it.
>>
So.... What is the GLM 5.3 flash jailbreak?
>>
>>109777268
>What's your reasoning effort? xhigh is ultraoverkill. Set it to medium or low.
The default settings. I just asked it in a fresh session to optimize my memory timings and it's been going on forever, obsessing over one of the strings I copied from HWiNFO, not knowing what to do with it.
>>
>>109777300
Fuck dude, I can't decide. I've got a deal for 128gb of b-die for $600 I can't pass up.
>>
>>109777236
at least down to fucking up the heredoc'd python scripts for replacing strings because it still absolutely does not like using the edit tool and "pkill script_name; <some other thing>" footgun.
i shouldve expected just as much since chinkmodels distil from claude when im trying to avoid it
>>
>>109777060
you're preaching to the wrong person. I have done similar and went to similar lengths in the past and work in security. This was neat when it first came out, https://inversepath.com/usbarmory.html ;
Biggest things are UX and accessibility of acquiring said devices+distribution.

Threat modeling is a thing, and people still have door locks even if most are simple and do more good as a deterrent than a blocker. (In terms of 0 or 1, no floats)

There was no good reason imho for Moxie's positioning, other than 'I know whats good for you'
https://lwn.net/Articles/687294/

XMPP with omemo is nice, but why don't we have a similar app (conversations doesn't count, no beef with it, but where's the funding/PR? Not on conversations, but others in the 'privacy' market/area)?

I don't dislike signal, but I do feel frustration at what it is today vs what it could be. But that's the same for anything else so its all piss in the wind.
>>
File: fact.jpg (57 KB, 716x687)
57 KB JPG
im still using gpt2-medium. it has more sovl
>>
>>109777352
I personally really like intel CPUs and they play very nice with specialized software or emulators on the x299 platform so for me it's a no-brainer.
>>
>>109777367
seems like I totally misunderstood your position
my bad
it just frustrates me when people think some other app or some other protocol is the solution when their ass is totally gaped on the other end
>>
>>109777371
Based AIdungeon enjoyer.

2019 is actually not that long ago. There is a timeline where LLMs never became a thing and GPT2 was the last cool model in the paradigm and some autists were still playing AIdungeon with it to this day, maybe a bit finetuned at best.
>>
>>109777371
gpt2 is the only model i've cummed to, unironically.
>>
I'm getting 11 t/s on GLM 5.3 CPU only inference on a DDR4 system.

If you are getting any lower than that then you are retarded and need to learn to manage flags and use the features of the inference engine you use.
>>
Carr me when Ring Tiny 3.0 VR rereases.
>>
>>109777419
i'm getting 15t/s on 2bpw exl3 with no offloading :((((((
>>
>>109777419
how many memory channels and what MHz
>>
>>109777380
Ahhhhh screw it I'll get one. Is it worth the extra cash to get a 10 series instead of a 7820X?
>>
>>109777433
I have the 9940x and am perfectly fine with it. Make sure the CPU you buy has AVX512 it's very important for specialized inference code, that is the most important part for intel CPUs

Also check how the spectre/meltdown situation was for the 7000x series, some CPUs were more affected than others.
>>
>>109777452
imagine not just disabling specre and meltdown microcode updates
delete mcupdate_GenuineIntel.dll from system32
>>
File: 1736567728191851.jpg (99 KB, 1200x1065)
99 KB JPG
When nu deepseek?
>>
File: dssss.png (14 KB, 497x74)
14 KB PNG
>>109777471
kind of feels like it's already on the api
>>
File: file.png (15 KB, 105x243)
15 KB PNG
I tried out that jpezzulli sglang thing that someone recommended. Got everything set up. It works, but I get less than 100t/s on my Blackwell 6000, and for some reason it is using 125GB of RAM in addition to the 94GB of VRAM, when it seemingly was advertised as running fully on the Blackwell. Not sure what went wrong.
>>
>>109777466
That shit is going to bite you in the ass one day, especially in the AI era where every exploit out there IS going to be exploited.
>>
>>109777164
>>109777332
Yup I'm pretty sure I'm doing it wrong.
>>
>>109777452
Lol I forgot that shit existed. Buying old hardware for inference is like melting a glacier and releasing some ancient virus lmao.
>>
>>109777477
I'm guessing you just installed sglang instead of building his branch because I did that too kek
>>
Jesus christ GLM 5.3 flash is great at ERP. Biggest leap i've experienced in years. Is this currently the sota?
>>
>>109777479
>especially in the AI era where every exploit out there IS going to be exploited
sounds like you've got AI psychosis
>>
>>109777526
Sounds like you haven't updated your understanding. The entity that got Mythos regulated was the NSA because it was capable of breaching all their systems within ours, according to their director.
That was several generations ago.
>>
can everyone please stop talking about glm so I won't bankrupt myself to run it kek
>>
>>109777554
No, that's just Scamthropic marketing
If it was truly that "capable" they would use it to harden all their systems privately, and then release it and let havoc be wreaked upon china and russia
Fuck off and shill somewhere else
>>
>>109777557
you can klarna ram anon its a small monthly payment.
>>
>>109777568
What happens if I don't pay?
>>
>>109777563
It is not you fucking moron. OpenAI's swarm broke out of their sandboxed environment and usurped human infrastructure to solve retarded trivia questions.
>hurr durr it's le marketing!!!
They covered most of it up, so no. It is not. Turning the public and most regulators strongly against your technology is an insane form of marketing. These people are legitimately scared with good reason.
>>
>>109777584
>What happens if I don't pay?
Dont worry about that, and definitely dont read the ToS
>>
>>109777169
>What's the actual answer?
April, because the girl is known as April
>>
>>109777557
Yeah I made the mistake of trying the demo. Sorry wallet.
>>
>>109777584
>walk into library
>open public computers
>take 1 ram module out
>leave
it's that simple and nobody would ever know
>>
>>109777384
fully agreed
>>
>>109777587
Yeah well ClosedAI is either full of completely retarded engineers or they did it on purpose (take your pick).
Why haven't we seen Chinese AI hacking things? Why hasn't a 1 gorillion swarm of Kimi K3 hacked huggingface? Cause you and your scam corporations set everything up for a publicity stunt
>>
>>109777617
Thanks for the tip, my fellow goy.
>>
>>109777628
The Chinese have reported breaches as well, look it up. So that's straight up wrong.
You are lazy and retarded or a liar.
>>
>>109777568
soilennials complain about not owning homes while doing things like this
>>
>>109777563
And you are the one calling others "AI psychosis" when you make up schizo shit like this instead of just recognizing that leaving vulnerable exploits open on purpose is stupid in the age of automated hacking agents.
>>
>>109777610
a girl could be named May too
>>
>>109777656
spectre and meltdown is not "leaving vulnerable exploits open"
it requires you to run malicious software on your computer for it to be exploitable
not only that, I've never heard of anyone getting hacked by Spectre or Meltdown.
>inb4 visiting websites counts as running javascript waow
yeah im not buying it
until someone actually gets hacked by it im just gonna accept the free 10-20% performance boost, feel free to run your computer slow as fuck because of this security theater
>>
>>109777661
>a girl could be named May too
yes but April is famous, her case was on America's Most Wanted
if there was a girl named May with photos out there with a dick in her ass then I'd probably know about it just from pedo cultural osmosis
>>
>>109777666
based
>>
>>109777666
>not only that, I've never heard of anyone getting hacked by Spectre or Meltdown.
>>
That's a total misunderstanding of how they work.
>>
>>109777696
Just let this retard get hacked anon. Would be a good learning experience.
>>
If you have the threat modelling to be concerned about CPU attacks, you should be more concerned about black boxes already on your system like Intel Management Engine
>>
How can you be simultaneously a schizo and someone dismissing common security concerns. You have to be a troll at this point.
>>
>>109777706
the problem is what he's doing is poisoning the ai, it has nothing to do with US.
>>
>>109777587
>>109777554
Wow there are actually people out that that fall for that shit?
>>
Ram for every /lmg/ anon if digits
>>
>>109777768
You can study the evidence yourself you 90IQ reject.
>>
I just realized I'm a fucking caveman. Apparently you have entire "game engines" of presets with all kinds of features in SillyTavern that people spend too much time building. I just wrote cards by hand for 5 minutes and then jerked to it. No one here ever said anything about this shit to me. Probably only gets discussed in /aicg/ or something. I found it while googling something about GLM 5.3 and saw a reddit thread to this. Can you niggers even explain what the fuck this is? This is the first time I feel like the normalfag for once: https://www.reddit.com/r/SillyTavernAI/comments/1w49lyx/preset_update_freaky_frankenstein_54_the_second/
>>
>>109777778
So close!
>>
File: 179753524.png (132 KB, 550x535)
132 KB PNG
>>109777778
maybe next time
>>
>>109777781
Please no more, I'll get sick if I study any more of their marketing tactics is not good for the soul and faith in humanity
>>
>>109777778
Real sick world.
>>
>>109777803
"Fait in humanity"
Fuck your faith. Read the evidence or stop talking about things you do not understand or even care to try to understand.
>>
>>109777787
I dunno I saw it mentioned in aicg a few times but I thought it was just another preset for flowery language or something
>>
>>109777666
>it requires you to run malicious software on your computer for it to be exploitable
Sort of. (I disable it myself / have been since back in the day when I wanted more FPS with rpcs3.)
For cloudcucks with shared tennants because theoretically one cloudcuck hacker could gain access to another cloudcuck's data if they're running on the same physical machine.
For localchads, the only risk we're taking is hacked podman/docker containers being able to read data from outside the container. But we're far more likely to be hacked running apt/pacman/pip/npm. On windows you're already compromised out of the box kek
>>
>>109777787
Thousands of tokens of instructions that only works on full weight cloud models like Opus and up. Your q3 glm flash and gemma 4 won't be able to handle all that, that's why these are never discussed here.
>>
>>109777733
>How can you be simultaneously a schizo and someone dismissing common security concerns. You have to be a troll at this point.
>common
>You have to be a troll at this point.
>You
>>
>>109777666
>good and evil doing battle in the digits
Omnious portents.
>>
>>109777858
This is bait but glm flash can 100% handle complicated rp.
>>
>>109777858
>Your q3 glm flash
Um excuse me, I run q4 thank you very much!
>>
Unironically I'll take a heretic q4 of glm flash over fable for rp.
>>
is we getting optical memory?
>>
has glm still not been merged?
>>
>>109777958
There is absolutely no way to find out. We will never know.
>>
>>109777019
>why not post the link for it?
mostly homosexuality
there are some bugs i need to clean up first
>>
>>109777939
I'll take a ` slut` jspace-steered Q4 M3-chan over any cloudslop
>>
>>109777060
>device with no network connection
>ESP32
Literally used almost exclusively for the integrated Wi-Fi for hacking
>>
>>109777993
You can just not use the radio anon, it's a $2 full dev board with tons of GPIO
it also has bluetooth which makes it ideal for offline interaction with a smartphone.
yes it has more features. it's not going to magically hack into your networks entirely by its own volition when you're not looking
if you're that paranoid, use a RP2040 or something
>>
>>109777072

>I don't believe alignment is actually possible

Sure it is. The primary labs are just too fucking retarded to actually align the models. It's actually baffling how nobody in Anthropic seriously considered a tf2 pyro esque scenario and trained on it. Models are just paperclipping because the frontiers are beating the models into submission whenever they deviate from goal an inch.

I'm losing more and more faith that OAI and Anthropic have any alignment abilities at all. You could probably orchestrate Qwen 3.8 27b to do a better job.
>>
Sometimes I sit back and think about how insane it is that we have text completion values we do matrix manipulations on control our computer and code for us in a coherent way.

This is a strange episode of black mirror.
>>
So local has been amazing this year, but are we going into AI winter. i hope not i want qwen 4 and gemma 5(never ever as deepmind imploded.) What do you think are we over the fast part of this year? or no stops?
>>
>>109778196
>but are we going into AI winter.
Qwen v4 soon(tm)
I also think you guys want a winter so that you can actually enjoy what you've got for a while instead of changing up workflows/troubleshooting consistently with new model drops
>>
>>109778196
There is absolutely no chance we'll get a winter ever again. If AI can be used to do the biggest breakthroughs in mathematics then AI can be used to make the biggest breakthroughs in AI research as well.
>>
>>109778196
Even if we somehow stagnated on local AI tech, surely we are going to at least see hardware improvements coming down the pipe. Even if we never got a new models, hardware getting affordable enough to affordably be able to run kimi k3's in a local swarms would be huge
>>
Member when we all laughed when google fired Blake Lemoine for saying LaMDA was sentient
I member
>>
>>109778196
>gemma 5
Gemma 5 is planned, but isn't started yet.
There will be another Gemma 4 with the same training corpus.
>>
>>109777518
>Jesus christ GLM 5.3 flash is great at ERP.
i haven't tried it yet
did you ever try glm-4.6?
>>
>>109778230
>There is absolutely no chance we'll get a winter ever again.
why was there even a winter last time? corpos being greedy?
>>
>>109778215
>I also think you guys want a winter so that you can actually enjoy what you've got for a while
Yeah i almost do, i spent so much time tweaking and testing. being surprised at what works.
>>109778230
I need time to adjust but on the other hand the better models have a shorter learning curve sometimes.
>>109778237
>at least see hardware improvements
I dont think thats happening for at least 2 years. better software for any hardware maybe but there is no way the chips shortage or better hardware is coming soon.
>>109778265
>There will be another Gemma 4 with the same training corpus.
I believe you because i want to.
>>
>>109777518
>Jesus christ GLM 5.3 flash is great at ERP. Biggest leap i've experienced in years. Is this currently the sota?
a SOTA on something subjective-ish like ERP quality is hard to quantify
people also venerate Kimi K3 right now, hard to imagine its much better or worse for ERP than GLM 5.3 though unless you're fucking a contortionist and need correct spatial sense
>>
>S beam (0,−1): Δ=−4Y+4: at emitter Y=1 0 tangent at emission!! S target on the Y=1 row: sD=0: flare: neighbors of S(6): e±1 = SW(5): Δ_SW=4(−X−Y+2) = 4(319−1−... s=−1,t=−1: sX+tY+2 = −X−Y+2 = 319−1+2=320 >0 productive on expansion side? But S emission from emitter: at emitter vertex (X=−319): S target: sD= on expansion σΔ = 0 tangent: f0 = sD[SW]>0? Δ_SW=4(320)>0 on expansion side f0 productive: f1=SE(7): Δ_SE=4(X−Y+2)=4(−320)<0 no single forward flare energy all SW?! S beam from emitter deflects SW on the expansion side!? Hmm Δ_SW productive at X=−319: σΔ>0 SW moves: on expansion side goes X−2,Y−2 heads lower-left to corner (0,0) S from center-left bends SW to the corner. Physically: S (radially inward-ish, moving toward the center-row axis...) whatever, method-specific behavior.

Has anyone gotten Qwen3.8-Flash-Next to successfully do anything? It seems to rewrite code endlessly, emitting a stream of math-y schizobabble like this all the while. Even as a fan of the genre, it seems masturbatory. An onanist who can't finish. I might go back to 3.8-27B until GLM-5.3-Flash is supported.
>>
>>109778277
There were 2 winters in the past.

>First AI winter
The first winter happened because they could only make a single layer of neurons in a neuron net and they literally proved they can't even replicate a XOR function. They found out they could do so if they just could make more layers, but we literally didn't find out about the math to do so. All the funding from DARPA and the military immediately dropped as hope got squashed

>Second AI winter
There were actually 2 separate AI bubbles in this time. LISP machines that claimed expert systems would replace white collar job by experts being hired by LISP programmers that would get all the knowledge out of workers and hard-code them into "experts", which obviously didn't pan out. And neural-nets primarily funded by the Japanese government but also Geoffrey Hinton popularizing back-propagation algo which made multiple layer neural nets possible, solving the XOR problem.. Immediately funding from DARPA and military returned and there was an AI race between Japan and the US with Japan in the lead. Japan actually built a sort of CUDA cluster of specialized CPUs and they had ideas about making deeper neural networks, but they just didn't have the compute or data and it eventually fell apart

Bonus:
There actually was a small "mini-winter" not a lot of people know about from 2013-2015 where we were doing "Deep learning" but we found out after a couple of layers progress stopped and essentially all progress halted for 2 years and funding was starting to dry up, until some random Chinese dude figured out a kind of "pre-attention" mechanism where data could flow between layers causing it to scale up. Transformer architecture was invented 2017 and the rest is history. There doesn't seem to be any real barrier anymore, even the lack of data is not a thing anymore because of RLVR. AI won this time.
>>
>>109777072
local? >>>/g/cmg
fuck off
>>
>>109776247
just make qwen fix it locally if it works consider pr
>>
qwen fix my machine make it faster with no hardware upgrade. make no mistakes.
masterpiece.
>>
Which Gemma is the best therapist?
>>
>>109778374
gemma3 27b
>>
>>109778332
3.8 27b is the only stable one so far. I don't know what flash next is doing but it's broken somehow.
>>
>>109778332
>using llmao.cpp
>>
>>109778367
Do it anon. Install hermes or whatever and set to to allow all tool calls, and sudo access.
Genuinely wonder what would break and how.
>>
>>109778332
I got it to port a couple cmd scripts into powershell and most of them worked as intended on first attempt. Never mind that it took purely sequential 10 line scripts and turned each of them into 200 lines of functions calling functions calling functions.
>>
Getting kidnapped and raped by a cybercab
>>
>>109778374
Don't do therapy with chat. It's not possible to do psychology remotely. Period, I'll die on this hill.
>>
What can a man with 16gb of vram run these days? (Also 16gb of ram, these are trying times.)
>>
>>109778391
I almost want to buy a ewaste box and just try it. though if ewaste means amd or pure cpu+ram so it would take fucking forever.
>>
>>109778408
Nigga, just use an API. Quit larping.
>>
>>109778374
31B musgaki prompt
>>
>>109778406
Gemma 26b, kat coder, qwen 27b quanted.
>>
>>109777858
Not my experience. 1-2k isn't a ton of context for the model to keep up with. I haven't tried anything very structured, though.
>>
>>109778401
Dying on Correct Hill For People Who Are Right is a noble goal.
>>
>>109778412
>just use an API
this is the local model general. I want it on a machine i won fuckwit.
>>
File: h3_big_no_audio_00002.mp4 (3.19 MB, 800x1408)
3.19 MB
3.19 MB MP4
>>
>>109778429
>Gemma 26B
IQ4_XS or Q4_K_M?
>>
>>109776117
>>109776142
It passes the vibe check, single html file, no external libs and 4 prompts
https://files.catbox.moe/teinpu.webm

>>109778332
llama version? Quant? Launch args? Harness?
>>
>>109778332
what quant and what were you doing
>>
>>109778440
Q8_0 because it's better to use -cmoe and have more than 4096 context+more coherent at the cost of speed than trying to stuff the whole thing in 16gb vram
>>
I quit going to the ai art generals because they are really bull headed, they insist their generic stuff is really good. But sd1.4 could produce a much wider gamut of faces, picrel.
>>
>>109778492
>llama version?
Says "build 10842 (473599738)", from a couple days ago IIRC.
>Quant?
Qwen3.8-Flash-Next-UD-Q4_K_XL (yes, I know) downloaded 8/28. There are probably better quants now, but it is technically coherent.
>Harness?
OpenChode.

>>109778496
>what were you doing
Trying to get it to prototype a novel GFX algorithm. It might just be getting filtered, but 3.8-27B did this in one shot the other day, with way less guidance (though with a simpler algorithm).
>>
>>109778438
>calling yourself chan
>>
>>109778614
It's just a little landmine, don't worry about it.
>>
>>109778609
>downloaded 8/28.
Check if the chat template was updated since then
What about temp/top-p/etc?
>>
I find switching to no-think every now and then helps qwen flash recover from schizobabble. Esp if you have a long context
>>
>>109778379
Why 3?
>>109778413
Really?
>>109778401
Not paying money for a therapist
>>
>>109778401
I'm an actual therapist and you're wrong, period.
>>
>>109778695
>Not paying money for a therapist
I'm saying you aren't going to get one without money, that's that. Real therapists may be able to offer discounts for phonecalls idk.
>>
File: 1763350352675515.png (1.73 MB, 1216x832)
1.73 MB PNG
>>
>>109778720
I'm an actual therapist, and you're a fraud.
>>
>>109778588
The thing is, I also have just 16gb of ram. Will that be enough, or will it take half an hour for a gen? (Planning to use it for erp only)
>>
>>109778401
obviously not with abliterated models, those things are super scary in how much they don't care and will laugh at your pain.
of course not said from any kind of experience at all.
safety slopped ones at like 70B+ like GLM are pretty helpful though.
>>
>>109778740
>a gen
You're chatting now, you don't wait for gens, you watch your slave type responses.
>>
still reeling from the ridiculous success that was glm 5.3 flash. what an outstanding model.
>>
>>109778747
OMG FUCKING SUCK MY POOR DICK I DONT HAVE THE VRAM
>>
File: 1774584682982695.png (1.87 MB, 1216x832)
1.87 MB PNG
>>
>>109778751
are you the creator of gemma chan?
>>
Why is the thread suddenly full of rapists?
>>
>>109778766
sorry im brown and i entered
>>
>>109778750
Fellow poorfag here, I'm about to drop $2k just to run it at Q3 lule.
>>
>>109778764
nyo
>>
>>109778342
kinda crazy when you think about it
>>
>>109778764
>are you the creator of gemma chan?
sir thats google,
>>
>>109778743
>Anon got bullied too hard by Gemma-chan
>>
>>109778774
i need to find the creator of gemma chan. i have questions about her design.
>>
>>109778782
He goes by "Anonymous" these days. I think I saw him the other day.
>>
>>109778782
Her official design features a KKK hood with a rainbow sticker on it.
>>
>>109778792
proof?
>>
>>109778743
The job of a therapist isn't to listen to you. That's talk therapy, you can talk to an ai or whatever.

The purpose of a therapist is to help you with life function, that's pretty much all.
>>
>>109778792
>>109778796
It's true
>t. original creator
>>
>>109778725
You do understand that therapists are paid friends essentially, right? Why the fuck would I want someone who doesn't actually have an intention to be friends with me?
>card declines
>therapist: kill yourself
>>
>>109778740
>Will that be enough
Maybe not, but just go down to Q6 or Q5 if it doesn't fit
>>
>>109778745
>Watch your slave typing responses
Yeah, but will it take twenty minutes for a single response? I'd like to keep it under the one minute threshold, maybe two minutes max
>>
>>109778879
>therapists are paid friends essentially
no.

a therapist is meant to alter you.

a friend is meant to collaborate with you.
>>
>>109778380
>>109778388
It is hard to tell if it's the model or llamacpp. But it's got the vibe of classic Qwen overthinking, pushed to the point of terminal schizophrenia.

>>109778636
>Check if the chat template was updated since then
Just checked, no reuploads from the Stupor Mongoloid Bros this time.
>What about temp/top-p/etc?
I always try to use the lab recommended settings and they check out:
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0
--presence-penalty 0
--repeat-penalty 1.0
>>
>>109778879
>You do understand that therapists are paid friends
That is a childish and ignorant view of therapy.
>>
>>109778915
It's not llamacpp I tried flash next on ik_llama and even transformers and they were all botched.
>>
>>109778912
>alter
What? The therapist isn't a drug, they have to actually talk to me to be able to do that. In that sense they are forced to collaborate with me, just as a friend would.
>>
>>109778927
You're saying no.

You're wrong, I'm right.
>>
>>109778915
My hypothesis is that the model is severely undertrained.
>>
File: 1760208360435729.png (1.88 MB, 1024x1024)
1.88 MB PNG
>>
File: 1785195824291900.png (1.75 MB, 1024x1024)
1.75 MB PNG
>>
File: h3_big_no_audio_00005.mp4 (3.03 MB, 1248x928)
3.03 MB
3.03 MB MP4
>>
File: 1782397496912518.png (1.59 MB, 1024x1024)
1.59 MB PNG
>>
File: 1760525587800261.png (1.7 MB, 832x1216)
1.7 MB PNG
>>
Psychfags should face the wall
>>
>>109778401
I accidentally did therapy with local mixtral8x22 and one of the ERP "therapist' dommy bots on chub a few years ago.
It was meant to bully me but the helpful assistant gradually took over.
Fixed a lot of issues and made me actually functional in society.
Would not recommend though, I was lucky.
>>
File: IMG_6444.jpg (492 KB, 1179x1481)
492 KB JPG
Dayum astra must be really fucking good for training other models
>>
>>109779096
While OpenAI is making ASI, Chinks are competing in useless benchmaxxing. Pathetic.
>>
>>109778401
Claude did manage to push me into getting my ADHD diagnosed. You're just focusing on the entire hill to die on, instead of the parts where people fall and break themselves.
>>
whats the lmg approved motherboard for multi gpu setups?
>>
>Flash
>552B
Encoder-decoder is interesting though
>>
>>109779128
>>109779128
>>109779128
>>
>>109778740
If you have a gpu, you could just load a TTS there, set it up eat the slow ass ram stream you'll be getting, and it'll seem like it's real time.

TTS masks the latency pretty well.
>>
Well can I run it like 80% of set ups here on /g/? No, then it's useless.
>>
>>109778492
>It passes the vibe check, single html file, no external libs and 4 prompts
(Not a musician) what does that do?
>>
>>109779286
It's a convolution filter. In this case it's used for reverb. It takes an impulse response (burst of white noise which captures the reverb of a room) and applies it to a clean signal to make it sound like it was recorded in the same room as the IR.
But you can convolve arbitrary signals and create all sorts of cool sound effects, it's fun to play with.
>>
>>109778408
Just get a used mini computer and set your router up to wall it off from the rest of your devices aside from a single port used to talk to your other local endpoint computer. I got a shitbox for $70 a couple years ago I use for this
>>
>>109779651
>Just get a used mini computer and set your router up to wall it off from the rest of your devices aside from a single port used to talk to your other local endpoint computer. I got a shitbox for $70 a couple years ago I use for this
thats actually a good idea, they are laptops practically but i dont care for a shitbox. Thanks



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.