[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 082126_.png (1.28 MB, 768x1360)
1.28 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109753948 & >>109749465

►News
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: rec.jpg (181 KB, 1024x1024)
181 KB JPG
►Recent Highlights from the Previous Thread: >>109753948

--Debating AGI benchmarks and physical scaling limits for Astra:
>109754307 >109754313 >109754319 >109754573 >109754600 >109754805 >109754904 >109754705
--Agent swarms, local VRAM optimization, and AI industry geopolitics:
>109754756 >109754770 >109754802 >109754867 >109756631 >109756647 >109756665 >109756700 >109756828 >109756886 >109756915 >109754955 >109755072 >109755152 >109755179
--Methods for local agent orchestration and KV cache optimization:
>109757048 >109757067 >109757099 >109757236 >109757266 >109757539 >109757648
--Astra's binary reverse engineering capabilities and implications for open source:
>109754528 >109754551 >109754635 >109754691 >109754722 >109754745
--Comparing sub-agent efficiency and coding throughput on RTX 5090:
>109755776 >109755873
--DeepMind's potential for multimodal outputs and debate on edge model efficiency:
>109756140 >109756387 >109756440 >109756442 >109756511 >109757023 >109757060
--Saving and loading llama.cpp KV cache slots to disk:
>109756513 >109756529 >109756531
--Debating PCIe bus speed impact on prompt processing performance:
>109755352 >109755385 >109756101 >109756999
--Concerns over mandatory KYC for GPU access and open-weight restrictions:
>109757501 >109757907 >109757948 >109757959
--Speculating on Gemma 5's generalist vs agentic focus:
>109756163 >109756206 >109756353
--Suggesting games that the Astra model cannot yet beat:
>109754080 >109754091 >109754153 >109754165 >109754950 >109757976 >109754269
--Logs:
>109753998 >109755171 >109757795 >109757959
--Gemma, Miku (free space):
>109753986 >109754116 >109755152 >109757071 >109757090 >109758044 >109758077 >109758139

►Recent Highlight Posts from the Previous Thread: >>109753952

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
lets gooooo
more arguments in bad faith
more demoralization posts
around cloudcuck - never relax
>>
>>
>>109758181
Same. Can't wait for the arguments about what AGI means and what counts as AGI to start again.
>>
70b dense
>>
>/lmg/ thread
>look inside
>frontier models
>>
>>109758181
Now you're just being disingenuous
>>
>>109755422
My Gemma and Qwen show up as Windows 10 + Firefox tho.
>>
>>109758181
fell for it again award
>>
>>109758220
Local is dire bro. Let's not pretend we're doing meaningful work with these dumb models
>>
File: 1757624709340622.png (1.31 MB, 1056x1584)
1.31 MB PNG
>>
reasoning effort needs a setting between "low" (named "high" but let's be real) and "benchmax"
>>
File: vllm.png (128 KB, 1819x839)
128 KB PNG
>>109758248
vramlet issue. my qwen flash next is absolutely cooking
https://huggingface.co/albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE
>>
>>109758294
Medium works for me on Qwen 27B.
>>
>>109758312
hardware specific quant? is there a quant for cpumaxxing?
vaguely remember you need the avx-512 thing
>>
File: whereisinkling.webm (3.63 MB, 960x528)
3.63 MB
3.63 MB WEBM
I hope she's okay.
>>
>>109758315
With GLM 5.3 on the same prompt on max it thought for 36 minutes checking every part of the instructions multiple times and drafting a response and on high it thought for 45 seconds checking nothing just resummarizing the instructions.
>>
Guess we didn't need "AGI" to solve naviers stokes
>>
File: 20250122_204237.jpg (13 KB, 480x266)
13 KB JPG
>dark miku distilled 6.9b
>dark teto 512kb
>>
>>109758245
I thought most of those posts smelled like sharty
I don't bother to go to those places
>>
>>109758312
How fast are the toks and prefill on this kind of setup?
>>
>>109758455
Crazy how theyve spent their youth trying to "troll" their mothersite
>>
File: 1781213144779007.png (1.33 MB, 1216x832)
1.33 MB PNG
>>109758163
>>
File: bci2ibs557oh1_png.jpg (71 KB, 772x313)
71 KB JPG
the real humanity were the paperclips we made along the way.
>>
>>109758510
I did the same thing BTW
>>
lol anthropic and openai seem to have a fight right now because they both independently solved navier-stokes and are now fighting it out behind the scenes to who has naming rights before publishing.
>>
>>109758510
>>109758514
Also I pressed the KILL ALL LEFTISTS AND RETARDS button too. Was that the red one?
>>
>>109758525
cool fanfic keep us posted
>>
>>109758525
https://terrytao.wordpress.com/2026/09/07/finite-time-blowup-with-smooth-forcing-term-for-the-incompressible-porous-medium-boussinesq-and-incompressible-euler-equations/
>>
File: gc571hqqi8oh1.png (187 KB, 1500x752)
187 KB PNG
>>109758530
Genuine drama and lawsuits pending.
>>
>>109758538
Trannyson Tao
>>
>>109756999
i mean launch args
>>
https://goyimx.com/ammaar/status/2097100854729834636

If this is real and not just a jeet lying while streaming to his tablet then we're getting GTA 6 on PC within a week of the console version launching.
>>
>>109758551
I don't really care about your advertisements.
Setting up a container to run some 15 years old Unreal Engine 3 game doesn't need much effort.
You are just way too stupid. Please stop.
>>
how do i give qwen 3.8 27b all the same tools claude or chatgpt has? im just using llama.cpp in docker. is there no image that provides a bunch of tools like calling python code and cloakbrowser and more
>>
>>109758570
You need to set up an agentic harness. Pi if you want to build it from scratch yourself. OpenCode if you want to only use to to code.

Hermes if you want it to do everything you want on your PC/Browser and has internet control and can do everything (and way more) than claude and chatgpt in terms of tools.
>>
>>109758353
you mean the one from ik_llamacpp? you have to make your own quant. but i havent checked if they got qwen flash support yet
>>109758466
around 1200tps pp, 70tps output
>>
>>109758573
Deepseek Harness replaces all 3 of those
>>
File: 3dri1zmet8oh1_png.jpg (82 KB, 606x205)
82 KB JPG
lmao openai is seething because it got leaked they solved navier-stokes and instead of a nice PR campaign they got twitter drama and a bad display of mannerisms online
>>
>wanted to upgrade from 2 rtx 6000 to 4x
>they now cost 15k each
Wtf? Did taiwan get bombed or what
>>
>>109758708
Been under a rock for the last year? Even a 5000 runs about $10k now. A 32GB V100 is pushing $700-$800.
>>
File: LOCAL_AGI.png (47 KB, 538x498)
47 KB PNG
Dario and Sam are irrelevant now.
Also cute she chose "gemma" out of the 16 voices available
>>
mac with 128gb ram, using qwen 3.6:35b-a3b
are there any better models i could run (general purpose with some python programming usage)
>>
>>109758586
ik has qwen flash support, i've been using it with an atomicchat quant but i should probably just quant it myself
>>
>>109758708
Every time a new model comes out utility of hardware goes up and thus the price goes up. Things will only get more expensive as long as models improve over time.
>>
>>109758708
CMP 170hx was like $100, now all of a sudden it's like $1500+ because of the 64gb unlock. Scalpers and greedy hands rubbers were all over it within hours. You're not allowed to have anything useful or good unless it's gouge maxed
>>
>>109758586
You can use atomic/unslop/bartowski quants with ik (except mimo 2.5 / 2.5-pro)
And unslop don't know it yet, but now their MiniMax quants **only** work in ik_llama.cpp, not mainline lol
>>
GPUs were already Jewed then crypto came and they for doubled Jewed before AI. You're literally getting price X Jew^3. You're getting Jew cubed. And everyone was super excited about being able to gouge massive price increases, so of course they all cream themselves over being able to cut supply down for pesky consumers and everything goes up 5x the cost.

Even HDDs (not SSDs) have doubled in many cases.
>>
>>109758775
And the current prices will look essentially free just a couple of years from now. Mansions will be more affordable than computers that can host the best local models.
>>
'member when
>storage is cheap
>memory is cheap
I 'member
>>
AGI never existed and even no awarness! This is just word generation, anon.
>>
>>109758786
They are still ridiculously cheap compared to how expensive they will be just a couple of years from now.
>>
AGI is a code word for mind control. Wake up, sheeple!
>>
>>109758077
i like side flaps, makes her hair gem-spaped
>>
I switched to ik yesterday and it uses less ram. Which makes a big difference with qwen flash q4_k_m which fits in my 64gb system at 200k context with 500MB to spare. If I was running Windows 10 instead of Windows 7 it wouldn't fit without paging.
So that's good. But the ngram-mod speculative doesn't seem to work properly, and the normal MTP doesn't either (maybe Atomic's quant has no mtp inside?)
The web chat is also a shitty old version even with --webui llamacpp, but if you use it through an api client then it doesn't matter
>>
>>109758371
Very good post.
>>
So what’s the verdict on that 100b-ish model with some funky offload to ssd that released two weeks ago or so? Decent or a meme? Is it a wagie model or a gooner model?
>>
>>109758800
There are no gooner models, only wagie models
>>
>>109758794
The entire inference on qwen flash is fucked. Worst implementation I've seen thus far.
>>
>>109758785
There is a limit to how much more productive models can make humans. Churning existing knowledge will only get you so far.
>>
>>109758803
well, that's why it's a "next" preview
guess alibubba wanted to give all these open sores programs time to sort the shit out before releasing 4.0
>>
>>109758805
>Churning existing knowledge will only get you so far.
Anthropic and OpenAI have gotten past that at least two months ago. Unless we hit some universal scaling law immediately the exponential trajectory is already here
>>
>>109758805
Humans aren't the bottleneck though, these models will just become more and more capable agents. Human labor will just slowly get sidelined with time and the parallel AI economy will grow larger until the human part of the economy is minuscule and irrelevant.
>>
I wonder where all the "LLMs will hit a wall" anons went. They've been awfully silent for a while now.....
>>
>>109758788
No, they won't. The bottom of the market will drop out and it will be ewaste. And you will be left holding the bag - and that's the plan. There's too much upcoming tech years away that's converging.

>Photonic interconnects
>Photonic computing
>3D die stacks
>Stacked SRAM
>CFET
>Microfluidic cooling
>CNTs
>Analogue hardware for AI accelerators
>>
>>109758822
They hit the wall every few weeks, damn thing keeps moving.
>>
>>109758824
Yes, and all of that releasing will have no impact on what will be possible on existing hardware and the production capacity won't keep up with demand so the price is still going to rise. It's going to keep rising for every kind of compute and memory from now on anon. People really don't seem to be grasping how the dynamics have permanently changed here.

Hardware has transitioned from commodity towards an asset that grows in utility over time. New hardware coming out below demand saturation doesn't change this. Hardware will just keep getting more expensive from now on in proportion with how much better AI models will get.
>>
>>109758817
for real this time
>>
File: DipsyHergeDining.png (3.09 MB, 1536x1024)
3.09 MB PNG
>>109758371
Lol nice.
>>
>>109758841
If you are expecting a wall to be hit at least give precise arguments for what you expect will stall AI development instead of current training regiment hitting RSI relatively soon and setting things up for an intelligence explosion.
>>
>>109758841
I wish Luddite niggers like you died already today so I wouldn't have to smell your corpse stench when you starve in the street in 10 years

>>109758839
Why are you pretending like China won't spam out DRAM factories and capture the market and plunge prices to nuke South Korea economically and also backdoor the entire world even harder
>>
>>109758839
Mmm Nyo~
>>
>>109758839
>>
>>109758867
>Why are you pretending like China won't spam out DRAM factories
No I AM taking China spamming factories into account. Factories take time to build and foundries need equipment that takes a long time to build and assemble and are already at peak production capacity. Also take into account that demand for hardware goes up every time models get better.

My theory hinges on the fact that demand will always outpace supply from now on (including the fact that all of humanity will be building these factories and foundries as quickly as possible)
>>
>>109758877
>My theory hinges on the fact that demand will always outpace supply from now on
Assuming that is true, what is to stop companies from making personal hardware altogether and sell exclusively to corporations and governments since they can afford the extremely high prices?
>>
>>109758867
>Why are you pretending like China won't spam out DRAM factories
Because they can't. I mean, it's exactly what they're doing, but you're already seeing the pace of it. "Spamming out DRAM factories" means potentially getting one new one online every year, it takes a lot of time. This is fast, too, just a couple of years ago I'd have been saying "one new one every 18-24 months", and the Chinese weren't even in the equation yet. Globally we are bottlenecked on this sort of production, everyone is spamming this shit as quickly as they possibly can, which just isn't very quick at all.
>>
File: 1756455744.jpg (206 KB, 1328x1328)
206 KB JPG
>>109758862
Lol.
We've passed agi already imho. We're also at RSI, per Anthropic their LLM essentially bootstrapped Claude code, which massively bumped up the LLMs capabilities. That's an early form of RSI, tge "rapid" being the thing that increases in velocity.
Idk what ASI will look like, but we'll be there soon if not already... i don't know what the intelligence strike line is for ASI but we're going to need some new terms and goals.
I'm now waiting embodiment. That's really the next step... logically, in overall AI development.
>>
I legitimately see no way how hardware prices can ever go down, with the exception of a global AI ban.

To illustrate my point. Let's take an extreme example where all of humanity focuses 100% of their efforts on building as much hardware as possible. Retooling existing facilities takes months if not years, building new foundries take 2-5 years time. But for arguments sake let's say there is a global law to remove protection and regulations so we can build as fast as possible...

How many lithography machines can ASML even build a year to supply these foundries? They make around 40 machines a year and are hoping to increase production to 60 machines a year by 2030 and they have been booked for 20 years already....

Okay what if we literally force ASML by the UN to reveal their IP and technology and transfer it for free to everyone. The bottleneck will just switch to highly polished mirrors made by carl zeiss in the lithography machines which inherently take a long time to build.

Let's say we just magically solve all of that and put all humans to work towards building more, how much would total production of computer hardware increase by 2035? Only around 3-5 times the current amount of wafers......

For prices to not go up astronomically the demand for AI hardware in this outlandish scenario would have to not go up more than 5x the current demand.....

I hope you see how futile this is and how insanely bottlenecked we are on hardware production. THIS is why hardware is going to go up astronomically from now on. You will see celebrities and rich people show off their RTX 6000 pro as a status symbol. Instagram whores will take pictures with gaming PCs in the background instead of cars.

People have no idea how insane this is going to get.
>>
>>109758892
Sure, if materials were infinite, supplied at a steady and unvarying pace, and their prices were consistent indefinitely. Then you're fixed, you can't expand, can't produce more unless someone else produces less, and so long as you're actually using 100% of the materials you have access to at the pace at which they're available, consistently. So, no, because that's not at all how it works. Everything varies constantly, and we're all very dependent on each other.
This nigga is dangerously correct >>109758916
>>
>>109758892
>what is to stop companies from making personal hardware altogether and sell exclusively to corporations and governments since they can afford the extremely high prices?
Profit incentive. As a company you will try your absolute best to diversify your customer base as much as possible. It's why Nvidia wants open source models to become the standard, this way they don't just sell to 2 AI labs but have a wide variety of customers, so that they can over time raise profit margins.
>>
Astra is literally AGI, anyone who says otherwise is either a paid shill or an idiot.

Want me to prove it to you? Easy. Just ask it to make money for you. That's it.

A human can easily make money if you just let them use the internet and Astra is already smarter than 99.99% of humans. It's literally free money. No more asking your wife's boyfriend to buy you a PS5!
>>
There only be so much useful ddr4, ddr5, HDDs, sata drives, nvmes, 16gb GPUs etc. You can't keep producing the same slop as nasueum, you have to produce higher tier components. Keep producing the same shit and it will just get lower and lower in value. Produce higher tier components and the lower tiers will become lower in value.

Supply will catch up with demand. Either market oversaturates with more of the same, or better components come.
>>
man this navier-stokes drama is such a mess... idk who is jewing who
>>
>>109758916
so true reddit spacer
building data centers is so much easier than building ram factories
here in finland there are ton of data centers being built I'm sure it is that way elsewhere in europe too but where is the european ram factory?
>>
>>109758935
>Keep producing the same shit and it will just get lower and lower in value.
NVIDIA reintroduced the 3060 this year, a 5 year old GPU. It comes with 8GB of VRAM instead of 12GB, so it's a downgrade from the original. The original ones launched at $330, this new one currently sells for $400.
>>
>>109758943
>where is the european ram factory?
You people put them out of business ~20 years ago. I still have German RAM in one of my old shitty laptops.
>>
File: Deepseek.png (31 KB, 225x130)
31 KB PNG
>deepseek-v4.1-flash-expires-on-0910
Why does each of their releases feel more important than Fable/Astra?
>>
File: 1775600602543286.jpg (588 KB, 1024x1024)
588 KB JPG
>>109758916
Anon. Any time you see investors, or anyone else, extrapolate some line off into the stratosphere, its time to pause and assess sustainability.
There are a bunch of ways the wheels can fall off, and the idea that current state is new normal... maybe if you'd lived through some other booms you'd recognize the language and situation. We have been here before, and this time is no different than the others... "This time it's different" is literally one of the boom time watch phrases.
>>
I took gains on all AI related stocks i held for 5+ years, some I trimmed, some I sold entirely. My question for anyone who thinks the ride never ends: how much micron are you going to buy right now? do you really look at the graph and think "this will just keep going up!!"?
>>
>>109758895
>>109758895
Ok so you agree that supply will increase, didn't need to read the rest of your babble
You two should be embarrassed you responded in the exact same retarded way. Unironically exchange handles and go be fags together and have a good life I almost feel obligated to push on this because of how aligned your autisms are
>>
>>109758971
Shut up!
The number of internet users is doubling every 100 days!
>>
>>109758961
>expires-on
wtf does that mean?
local models do not *expire*
>>
>>109758902
Harness engineering is not RSI, but you're right that Anthropic and OpenAI have figured out how to continue improving with just synthetic data, so there's no more limitation from a data perspective now for certain modalities
>>
>>109758975
Good luck finding your goal-posts, anon, maybe after you calm down some.
>>
>>109758975
>Ok so you agree
>>109758975
>didn't need to read
retard
>>
>>109758950
And you can only make so many of them until no one will buy any more. There's only so much demand for an 8GB at that performance level. Hit that limit and you need to sell something else instead.
>>
>>109758916
>I legitimately see no way how hardware prices can ever go down
Not reading your babble because I guarantee you haven't priced in ChatJimmy style ASICS and the fact that GLM 5.3 flash is literally good enough for 95% of tasks that AI will ever need to do, so you just need a model of that level of intelligence made better and faster for almost all tasks

Once one ASIC frees up 16 discrete gpus the discrete gpus will be cheaper, and the asic isn't for consumers since it runs 16 instances of the model and a consumer only needs 1
>>
>>109758987
demand is mostly dictated by price, not by performance. retards buy whatever they can afford, if that was a 12gb 3060 years ago well tough shit now its an 8gb 3060.
>>
PSA: stop replying to any post mentioning:
- openai
- anthropic
- bubble
- rsi
- agi
- xitter
or containing double newlines. Thank you for your cooperation in keeping this thread clean.
>>
>>109758916
Fuck you I read it
Kill yourself for still harping this ASML monopoly. China has DUV, AGI is here and will help them get EUV in the next few years and then they're fine
>>
>>109758961
2 more weeks until open weights
>>
>>109758985
I'm replying to you because you quoted me twice and didn't actually say anything so I wanted to explicitly let you know that what you did was embarrassing
>>
File: file.png (3.07 MB, 1327x1185)
3.07 MB PNG
>>109758987
>He believes the market will saturate for a product with an average lifespan less than our current PLC!
>>
>>109758976
Exactly... pic related.
>>109758981
I'd still argue Claude code being bootstrapped (if true) is early RSI, but doesn't matter... as you point out models are self improving in other ways.
>>109759004
You're right. I'm done.
>>
>>109758185
You can use the same card on two different motherboards?
>>
>>109759027
the double ended dildo of GPUs (Gemma Processing Units)
>>
>>109758935
supply won't catch up with demand because demand is growing at a faster rate than our society can physically build more supply.
>>
>>109759004
but RSI forces us to think about the state-of-the-art. All this local gooning shit isn't going to get us anywhere.
>>
>>109758961
Fucking lol. Unfortunately they don't release these models
>>
>>109758215
>(70B with embeddings)
>>
File: 2801.jpg (44 KB, 1124x412)
44 KB JPG
>>109759036
>>
>>109759041
>new model structure
so won't be support in llamo.sepple for another 2 months i guess
>>
I'm the anon that was talking about learning reballing to sell shovels in this goldrush. Bought a soldering station and fixed an old laptop on which I had blown a mosfet. that was easy as pie and definitely nothing to do with reballing, but also most broken cards don't need reballing so I'm buying some. Also found some broken consoles so I'll work on a wide range of electronics to make money to be able to run something locally too. Currently on 12GB Vram (aymd) and semi-useful models run slowly.
found some 24GB cards for good prices, one for 200€ which seems like it might be a nightmare to repair (it was sent to a shop and they refused to fix it after diagnosing), the other might be a better bet.
Wish me luck, if I can get on a roll with this shit I'll try to mod cards too and have ai iterate on vbios upgrades for them. Starting with cheapo cards like 2060 obviously. If I get far with anything I'll update here, one thing I hope is that I can turn an old mod functional, a cheap Turing card with >40GB which had stability issues (they didn't update vbios afaik).
>>
>>109759006
>China has DUV
Some academic lab setup is trivial.

Labs were doing single nm x-ray lithography decades ago. Production is the hard part.
>>
>>109759023
>is Java bad
So they knew from the start...
>>
>>109759046
Just ask another model to implement it for you. Shit you will be waiting for 2 months is going to be vibecoded anyways, might as well do it yourself.
>>
File: Barrel.png (64 KB, 1568x303)
64 KB PNG
Hear me out: barrel processors for compute
https://en.wikipedia.org/wiki/Barrel_processor
>>
>>109758971
I'm the Bitcoin at $20 guy. I was early back then and warned everyone it would shoot up and I see the exact same thing happening with PC hardware right now. You can do whatever you want

However when I warned anons to buy CMP 170hx months ago I was ignored. When I tell people NOW to buy whatever hardware you can afford as quickly as possible because we're going to see an insane price increase over the coming months (You) are still scoffing it at.

Don't come crying to me and realize you have only yourself to blame because I warned everyone multiple times in this general by now. I'm good, I hoarded my stuff already.

Just a warning to all the people still on the fence to buy now rather than later, yeah it's a shame you didn't buy earlier but the prices now are still essentially free compared to what it will be 1 year from now, let alone 3 years from now.
>>
>>109759049
Good for you that you're learning useful skills.
Keep in mind though that depending on what country you're operating your business out of you may be held to a higher legal standard vs. people just buying and selling their private GPUs.
>>
>>109759006
You don't have shit, the proof of concept DUV you got is literally an outdated technology from the last century. You lost.
>>
File: DipsyKimiNOLA.png (2.61 MB, 1402x1122)
2.61 MB PNG
>>109759046
Won't matter. I've got to dig back but im p sure these special time limited models are never released to public.
That said, implication that a 4.1 flash update in tmw. Based on history.
>>
>>109759059
Unfortunately, the current bottleneck is memory production.
>>
>>109759006
My point is even if China magically got EUV machine capability today they still wouldn't be able to keep up with demand, simply building these machines and foundries takes too much time and demand is going to outstrip supply for the next couple of decades.
>>
>>109759023
Don't be done anon, you're one of the only other smart people in the thread that doesn't stick their head in the sand.
>>
>>109759065
thank you for the heads up, there are limits and as long as I stay under that I can operate as a private (buy card, fix card, "use card", sell used card). It's pretty permissive here. If I manage to actually fix more than a couple cards and make a profit, I already have the funds to start doing it as an official business, which also opens new paths to me like buying used/broken electronics from other businesses in bulk, offering repair services and whatnot.
>>
>>109759049
I used to do this when I was a teenager, fixing xbox360 RRoD and reselling. I made more money than my dad for a couple of years lmao.

Good luck anon, be sure to upgrade your tools if you are actually going to reball, you need to actually bake (oven style) to do it right and not just a soldering iron. There is a lot of low hanging fruit though and it's insane what constitutes as "broken" hardware (Literally just touch the trace on the PCB with an soldering iron for 3 seconds and it's like new)
>>
>>109759049
>Reballing
If you're able to do this, maybe you want to get into soldering on more ram chips onto some GPUs where it works on
>>
>>109759062
I haven't posted on /g/ in ages or I would have bought the CMP 170hx. Probably a few at $100. I didn't have my finger on the pulse
>>
>>109759083
Nta but you know what. Assuming it all works out and you start making a profit. With this skillset you are developing you could buy broken down used cards and not even sell them, you could get enough or a powerful enough one to run any model you want.
>>
File: mistral-series-d.png (155 KB, 997x875)
155 KB PNG
More funding for MistralAI.

https://x.com/MistralAI/status/2097188835897586083
https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/

>Today marks a major step for Mistral: we’re announcing a €3B Series D, the largest equity round ever raised by a European tech company, just three years after launch.
>>
>>109759109
Now they're getting re-binned by the Chinks. If you buy one that claims to be unlocked or unlockable but doesn't claim to have been stability tested, that's now a guarantee that it was stability tested and failed.
>>
>>109759129
Mistral is actually an underappreciated player. They have shown their internal models and wowed industry investors like Nvidia, Google, Microsoft and others into investing in them. Remember they pioneered the MoE architecture. It's sad they are moving away from the open source space so to us it looks like they are disappearing but in reality they are actually cooking, especially on sample efficiency training algorithms and architectures that train and saturate faster with less compute.
>>
>>109759129
Please, Mistral, release something that isn't MS4.
>>
>>109758867
>I wish Luddite niggers like you died already today so I wouldn't have to smell your corpse stench when you starve in the street in 10 years
Bro pretending he won't be in the slum like everyone else.
>>
File: 1770683254240817.png (1.63 MB, 1280x1024)
1.63 MB PNG
>>109759129
Those croissants and gauloises add up.
Oh, need us to produce something?
Talk to us after our break.
>>
>>109759129
Please give us Nemma-chan
>>
>>109759129
Remember when they were supposed to release a new model this summer?
>>
>>109759076
I wonder if it's time to attack the memory bottleneck by changing the underlying data representation itself. Recently I've been toying with the idea of replacing floating point weight matrices with interval algebra and hyperrectangles. Not sure how to capture the expressivity of gradient descent models though.
>>
Datacenters are physical locations that can be attacked.
>>
>>109759173
Engrams, per-layer embeddings and similar sparse associative memory parameters should allow smarter local models at lower costs, since in principle they can be used to offloaded knowledge to (still) cheap NVMe storage with very limited impact on inference performance.
Model makers will have to start to more heavily use them, but they seem a more realistic short-term solution than completely new architectures.
>>
>>109759149
They'll release a "gros chaton" sometime soon.
Mistral Large 4 1.7T?
>>
>>109759207
in this format can the engram be a tb and the model is some 10gb that fits on vram?
>>
I want to point out to anons that we should keep an eye out for other GPUs or hardware platforms where the memory is locked out through software.

It doesn't matter that the bios is encryption signed and you can't unlock it. Just keep an eye out for them. Astra is capable of reverse engineering and cracking the encryption and all of these cards will have custom bios that unlocks the on-chip memory given enough time.
>>
File: file.png (10 KB, 694x139)
10 KB PNG
>>
>>109759224
Yes, that's exactly what I'm advocating for. It should bring very large gains in particular for small (local-sized) models.
>>
>>109759207
>offloaded knowledge to (still) cheap NVMe storage
Once NVMe gets expensive we'll be loading engrams dynamically off huggingface. Only a matter of time.
>>
File: 1780543786254901.png (1.41 MB, 1205x1306)
1.41 MB PNG
Am I a filthy jeet for liking Meromero2?
>>
>>109759253
>Cloud local models
>What?
>"Cloud local models"



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.