AI
3 sources
3 sources
10 non lus sur 10
r/LocalLLaMARetirer le filtresubmitted by /u/Niceyyc [link] [comments]
Planning to build a PC mainly for local LLMs/coding agents. I keep seeing 3090 + Qwen 3.8 Flash Next benchmarks, but could not find enough info for the 7900 XTX 24GB. 3090s are hard to find where I live, while newer Nvidia GPUs are too expe
Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models. For the open models, GLM-5.3 is in a leag
https://huggingface.co/security.txt submitted by /u/Nunki08 [link] [comments]
submitted by /u/jinnyjuice [link] [comments]
I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations. Introducing Cherenkov, an inference engine for Apple Silicon combining predictive e
After noticing that it is ranked among MUCH larger frontier models in the EQ-Bench Creative Writing benchmark and the Hemingway-bench, I decided to give it a try and was very impressed. I didn't do very formal testing, but I did ask it to e
A few years ago a 100GB was considered a very large language model. What do we call under 100GB models now? Tiny models? haha submitted by /u/Terminator857 [link] [comments]
Github Repo. Blog post. 💡 TL;DR (from the Github Readme) Spend less without making the agent do less useful work. SoL-Pi is a standalone extension for Pi that packages four reusable efficiency mechanisms discovered through scaled auto-rese
Title says it all. https://github.com/gjabdelnoor/Day1DeepseekV4.1-CPU The goal is pretty simple, I like having infinite slow tokens from the bioinformatics machine in the lab to run overnight or over-week agentic jobs, paired with a watche