注册并分享邀请链接,可获得视频播放与邀请奖励。

Unsloth AI 的个人资料封面
Unsloth AI 的头像

Unsloth AI (@UnslothAI)

@UnslothAI
0 正在关注    0 粉丝
DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️ DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change. DeepSeek-V4-Flash-0731 can reach at 120 tokens/s. GGUFs: Guide:
显示更多
0
32
494
47
转发到社区
DeepSeek V4 Flash 0731 can now be run locally! 🐳 Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM. V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today. Guide: GGUF:
显示更多
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs:
显示更多
0
143
3.2K
402
转发到社区
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5 We gave 3 models the same prompt and compared one-shot outputs. The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s. Which output do you like best? GGUF:
显示更多
0
164
3.4K
379
转发到社区
GLM-5.2 can now be run locally!🔥 The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Guide: GGUF:
显示更多
0
263
7.1K
844
转发到社区
Qwen3.6 now runs 2x faster with MTP GGUFs! Run locally on just 18GB RAM. ⚡️ MTP enables Qwen3.6 to generate ~1.4–2.2× faster with no accuracy change. Qwen3.6-27B MTP runs at 160 tokens/s. 35B-A3B reaches 240 t/s. GGUFs: Guide:
显示更多
0
115
2.1K
248
转发到社区