注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Gemma」相关的搜索结果

Gemma 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Gemma 的内容
Introducing OUI-1: the first open-weights model for Generative UI 71.7% on Generative UI Bench at 4B params. Beats Gemma 4 31B with 8× fewer active params, and scores 5.5× the base DiffusionGemma it was fine-tuned from. Methodology, weights, and full benchmark results in the blog 👇
显示更多
0
38
1.1K
87
转发到社区
CS329A Self-Improving AI Agents 斯坦福大学一门关于「能够通过与自身和环境的交互持续自我改进的 AI Agents」的课程。 讲师阵容 Azalia Mirhoseini:AlphaCode 与 LLM 大规模训练调度优化("Circuit Training")的核心人物,现为 Reflection AI 联合创始人 Aakanksha Chowdhery:PaLM 2 与 Gemma 的负责人,同样在 Reflection AI 嘉宾名单 Misha Laskin(Reflection AI CEO) 多位 Google DeepMind 研究员(Denny Zhou、Thang Luong、Melvin Johnson)、Physical Intelligence(Danny Driess,机器人) 课程主线:一条完整的自我改进技术栈 1. 推理时自我改进(Test-time) 2. 从反馈中学习(Feedback → Learning) 3. 开放式进化与搜索(Open-endedness) 4. 工程化与瓶颈(Agentic Engineering & Bottlenecks) 课程详细信息:
显示更多
0
61
497
112
转发到社区
When we first introduced @GoogleGemma, our family of open models, our goal was to give developers the tools to build responsible, innovative AI applications anywhere. Today, Gemma models have surpassed a billion downloads. 🎉 Over the past two years, developers have published over 100,000 Gemma model variants and built a thriving innovation ecosystem we call the Gemmaverse. We’re taking a look at how the Gemmaverse is making an impact across the globe ↓🧵
显示更多
0
63
1.3K
109
转发到社区
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
显示更多
0
159
4.3K
580
转发到社区
可能很多人对27B模型没什么概念。 27B听起来好像也没多大啊,现在普通玩家128G内存都开始普及了。 但我讲几个冷知识。 Google训练Gemma 3 27B的时候,预训练数据是14万亿Token。 不是140亿,也不是1400亿。 是14万亿。 训练它的集群用了6144颗TPU v5p。 而27B这个数字,只代表它有大约270亿个参数。 光把BF16原始权重放进内存,大概就是54GB。 32K上下文跑起来,加上KV Cache,大概72.7GB。 注意,这还只是推理。 真正训练的时候还要塞梯度,优化器状态,激活值,各种通信Buffer。 所以你看到一个27B模型只有几十GB,很容易产生一种错觉。 这玩意我电脑都装得下,我是不是也能训练。 完全是两回事。 还有一个更离谱的地方。 14万亿Token是什么概念。 假如一个人一天非常夸张地读10万个Token,而且一天不休息。 读完这些训练数据,大概需要38万年。 而机器要做的也不只是把这些字读一遍。 它要一层一层做矩阵运算,算损失,反向传播,再一点一点修改270亿个参数。 所以我现在越来越觉得,本地能跑27B,甚至70B,是一件非常牛逼的事情。 不是因为我们的电脑已经可以训练它们了。 而是因为人类把一个需要几千颗AI芯片训练出来的东西,压缩,量化,优化之后,居然可以塞进一台桌子下面的小电脑里。 这个过程本身就挺赛博朋克的。
显示更多
0
72
867
81
转发到社区
Grok 4.6 multimodal is a step change from Grok 4.5. It’s one of the under-discussed improvements, and I’ve been very impressed by it. My daily work includes reviewing lots of videos and understanding the context; Grok 4.6 improves the productivity of such workloads by at least 10x if not 100x. Such workflows may not be captured by common VLM benchmarks, but in my use cases it outperforms Gemini 3.5 and Gemma 4, which is considered the SOTA of VLMs in my opinion. Hats off to the multimodal teams—you did a great job.
显示更多
0
43
507
27
转发到社区
兄弟们!想靠 AI 提升收入,别先迷信“工具清单”。 真正够用的配置就 5 个: • Codex应用(5.6 sol中等) • Hermes Agent(由Qwen本地模型驱动) • OpenClaw(ChatGPT 5.6 oauth) • Mac Mini上运行的Gemma 4 • ChatGPT语音助手,边走路边工作/健身 • Claude Fable 5用于规划 • Claude Design用于前端设计 • ChatGPT图像生成2,超乎想象 • Spotify播放lofi音乐 • 第二台显示器,24/7运行这些代理
显示更多
┏━┳━┳━┳━┳━┳━┳━┓ ┃本┃日┃2┃1┃時┃ま┃で┃ ┗━┻━┻━┻━┻━┻━┻━┛ 【完売目前…‼︎ お見逃しなく🔥】 お得な前売り券🎫⇢ 当日衣装の郵送チェキ💌⇢ ◻︎じっくり撮りたい方へ 各モデル毎の「順番撮りチケット」も当日販売🎫 限定枠なのでご購入はお早めに✨ 【前半TEAM-1】 益田アンナ(@anna_masuda) 百合川サシャ(@sasha_chu_) 【前半TEAM-2】 山口小雪希(@Koyu_221) 河野亜季子(@akiko_kono317) みさ(@_misadayo) 【前半TEAM-3】 雪村花鈴(@yukimura_karin) 月城レミ(@remiremi_38) 【前半TEAM-4】 叶野僾(@honiiiiichan) 成瀬いな(@ina_coscos) 峰尾こずえ(@kozurin69) 【前半TEAM-5】 大城かえ(@oshiro_kae) 琴里ここ(@coco_p1que) 【後半TEAM-1】 西野夢菜(@YumenaNishino) 阿久津こてつ(@kotenanoda4) 【後半TEAM-2】 薺かれん(@udon_chan8) 百瀬せいな(@momose_seina) 仲村まひろ(@mahironakamura) 【後半TEAM-3】 ジェマ・ルイーズ(@gemmalouisejpn) 日下部ほたる(@hotaru_kusakabe) 【後半TEAM-4】 宇佐美彩乃(@ayanon_usami) 清宮みいな(@miinyaaa23) 月森なな(@yokkyu530) 【後半TEAM-5】 波崎天結(@ayu_hazaki) 一ノ瀬ことね(@Kotone_ResChu) ⇢ #フレッシュ撮影会#
显示更多
0
0
103
14
转发到社区
Meet the Gemma Translator! A fully offline device powered by Gemma 4 E2B built with @Antigravity. Running entirely on a Raspberry Pi 5 with a connected microphone and speaker, this highly portable prototype is housed inside a custom, 3D-printed case.
显示更多
0
13
236
26
转发到社区
Every few weeks, a new model tops a public leaderboard. But in finance and business research, the model isn't the bottleneck anymore. Context is. Continuous benchmarking at AlphaSense shows GPT-5.6 Sol leading the pack, Opus 5 underperforming its predecessor at 5x the cost, and Gemma 4-31B matching Sonnet 5 quality at 40x lower cost. The lesson: newer isn't automatically better, token price doesn't equal question cost, and per-task model routing beats any single-model strategy by 2.8x. Chris Ackerson and Daniel Campos break down what months of head-to-head model testing reveal about frontier AI, and the two engineering programs designed to widen the gap even further. Read the full article:
显示更多
0
11
333
16
转发到社区