注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Retro」相关的搜索结果

Retro 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Retro 的内容
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
显示更多
0
159
4.3K
580
转发到社区
J-3 avant de les retrouver !!
0
23
2K
153
转发到社区
THE FOUR-PAYMENT ILLUSION ────────────────── "Four interest-free payments" isn't a discount. It's underwriting with the numbers hidden. ────────────────── > someone posted their Klarna balance today: £17,000 - on a product marketed as 4 payments, no interest > BNPL providers don't share data with each other - split $2,000 across Klarna, Affirm, and Afterpay and none of the three sees what the other two already lent you > miss one installment and the "interest-free" plan can convert to late fees - and in some cases retroactive interest - on the full balance > unlike a credit card, most BNPL debt isn't reported to a bureau until it's already gone to collections, so there's no early warning for you or your next lender ────────────────── A credit card has to tell you the APR up front. BNPL sells you the absence of a number - then collects like there was one all along.
显示更多
Would you like to make pics in a retro photo booth? :p
0
51
7.9K
107
转发到社区
📂 HERMES ┃ ┣ 📂 VPS ┃ ┣ 📂 Petit Serveur Linux ┃ ┣ 📂 Utilisateur Dédié ┃ ┣ 📂 Accès SSH ┃ ┣ 📂 Pare-Feu ┃ ┣ 📂 Tailscale ┃ ┗ 📂 Sauvegardes ┃ ┣ 📂 Installation ┃ ┣ 📂 Installateur Officiel ┃ ┣ 📂 Python Et Dépendances ┃ ┣ 📂 Hermes Doctor ┃ ┣ 📂 Mises À Jour ┃ ┗ 📂 Service Systemd ┃ ┣ 📂 Modèles ┃ ┣ 📂 Nous Portal ┃ ┣ 📂 OpenAI Codex ┃ ┣ 📂 Anthropic ┃ ┣ 📂 OpenRouter ┃ ┣ 📂 Modèles Locaux ┃ ┗ 📂 Changement À La Volée ┃ ┣ 📂 Interfaces ┃ ┣ 📂 Terminal ┃ ┣ 📂 TUI ┃ ┣ 📂 Application Desktop ┃ ┣ 📂 Dashboard Web ┃ ┣ 📂 IDE Via ACP ┃ ┗ 📂 API Compatible OpenAI ┃ ┣ 📂 Gateway ┃ ┣ 📂 Discord ┃ ┣ 📂 Telegram ┃ ┣ 📂 Slack ┃ ┣ 📂 WhatsApp ┃ ┣ 📂 Email ┃ ┗ 📂 20+ Plateformes ┃ ┣ 📂 Contexte ┃ ┣ 📂 AGENTS.md ┃ ┣ 📂 Règles Du Projet ┃ ┣ 📂 Préférences Utilisateur ┃ ┣ 📂 Historique Des Sessions ┃ ┗ 📂 Contexte Propre À Chaque Canal ┃ ┣ 📂 Outils ┃ ┣ 📂 Terminal ┃ ┣ 📂 Lecture Et Écriture De Fichiers ┃ ┣ 📂 Recherche Web ┃ ┣ 📂 Navigation Web ┃ ┣ 📂 Vision ┃ ┣ 📂 Génération D’Images ┃ ┗ 📂 Synthèse Vocale ┃ ┣ 📂 Mémoire ┃ ┣ 📂 Profil Utilisateur ┃ ┣ 📂 Conventions De Travail ┃ ┣ 📂 Environnement Technique ┃ ┣ 📂 Rappel Inter-Sessions ┃ ┗ 📂 Recherche Dans Les Anciennes Conversations ┃ ┣ 📂 Skills ┃ ┣ 📂 Procédures Réutilisables ┃ ┣ 📂 Catalogue Communautaire ┃ ┣ 📂 Commandes Slash ┃ ┣ 📂 Skills Personnalisés ┃ ┗ 📂 Amélioration Pendant L’Usage ┃ ┣ 📂 Intégrations ┃ ┣ 📂 Serveurs MCP ┃ ┣ 📂 GitHub ┃ ┣ 📂 Bases De Données ┃ ┣ 📂 Outils SaaS ┃ ┣ 📂 Home Assistant ┃ ┗ 📂 Plugins Personnalisés ┃ ┣ 📂 Automatisation ┃ ┣ 📂 Tâches Planifiées ┃ ┣ 📂 Crons Avec Skills ┃ ┣ 📂 Webhooks ┃ ┣ 📂 Scripts Sans LLM ┃ ┣ 📂 Livraison Multiplateforme ┃ ┗ 📂 Alertes Seulement Si Nécessaire ┃ ┣ 📂 Délégation ┃ ┣ 📂 Sous-Agents Isolés ┃ ┣ 📂 Travail En Parallèle ┃ ┣ 📂 Sessions En Arrière-Plan ┃ ┣ 📂 Pipelines Multi-Étapes ┃ ┗ 📂 Retour Automatique Du Résultat ┃ ┣ 📂 Sécurité ┃ ┣ 📂 Utilisateurs Autorisés ┃ ┣ 📂 Appairage Par Code ┃ ┣ 📂 Validation Des Commandes Dangereuses ┃ ┣ 📂 Protection Des Fichiers Sensibles ┃ ┣ 📂 Isolation Docker Ou SSH ┃ ┗ 📂 Blocage Des Commandes Irréversibles ┃ ┣ 📂 Recherche ┃ ┣ 📂 Veille Automatique ┃ ┣ 📂 Recherche Multisource ┃ ┣ 📂 Analyse De Concurrents ┃ ┣ 📂 Synthèse De Documents ┃ ┗ 📂 Rapports Sourcés ┃ ┣ 📂 Développement ┃ ┣ 📂 Inspection De Dépôts ┃ ┣ 📂 Création De Fonctionnalités ┃ ┣ 📂 Correction De Bugs ┃ ┣ 📂 Tests Et Lint ┃ ┣ 📂 Pull Requests ┃ ┗ 📂 Revue De Code ┃ ┣ 📂 Contenu ┃ ┣ 📂 Recherche D’Idées ┃ ┣ 📂 Rédaction ┃ ┣ 📂 Newsletters ┃ ┣ 📂 Réseaux Sociaux ┃ ┣ 📂 Images ┃ ┗ 📂 Audio ┃ ┣ 📂 Business ┃ ┣ 📂 Tri Des Emails ┃ ┣ 📂 Comptes Rendus De Réunion ┃ ┣ 📂 Mise À Jour Du CRM ┃ ┣ 📂 Rapports De KPI ┃ ┣ 📂 Analyse De Données ┃ ┗ 📂 Préparation De Documents ┃ ┣ 📂 Monitoring ┃ ┣ 📂 État Des Serveurs ┃ ┣ 📂 Changements De Sites ┃ ┣ 📂 Actualités ┃ ┣ 📂 Prix ┃ ┣ 📂 Logs ┃ ┗ 📂 Alertes Dans Discord ┃ ┣ 📂 Utilisation Quotidienne ┃ ┣ 📂 Message Depuis Le Téléphone ┃ ┣ 📂 Travail Pendant Que Le PC Est Éteint ┃ ┣ 📂 Reprise Des Sessions ┃ ┣ 📂 Correction En Cours D’Exécution ┃ ┣ 📂 Résultats Livrés Dans Le Bon Canal ┃ ┗ 📂 Validation Humaine Avant Action Sensible ┃ ┗ 📂 Boucle D’Amélioration ┣ 📂 Mémoriser Ce Qui Compte ┣ 📂 Retrouver Les Anciennes Décisions ┣ 📂 Transformer Une Méthode En Skill ┣ 📂 Réutiliser Le Skill ┗ 📂 L’Améliorer À Chaque Usage
显示更多
0
45
2K
159
转发到社区
🕺🏽Make your code GLOW & Retro with the SynthWave extension in #vscode# #vscodeextension# #eighties# ➡️
0
9
177
18
转发到社区
Six months (and a few days) ago, a day after leaving Google, I was entering the OpenAI office for my orientation as a new hire. Today I am posting a retrospective of the first 6 months.
显示更多
0
9
276
14
转发到社区
This Drake Maye retro autograph is BEAUTIFUL (Chris Uhlar/FB)
0
14
173
8
转发到社区
Celebrating the 55th anniversary of Personal Computing. New Kenbak-1 Ruler layout (connectors on the left side). Playing with the Video Editor OpenShot ... #KenbakRuler# #KENBAK# #retrocomputing#
显示更多