注册并分享邀请链接,可获得视频播放与邀请奖励。

Karl Mehta 的个人资料封面
Karl Mehta 的头像

Karl Mehta (@karlmehta)

@karlmehta
3x Exited Founder/ CEO of tech cos, Chairman Emeritus- QUIN(Quad), former VC@Menlo Ventures, Author of 2 books, fmr White House fellow. All tweets personal.
3.2K 正在关注    144K 粉丝
Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (@StarTalkRadio) with Neil deGrasse Tyson.
显示更多
0
23
203
37
转发到社区
Mark Zuckerberg explains the 405B teacher-model flywheel that could make one giant AI the wrong end state "People are gonna wanna do inference directly on the 405 because it's, you know, by our estimates, it's gonna be about 50% cheaper, I think, than GPT-4o to do that directly." "Because it's open weights, the ability to take the model and distill it down to whatever size that you want, to use it for synthetic data generation, to use it as a teacher model." "Our vision is that there should be lots of different models. I think every startup out there, every enterprise, governments, they all kind of wanna have their own custom models." "Right now, as open source basically closes the gap, I think you're just gonna see this wide proliferation of models where people now have the incentive to basically customize and build and train exactly the right size model for what they're doing, train their data into it." "They're gonna have the tools to do it because of a lot of the partner integrations that the companies like Amazon are doing with AWS or Databricks or different folks like that who are building these whole suites of services for distilling and fine-tuning open models." The counterintuitive edge is that the 405B model may be most valuable as raw material, not an endpoint. The open model compresses into the right size, absorbs proprietary data, and turns one frontier release into thousands of company-specific systems. Distribution of intelligence beats centralization. - Mark Zuckerberg (@finkd), CEO of Meta, with @rowancheung
显示更多
0
45
234
26
转发到社区
Palantir CEO Alex Karp on why the model is not the moat. The application layer is: "To make them valuable in an enterprise, like a battlefield, regulated, or manufacturing context, you have to have what's called an application layer." "We have this thing called ontology that now everyone's copying. It takes a large language model and makes it safe and useful and precise." "Safe because it doesn't touch your underlying data." "Safe because it prevents the large language model from caching your data and replicating your business." The tell for any enterprise AI pitch: ask what sits between the model and your data. If the answer is nothing, you are not buying leverage. You are renting a fast way to hand your business to whoever owns the weights. Alex Karp, CEO of Palantir, on CNBC.
显示更多
0
28
195
32
转发到社区
Dario Amodei, the CEO of Anthropic, went on CNN and put a number on the thing most AI CEOs will not say out loud: Half of all entry-level white-collar jobs gone. Unemployment at 10 to 20%. Inside 1 to 5 years. "The most salient feature of the technology, and what is driving all of this, is how fast the technology is getting better." "A couple of years ago, you could say that AI models were maybe as good as a smart high school student. I would say that now they're as good as a smart college student, and sort of reaching past that." "I really worry, particularly at the entry level, that the AI models are very much at the center of what an entry level human worker would do." "What is striking to me about this AI boom is that it's bigger and it's broader and it's moving faster than anything has before." "People will adapt, but they may not adapt fast enough. And so there may be an adjustment period." The tell: this is not a critic outside the industry. It is the man building the models saying the quiet part on camera -- because he thinks the people whose jobs are on the line adapt slower than the models improve. High school to college in two years was the warm-up. The entry level is where it lands first.
显示更多
0
192
744
133
转发到社区
Thanks @DarioAmodei for your deep-thoughts on this critical and crucial challenge that we face for humanity. below - my response to it @AnthropicAI
Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap:
显示更多