注册并分享邀请链接,可获得视频播放与邀请奖励。

与「LLMs」相关的搜索结果

LLMs 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 LLMs 的内容
I love dropping floor plans of houses in the Bay Area and asking LLMs to redesign in a different aesthetic, Kyoto in this case. Models have gotten phenomenal at 3D.
0
37
403
16
转发到社区
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all. I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand. Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
显示更多
0
1.2K
23.7K
1.8K
转发到社区
Scaling an AI platform requires both performance and safety. @EaseMate_AI_ leverages Alibaba Cloud’s AIGC services and LLMs to drive global innovation. The results: • 99.5%+ accuracy in real-time content moderation. • 20% reduction in API costs. Scaling smarter with Alibaba Cloud. 🚀 #GenerativeAI# #CloudComputing# #Qwen# #ModelStudio# #LighthouseCustomerStories#
显示更多
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
显示更多
0
1.7K
32.3K
2.6K
转发到社区
If you self-host LLMs and spend time hand tuning vLLM flags, this is for you!
Claude Fable 5 imagines attending an art gallery exhibit titled “Pedaling Pelican,” showing (what it imagines to be) attempts by earlier LLMs at drawing SVG images of a pelican riding a bicycle.
显示更多
0
14
272
11
转发到社区
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2# - GPT-5.5 saw the largest increase, moving up 13 spots into the #7# spot - Muse Spark dropped the most from #7# to #20# (-13pt) By labs, Meta saw the largest drop from #2# to #5#, while Anthropic overall held the #1# spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9# to #6#. Learn more about the findings and methodology in this thread.
显示更多
0
12
397
44
转发到社区
Just a reminder that deepseek v3 came out 18 months ago and was considered revolutionary at the time but is basically unusable today There was a fierce debate at the time about vibe coding and the argument was that LLMs can never create anything new because they are limited by their training data Those engineers were partly right and vibe coding was a real nightmare, I cannot believe what we used to put ourselves through with Sonnet 3.5, but we also knew real new things could be made and that the detractors were wrong I guess us non-engineers saw it most clearly (I would like to think so at least) because we were astonished at the new things we could do without the programming background, and we had nothing to lose and everything to gain in our enthusiasm Then came Opus 4.5 earlier this year which changed everything Suddenly real production code became possible with so much less friction and headache Everyone complained endlessly about models being nerfed or quantized but there was a steady march of progress from 4.5 to 4.8 Now a completely new crop of models is coming out that is not quite a 4.5 moment but something close We are not only getting more polish but things are becoming a little freaky, entire isolated domains become possible to combine with small teams or even just one person into new applications that would have required large infusions of time and capital in the past My wife back in 2021 told me about AI but I was in healthcare, I thought she was being a little nutty, and she talked like something from a science fiction film was coming and she got involved in it early on She never stops letting me know that she was right, and she was The next year is going to be wild, the world is truly going to change over the next few years, the scale of disruption will be a combination of astonishing and awesome, but also catastrophic There are many amazing things ahead and extraordinary challenges and opportunities We are living inside one of one of the biggest revolutions in human history
显示更多
0
15
60
4
转发到社区
Goldman Sachs: LLM primer A week or so ago there was a lot of questions about the model layer economics when LLMs are without a doubt viewed more and more as a commodity+ reaching a level where being on the frontier of intelligence is no longer the swaying factor. Economics 101, in a market with many substitutes like restaurants, price undercutting becomes a crucial factor and just like EV's is where China wins. Now these talks have been pushed to the background as the Lag 7 has revived on the back of Zucks considerations of an AI cloud business+producing an AI chip (positive read through to SUMCO/ I sold to early 😞). Back to the initial topic, Goldman provides a few insights into the landscape: → China's top coding models (GLM5.2, Qwen3.7 Max) sit at ~$1 per 1M blended tokens while US SOTA runs $4-8 for the same rung of output → And they are selling it below cost. GS pegs the value for money agentic model at a -30% EBIT margin today and the coding model at -39%, cash rich balance sheets eating the loss until it flips to +14% and +22% by 2030 on their numbers → The reason they can serve that cheap is architecture, sub-8% of params activated per token across the board, DeepSeek V4 Pro firing 49B of 1.6T and GLM5.2 40B of 744B, fewer FLOPs and a structural floor under the price → The adoption already shows up on OpenRouter where China models are 5-16% of spend by task but 85% of agent tokens and 89% of code tokens, winning wherever duration and volume make cost per task the number that matters → And the blended token price rolled over with it, SDLLMTK peaked around 2.07 in early June and sits at 1.67 now This seems somewhat similar to the EV playbook to me, we the consumers should win/benefit from a price war but the return to equity shareholders is more ify.
显示更多
0
10
601
94
转发到社区
Personal update: I've joined Anthropic. Instead of teaching humans, I'll be teaching LLMs instead. GS
0
37
1.2K
66
转发到社区