注册并分享邀请链接,可获得视频播放与邀请奖励。

雪踏乌云 的个人资料封面
雪踏乌云 的头像

雪踏乌云 (@Pluvio9yte)

@Pluvio9yte
0 正在关注    0 粉丝
不要再研究什么无审查版模型了!qwen-image-2.1 拿来本地平替 GPT-image-2.5 才是最香的!完整部署教程来了 我不明白,qwen-image-2.1 本地部署完本来就没有验证和审核,大伙为啥都把重心放在奇怪的地方了。 这次 qwen-image-2.1 给我带来最大的震撼,还是其原生2K的生成质量,完全不输 GPT-image-2.5。最关键的一点是,它生成的图片不会出现GPT 经常出现的那种噪点,非常适合拿来做首帧参考图,拿去跑AIGC短片。 给大家带来单卡 4090 部署模型,并跑通2K图像的攻略:(直接发给Codex帮你部署就行) 不建议用梯子下载权重,直接走国内下 ModelScope 的分片,这样会快得多(国模最大的好处之一了,不用代理也下载快);diffusers 同理,可以不走github,用 Gitee 上的源码。 有个小坑,本来想直接加在现成的 Comfy 上。版本一查是 0.32,qwen-image-2.1 要 0.37 以上。这点大家要注意,用comfy的话一定要对上版本。版本不对不支持,服务起不来。 服务就做一个出图接口,一共33G 权重在启动时加载,我这边两分钟左右。不要等第一张图的请求来了再加载,那边会先超时。CFG 没开,官方默认就是关的,打开之后每步计算大概翻倍,2K 会更慢,没必要。 把这篇帖子内容喂给AI,它能直接帮你搞定部署流程,怕找不到的话大伙可以点个收藏。 最后再跑一下测试,我这边2048的尺寸,30步跑的话,耗时3分钟左右(CFG=0),成品如图所示,审核什么的完全不用担心,完全没有任何限制。 我有预感,本地部署 qwen-image-2.1 + Minimax-h3 很有可能成为接下来AI短片流水线的爆款流水线。
显示更多
0
17
169
10
转发到社区
让Codex调用Grok之后,我的工作流被大大优化了。 Grok Build是Grok的CLI,能够被目前的Codex最强的SOL模型这一最强大脑调用 Sol做规划或者思考大脑,然后调用Grok去X上能够搜索很多技巧或者帖子 调用方式直接用自然语言告诉Codex调用本地的Grok Build或者用下面这个skill
显示更多
正确使用 Tibo @thsottiaux 的姿势 点进他的主页,点开小铃铛 打开Codex,开启 Ultra 和 Fast 小铃铛响的时候换回到平时用的 sol-meduim ,关闭 fast
显示更多
0
58
174
17
转发到社区
正确使用 Tibo @thsottiaux 的姿势 点进他的主页,点开小铃铛 打开Codex,开启 Ultra 和 Fast 小铃铛响的时候换回到平时用的 sol-meduim ,关闭 fast
显示更多
0
73
50
1
转发到社区
来推特一年真是学到了很多有用的知识 1. 海外银行卡 最基础的,比如成功申请到了一些海外万事达卡(比如 Safepal、Bitget、Bybit,还有 N26 等国外银行卡),再也不怕 AI 被封号了 还去香港办了港卡,办下来了众安。有一个港卡十分方便 2. 海外电话卡 还办了几张海外的电话卡,比如英国的 Giffgaff、美国的 Paygo 紫卡等,再也没有去找过接码平台收验证码,省下了很多的精力 3. 自建了翻墙代理 之前一直都用付费服务,自从了解了很多 VPS 之间的区别,才知道了什么是优质的回国线路。也自己折腾了好几个 VPS,自建了翻墙代理。再也不用受制于人了,自建速度很快还很稳定 4. 用上了顶尖的AI Twitter 有一个好处,就是能够时刻看到很多开发者或博主分享一些模型的体感。有时候出了新模型,我不太容易去体验的,之后就能很容易通过看其他人的测评,及时地切换模型。 比如 GPT 5.5 出来的时候,我其实还在用 Claude。不看其他人的测评,根本就不知道当时的 GPT 5.5 到了何种水平,最后也转向了 200 美金的 OpenAI 订阅 5. 知道了自媒体是怎么运作的 我也渐渐不只是为了分享而发帖,同时也真正明白了“自媒体”这几个字的含义,懂了如何去做好自媒体
显示更多
众安实体卡到了,从申请到收到用时4天,真的快。
0
45
218
20
转发到社区
众安实体卡到了,从申请到收到用时4天,真的快。
0
17
19
0
转发到社区
Kimi K3做了五个前端 提示词: 在页面下创建五个文件夹 然后分别在五个文件夹中工作 创建五个不同风格的前端页面 尽可能发挥你能够展现的极限 感觉前三个风格非常不错,最后一个疑似蒸馏了Claude
显示更多
0
12
12
2
转发到社区
最近发现一个实用的开源项目:OpenCodex。 它可以让 Codex 不再局限于 GPT 模型,还能接入 Kimi、Grok、GLM 等其他大模型。配置一次后,原来的 Codex 应用和工作流基本不用变,需要时直接切换模型即可。 开源地址
显示更多
0
17
52
4
转发到社区
Kimi K3的游戏制作能力也在线啊,用K3制作了一个马里奥游戏,基本就是提示词直出的,没有修改过。 部署到vercel了,大家感兴趣可以试一下,地址在评论区。 还是带音乐的,有点带感。。。
显示更多
0
15
25
2
转发到社区
AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or @alayastd. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:
显示更多
接入anysearch后,我的agent搜索效率提高了 Parallel、Perplexity、Tavily 这些搜索工具给 Agent 用,有个一直没解决好的问题——返回的是链接加摘要,Agent 拿到之后还要自己点开网页、筛选内容、判断哪条有用。金融、学术、代码这些垂直领域的搜索质量更差。输出没有结构,光解析内容就要烧掉大量 token。 分享一个我一直在用的为 AI Agent 设计的搜索基础设施 —AnySearch,一个为 Agent 量身打造的「搜索 Skill」: - 实时网页搜索 - 支持金融、学术、代码、社会媒体等垂直领域搜索 - 直接把网页转成干净的 Markdown(extract) - 输出结构化,Agent 能直接拿来用 安装 AnySearch SKILL 超级简单:复制下方提示词并发送给你的 Agent,即可自动完成安装:
显示更多
我的视频复刻Skill在Sol模型的加持下,现在非常牛逼,只要你觉得某个视频是Remotion或者Hyperframes做的,都能复刻。 评论区还有视频样例和开源地址
显示更多
Codex 大更新 —— Chrome 内置最完全功能应用解析 1. 内置浏览器 + Chrome Cookie 一键导入 现在导入了 Cookie 之后,能够把抖音、小红书、推特的账号 Cookie 都导入。现在打开,直接就是登录状态。 2. 开发者模式 可以查 DOM、Console、Network 请求、性能等。 比如说,原来我开发网站如果遇到性能问题,让它优化就会麻烦一些。现在直接用内置的浏览器,速度非常快,这个优化速度。 下面是一些应用场景 内容创作与研究 让 Codex 带着你的账号去小红书、微博、知乎等平台“逛”,找选题、分析爆款、抓用户反馈。 开发者调试 开启开发者模式后,Codex 能实时 inspect 页面: 快速定位 console error、网络请求异常。 分析 DOM 结构、调试已登录态的内部工具。 提示词示例:“开启开发者模式,打开这个页面,检查 Network 请求里和用户数据相关的接口,并把 console 里的 error 列出来。” 测试需要真实登录态的 Web 应用等
显示更多
卧槽,Codex 这次更新,让 Chrome 扩展有点尴尬了。。。 现在 Codex 真的能带着你的账号逛互联网了! 一键导入 Chrome 的 Cookies 和 Password,Gmail、推特、小红书等各种网站打开就是你的登录状态 你可以直接让它: 查邮件和附件、跨平台找选题、巡检后台、下载报表、填表上传、核对订单...... 打开开发者模式以后,能进真实页面查 DOM、控制台和网络请求。 并且它在内置浏览器里跑,你可以继续用 Chrome,互不干扰。 各种做浏览器插件的,直接天塌了。。。
显示更多
0
22
154
23
转发到社区
这个抖音号15天破千粉,全部内容由AI生成。 从想法到成品,没有一条我动过脑子。 而我制作所有视频花费净时只有5小时🤣 等我测试到1w粉丝需要多久,不去尝试和探索真的永远不知道有多容易
显示更多
0
43
116
8
转发到社区
这里也列出一些我看到的利用自媒体快速变现的博主,从上百万到几十万的都有。 @Pluvio9yte (先分享个我不过分吧) @Zesee 我从Rachel这里学到很多认知 @Leobai825 最快最多变现 @libapi_ 我从来没想到还能卖硬件 @gkxspace 如何变现的大家仔细观察下 @bozhou_ai 28岁就已经辞职了 @Astronaut_1216 最容易复制的路径 @yaohui12138 商单变现顶级 如果你想学习怎么利用自媒体变现,不要只看他们的表象,学习一下他们的路径(跳出帖子,说多了要砸他们饭碗了哈哈哈),用心观察
显示更多
0
9
141
14
转发到社区
分享一个最近觉得很好用的提示词,从几万star的skill中抽出来的 开始之前,请遵守三条规则: 不要对我没有说明的信息做任何假设; 把你需要向我确认的问题列出来,按重要性排序,一次最多 5 个; 在我回答之前,不要输出任何方案或代码。
显示更多
0
15
144
12
转发到社区
论网感,找到了最低成本小红书起号的方式,最近没空写,下周末公开
0
21
71
2
转发到社区
最近咋都流行晒收入晒年龄 我,02年,上个月赚了23w(大概) 收入构成:去年外包尾款+商单+中转站+工资+付费咨询+培训 没有天赋,已经持续8个月一天工作14小时,全是一路的努力和辛酸
显示更多
0
94
279
6
转发到社区