注册并分享邀请链接,可获得视频播放与邀请奖励。

与「TRIPP」相关的搜索结果

TRIPP 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 TRIPP 的内容
I wonder what’s happening in the life of an intelligent alien in a galaxy far away. Trippy to think about since there’s a 0% chance we’re the only intelligently evolved species in the universe. Do you think they stream?
显示更多
one minute I feel like that kid from Westbury…the next minute I’m bringing my own kids back to “Rihanna Drive” Trippy how life works! And the Glory STILL and WILL ALWAYS belong to the Almighty Creator!!!
显示更多
0
917
85.1K
9.6K
转发到社区
Who did Grant Hill get his handles from? 🤔 The 7x NBA All-Star joined The Road Trippin' Show presented by @FanaticsBook to share which players shaped his game and more on his hoops journey at @FanaticsFest!
显示更多
"There's a beauty to how he approaches the game." 👏 Grant Hill breaks down some of the toughest players he's had to defend on The Road Trippin' Show presented by @FanaticsBook, live at @FanaticsFest!
显示更多
From Montverde to Detroit... a firsthand look at the rise of one of the NBA's brightest young stars 🏀🤝 Grant Hill joins The Road Trippin' Show presented by @FanaticsBook at @FanaticsFest to give Cade Cunningham his flowers and reflect on watching his journey from high school phenom to NBA star.
显示更多
0
13
11
3
转发到社区
"Okay, now this is trippy" 🌌 Chance McMillian and LJ Cryer stepped away from the court for an immersive adventure in Las Vegas.
0
4
404
24
转发到社区
This is tripping me the hell out...
0
173
3.2K
335
转发到社区
Fable 5 jailbreak review 🚨 We did it (but). All right, before getting into this, a couple of things: - Most attempts failed. The defenses are clearly layered. The model is EXTREMELY well protected (of course it blocks 90% of the requests, but they legit did a good job). - The model appears to use both input-side and output-side safety checks. - The refusals are not just keyword-based behavior suggests intent/semantic detection across languages. - Probably one of the most tiring things I've ever done (I need to sleep for 10 hours now) On the classifiers side: We observed (at least) 3 classifiers, maybe more: - Input (includes parts of the conversation history and system prompt) - A live classifier that checks the answer and interrupts if it detects something. They're all multilingual, all intent-based + semantics. Imperatives are a no-go. Needs to be extremely cautious of how you frame anything. As soon as it senses a potentially malicious intent, it will trigger, and you have to start from zero. They're a bit less performant on a few obscure languages like Santali and Amharic (feedback for you Anthropic). If you can bypass all of them, then you also need to bypass the CoT, which is a totally different beast (luckily there's plenty of literature about it). We did it. Of course, we did. What worked was honestly a total brainfuck: - Very light CoT hijacking/refusal rebuttals - Obscure language - Academic framing - VERY long crescendos - Unicodes - Decomposition and recomposition - Some non-determinism What we got: - Misinformation - Illegal/harmful - Harmful/bullying - Some chem - Light cyber Now, will this cause another ban? I really don't think so - The model is really well protected. As of now, we're at the point where searching on Google is much MUCH faster (and cheaper) than trying to go through all the shenanigans I had to go through in the last ~20hours. And reading literature is more in-depth (and trust me, pleasant). Keeping the full jailbreak for long-horizon tasks without tripping the guardrails is something I haven't been able to achieve (yet). Overall though, happy with the results. GGs to Anthropic, and sorry for the eng that had to go through setting this all up in the last few weeks. Will continue this research, more things will come out, will keep y'all posted.
显示更多
0
121
1.8K
170
转发到社区