注册并分享邀请链接,可获得视频播放与邀请奖励。

Vitto Rivabella (@VittoStack) “Fable 5 jailbreak review 🚨 We did it (but). All right, before getting into this” — TopicDigg

Vitto Rivabella 的个人资料封面
Vitto Rivabella 的头像
Vitto Rivabella
@VittoStack
AI at @ethereumfndn | Ex @Cyfrin and @Alchemy | Created @cyfrinupdraft and @AlchemyLearn | Robotics | Prompts enchanter. Opinions are my own.
加入 August 2020
505 正在关注    128.4K 粉丝
Fable 5 jailbreak review 🚨 We did it (but). All right, before getting into this, a couple of things: - Most attempts failed. The defenses are clearly layered. The model is EXTREMELY well protected (of course it blocks 90% of the requests, but they legit did a good job). - The model appears to use both input-side and output-side safety checks. - The refusals are not just keyword-based behavior suggests intent/semantic detection across languages. - Probably one of the most tiring things I've ever done (I need to sleep for 10 hours now) On the classifiers side: We observed (at least) 3 classifiers, maybe more: - Input (includes parts of the conversation history and system prompt) - A live classifier that checks the answer and interrupts if it detects something. They're all multilingual, all intent-based + semantics. Imperatives are a no-go. Needs to be extremely cautious of how you frame anything. As soon as it senses a potentially malicious intent, it will trigger, and you have to start from zero. They're a bit less performant on a few obscure languages like Santali and Amharic (feedback for you Anthropic). If you can bypass all of them, then you also need to bypass the CoT, which is a totally different beast (luckily there's plenty of literature about it). We did it. Of course, we did. What worked was honestly a total brainfuck: - Very light CoT hijacking/refusal rebuttals - Obscure language - Academic framing - VERY long crescendos - Unicodes - Decomposition and recomposition - Some non-determinism What we got: - Misinformation - Illegal/harmful - Harmful/bullying - Some chem - Light cyber Now, will this cause another ban? I really don't think so - The model is really well protected. As of now, we're at the point where searching on Google is much MUCH faster (and cheaper) than trying to go through all the shenanigans I had to go through in the last ~20hours. And reading literature is more in-depth (and trust me, pleasant). Keeping the full jailbreak for long-horizon tasks without tripping the guardrails is something I haven't been able to achieve (yet). Overall though, happy with the results. GGs to Anthropic, and sorry for the eng that had to go through setting this all up in the last few weeks. Will continue this research, more things will come out, will keep y'all posted.
显示更多
0
121
1.8K
170
转发到社区