注册并分享邀请链接,可获得视频播放与邀请奖励。

与「N-1」相关的搜索结果

N-1 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 N-1 的内容
保持总数10一定之下(一个元素集合,即空间),取映射f: 10→(N, 10-N),g: 10-N→N,使得这两个映射的某种复合为恒等映射。 即第一个操作是分割,第二个操作是取补集 为了保证取补的操作是合法的,补上一点空点,即无所谓的量 具体数学细节的补充是简单的 好的,如果你不是数学系的人多半是听不懂的,翻译成人话就是……你拿10个出来,全都翻个面,结束
显示更多
昨天忽然老板说业务不行团队要解散 老板不愿意N+1 我准备仲裁了(つД`)ノ 创业公司真的死得快 附两张美照增加流量! 还有哪个有意思的创业团队要招人吗
显示更多
0
94
241
2
转发到社区
Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated. Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method: In generation: 1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash 2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set. For watermark detection: 1. For each token, see if it was in the green or red set. 2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words. 3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark. I can tell you want to ask the following: 1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine. 2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4". 3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would need an exponentially large number of samples from the watermarker to reconstruct those sets exactly, but it's a risk if the detector is open to the wild (Watermark Stealing) Still, there are couple challenges that a frontier lab needs to overcome: 1. Their watermark needs to work token-by-token because they are streaming their text to users. Many watermark methods plan sentences or paragraphs at a time, or change the text after its entirely written, in order to make their watermark robust to paraphrasers, and a frontier lab cannot afford to do this yet (SemStamp, PostMark) 2. If the secret key leaks, the watermark is busted. To avoid a large blast damage from this, you need to have a couple secret keys in rotation. 3. There are some texts, like code, that cannot be arbitrarily changed, otherwise the code will break. In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms (i.e. like variable naming) - see SWEET, EWD, Invisible Entropy. 4. They will need to educate their users on how to deal with false positives and false negatives of a detector, which is a big challenge (one we put a lot of effort into) So, how do I see this playing out in the next 6 months? 1. If Anthropic releases the watermark detector publically, I think they defeat their own watermark. People find reliable watermark removal strategies by testing against Anthropic (AI detectors like GPTZero have an advantage here because they can train against these adversaries once they become popular). 2. If they keep the detector private to the government, like Google has done, it's "safer". However, there are some papers showing trained approaches that work robustly to zero-shot break watermarks without any data, simply because they try to write the text just like a human (Zhang et al. 2024, Watermarks in the Sand). Also, making your detector makes it battle-tested and stronger long-term (my experience). 3. In my testing, the watermarks don't survive intense paraphrasing (especially if you combine word choice and syntax attacks), or human text substitution (rewrite your AI text by plagiarizing human authors). The free paraphrasers I've tried have quickly bypassed Google Deepmind's SynthId for what it's worth. 4. All-in-all, frontier labs are likely okay with this because they expect most users to not attack the watermark, and also because they + European regulators likely don't care past a certain point - its good enough. 5. Overall, I think users of frontier LLMs will not really care about this, because 1) they don't realize watermarks are there, 2) EU will force everyone to conform, 3) this seems more like regulatory hoop-jumping than an earnest effort from frontier labs to expose LLM use Lastly, people's first concern shouldn't be watermarking, it should be AI detectors! If you're posting, "its not X, its Y!!", I don't think the watermark is going to make a difference :)
显示更多
0
121
2.5K
316
转发到社区
「n²が偶数ならばnは偶数」の証明で n = n² - n(n - 1) と変形するのを初めて見たとき流石に天才と思った。
0
14
1.8K
126
转发到社区
OpenAI公布十项数学与理论计算机科学新结果,全部由AI主导完成 OpenAI内部一个代号Astra的未发布模型,在球体堆积、编码理论、群论、算子代数、量子复杂度、格密码学和极值组合数学等领域,解决了十个悬而未决数十年的公开问题。每个结果都配有一份用Lean 4形式化的证明,证明代码公开在Github上( 这些数学论证本身是由模型生成的,人类只负责把论证整理成手稿、并完成Lean形式化。跟现在程序员用AI写代码的过程类似:AI写代码,人类程序员只负责review,甚至有的都不review,直接accept后做功能测试。 想想未来数学家、科学家像程序员一样,主要职责就是验证和Accept,未免有点太刺激。 OpenAI原文: 10道题目简介: 1. 高维球体堆积:多少个球能塞进一个盒子 想象在很高维度的空间里堆橙子,怎么堆最省地方、能塞下最多球,是一个跟晶体结构、通信编码都有关的经典问题。1978年提出的Kabatiansky–Levenshtein上界四十多年没人从根本上突破。AI生成的公式算出的指数约为每维度−0.604,比历史最佳的−0.599更紧。 2. 二元码与球面码:怎样让传错的信息还能被纠回来 给数字信号加冗余,让接收端即使传输出错也能纠正,是编码理论的核心问题——纠错能力越强(码字间距离越大),能塞进的合法码字就越少,这是一个此消彼长的权衡。这次的结果把已知的码字数量上界在指数意义上大幅收紧,缩小了理论允许的最好码和已知构造出的码之间的差距。 3. 非苏菲克群:所有“群”都能用有限的东西去逼近吗 群是数学里描述“对称性”的基本结构。一个自然的问题是:任意一个群,是否总能用足够大的有限置换群去足够精确地逼近(这类群叫“苏菲克群”)?这个问题多年没有定论。这次的结果构造出一个明确、有限表现的群,并证明它天生无法被这样逼近,即它是非苏菲克的。 4. Connes刚性猜想:换了一张“脸”,身份还是原来那个吗 每个群都能生成一种叫von Neumann代数的运算结构,有点像给群拍了一张“运算指纹”。Connes猜想认为,对某一类群来说,这张指纹应该能唯一认出原来的群,不会有两个不同的群共享同一张指纹。这次的结果构造出一个反例:两个结构不同的群,却生成了本质相同的von Neumann代数,说明这张“指纹”并不总是唯一的。 5. Permanent的计算下界:为什么有些矩阵运算天生就很难加速 矩阵有个大家熟悉的量叫行列式,靠消元法就能快速算出来。它有个“难兄弟”叫permanent,长得像行列式但去掉了正负号,恰恰就是这个符号差异,让permanent在已知算法里始终摆脱不了指数级的计算量。这次的结果证明,只用加法和乘法搭出的电路(不允许用除法这类“捷径”)去计算permanent,电路规模必须达到某个新的、更高的下界,从数学上解释了为什么这类计算“绕不开”。 6. 量子平行重复:把同一个博弈重复玩,输赢概率会怎样变 在双人合作博弈里,如果重复玩k次要求全部获胜,直觉上获胜概率应该随k指数级下降。这件事在经典(非量子)情形下早已证明,但如果博弈双方可以用量子纠缠来配合,之前的证明工具只能给出很弱的下降速度(大约是根号级)。这次的结果把这个结论真正推广到了一般的量子博弈情形,证明获胜概率同样会指数级下降。 7. 最近向量问题CVP:格密码学为什么“抗量子” 把很多点按固定的方向和间隔整齐排列在高维空间里,就构成了一个“格”。给定格外一个随意的点,找出离它最近的格点,这就是最近向量问题CVP。这次的结果证明,哪怕只要求“大致找到”一个足够近的格点(而不是精确最近),这个问题依然是NP困难的,也就是说没有已知的快速算法能保证解决它。这正是后量子密码学敢于依赖格结构的理论基石。 8. Ehrhart体积猜想:一个凸体最多能“胖”到什么程度 设想一个凸的几何体,它唯一包含的内部整数点就是它自己的重心,那么这个凸体的体积最多能有多大?这次的结果给出了每个维度下精确的答案,是(n+1)n/n!(n+1)^n/n! (n+1)n/n!。 9. 多色拉姆齐数:想在一堆点里完全避开“三角形”,要多少种颜色 拉姆齐理论说的是:只要一个系统足够大,规律和结构必然会自发出现。把这句话落到最经典的问题上:给一张完全图的每条边染上k种颜色中的一种,要让图里避免出现同色三角形,图至少要多大?这就是Erdős第183号公开问题。这次的结果证明,这个最小规模会随k呈“超指数”增长,也就是比任何形如 c^k 的指数增长都快,彻底解决了这道悬了很久的题目。 10. 极值图论中的两个反例:稠密程度能推出图有多“退化”吗 极值图论研究的是:给定一些限制条件(比如不能出现某种子图),一张图最多能有多少条边。这次的结果针对两个具体的公开猜想(Erdős第146和180号问题)构造出了反例,说明原猜想设想的“边数够多就必然导致某种退化结构”的推理并不总是成立。
显示更多
0
28
52
16
转发到社区
Subquadratic发布SubQ 1.1 Small:速度超快的LLM SSA将注意力计算从O(n²)降至O(n),1M token下比密集注意力省64倍算力、比FlashAttention-2快56倍。模型在12M token大海捞针测试中保持98%准确率,通用推理能力接近中上游前沿水平。面向金融、法律、代码库等需要完整长文档推理的企业场景。 官方介绍:
显示更多
资本家的嘴脸同出一辙 开水团10年员工被裁员拒绝接受N+1 被调岗到上海,拒绝后公司以未到岗解除合同
听说阿里大裁员 看来无招也在本次名单里 但是阿里又不想给n+1 他一人工资可以顶很多普通员工了。
0
29
126
1
转发到社区
离职谈判核心是明确N、N+1、2N赔偿标准、薪资基数、支付时效,彻底规避后续劳动纠纷。 谈判全程依托《劳动合同法》第47条、第87条法定规则,清晰界定合法补偿与企业违法双倍赔偿边界。 劳动者需提前留存完整证据、明确诉求,企业严守合规解除流程,双方依法平衡利益、高效达成协议。 其实最后就是赔偿钱的问题,扯来扯去的。。。
显示更多
建议:新手别滥用 viber coding 。 混用 中国和海外的大模型,viber coding ,我把所有研发坑都踩完了 : 1. 时区问题,中国的大模型看 北京时间 ,美国的大模型看 美国时间 ... 2. Error code 到底是 int 类型还是 str 类型, 傻傻分不清. 3. 软删除 是真的删除还是 标志为 'is_deleted' = true . 4. 状态码 ( 500 , 400 ... ) 分散在全项目各处都是 ... 5. 数据库慢查询设置 , N+1 问题 未考虑 ... 6. 短信验证码 、 实名认证的防暴力破解 未考虑 ... 7. 状态机 待会儿 8个状态 ,隔一天就10个状态,再隔一天 7个状态 ... 8. 同一个业务含义的变量,分不同函数名 或者 不同 变量名来写。 9. 写了而又不用的函数名。 10. 不用 ORM ,而是裸用 SQL ,带来 主键 自增长问题 。 11. 数据库设计了 JsonB ,导致 数据查询 和 数据检索的 困难问题。 12. 高危严重的水平越权漏洞 和 垂直越权漏洞 泛滥 ... ...
显示更多
0
16
52
5
转发到社区