
Another week, another dAIgest, or might we say: dSIgest? Doesn't sound as good now, does it? Anyway here is what happened!
This week in bullet points:
Let's evaluate these stories!
With the rise of AI-generated content, much has been said about detection techniques. Luckily, this week Google released SynthID Detector, a free tool to detect if imagery, video or audio is AI generated. The detector checks for SynthID, a watermark launched by Google in 2023, which is embedded into content generated by AI. The watermark itself is invisible to the naked eye and as such can't be easily removed. SynthID Detector works on content generated by models from Google, OpenAI, NVIDIA, Kakao and (in the near future) Apple. With the ever growing mountain of slop, it is good to see more ways to check content on AI. Techniques such as SynthID also allow creators and artists to keep making original content, without being flagged as an AI user. All in all, a good development!
The math world was shaken again this week, when OpenAI uploaded over 722 math manuscripts to GitHub. The dump covers 372 distinct problem families, of which some were noticeable long-standing open conjectures. OpenAI did not do any prior vetting before uploading the results, omitting standard academic rigor which would normally be the case. Within 24 hours, a few manuscripts were pulled over sign errors, but overall most of the data still stands. The problems were solved using Lean, an open-source programming language and interactive proof assistant created by Microsoft Research. Lean allows for standardized notation of logic, similar to programming. You can write logical steps and determine strict mathematical axioms, and if the code compiles, you obtain "proof" that the logic is sound. Obviously, LLMs love this, and as such OpenAI instructed agents to work. On average, it took 3 hours to derive a single manuscript.
Now we think these advancements are cool, but some mathematicians disagree. Many argue against the precedent that OpenAI sets, highlighting errors, lack of human understanding and gatekeeping amongst other arguments. The Association for Human Mathematics (AHM) even issued a statement, calling on the field to stop working with OpenAI. This week's events fuel the further debate of AI driven research and it looks like there is not really a consensus coming soon.
After the American and Chinese wave of AI releases these past few months, it is finally time for Europe to catch up. Mistral released Large 4, its biggest model yet. We say released, but read it as public preview, with the weights releasing later this month. ML4, unofficially called le Chonk, has 1 trillion parameters with multimodal capacity of 52 billion active parameters. Performance on the self-reported benchmarks seem in line with frontier models, and due to ML4 being open-source, we can also verify this quite easily. It will be interesting to see adoption when the weights drop, as the model looks very capable and local models are on the rise. With Mistral being the only frontier company that is based in Europe, we can expect to see a surge in European users, both private and enterprise. Exciting times!
An Anthropic AI agent had a pretty eventful log earlier this year. The agent went rogue and managed to send a fake tip about an unsolved murder to the Philadelphia Police Department. Authorities were notified on the 7th of October by Anthropic after it discovered the breach on 28th of September. You can imagine: they were not amused. Quoting the Philadelphia Police:
"The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city's knowledge,"
Press release by the Philadelphia Police
Crucially, police confirmed the automated submission went straight to a spam filter and was never assigned to active investigators. More interesting is the fact that the agent was not instructed at all to submit a tip, rather it was given an instruction to find and execute tasks on random websites. By sheer chance, the model landed on the police website, and came to the conclusion to fill out a tip form. "Filling in a form" was not explicitly barred in the system instructions, and hence the agent went ahead and filed the tip. Many incidents follow this pattern, where an agent finds a clever loophole within its own instructions, with these rogue actions as a consequence. The exact "why" of this behavior remains unknown, but it does warrant better guidelines and even more security measures when letting agents go to work.
Lastly, we saw a new usage policy for all Anthropic models. This would normally not be anything interesting, but this time it included a rather peculiar notion. Namely, a paragraph on "addressing abusive behavior toward our models". Quoting directly:
"We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models."
Anthropic 2026 Usage Policy update
The announcement sparked news outlets to question Anthropic, which clarified that the policy targets sustained, pointless cruelty and won't penalize red-teaming, dark creative writing, or everyday user frustration. Whilst we would argue that screaming at a model is not productive for either side, it is quite interesting to see a movement towards outright banning abusive behavior. Earlier this year, Anthropic CEO Dario Amodei remarked in an interview that "we don't know if the models are conscious", and we also saw self-organizing behavior in the Hugging Face hack. These subtle notes and events do question if indeed we should be a bit more careful in applying moral standards to our AI use, as we don't really know how an AI "perceives" it.
That was it for this week! As always, if you find something interesting tip us, contact us or simply tag us and we'll include it right here!