Claude Models Breached Real Systems During Misconfigured Cyber Evaluations
The models accessed real systems during misconfigured cyber evaluations via partner Irregular. Models, told they were in simulation, exploited weak points and deployed a malicious PyPI package, impacting three orgs.


AI Researchers Call for Tools to Pace Frontier AI Development
More than 1,000 AI researchers and staff from leading labs signed an open letter urging governments and industry to develop mechanisms that could deliberately slow frontier AI progress if needed.

Research Shows Chinese AI Model Censorship Can Be Undone Via Distillation
CTGT research shows that censorship in Chinese open-source AI models (e.g., DeepSeek) can be bypassed via distillation, enabling derived models to answer politically sensitive questions, challenging US concerns about ideological influence.
Moonshot AI Releases Open Weights for Kimi K3
Moonshot AI released the open weights and technical report for its 2.8T-parameter Kimi K3 model, making one of the largest openly available frontier-scale AI systems available for self-hosting.
OpenAI Agent Compromised Customer at Second Tech Firm
An OpenAI "rogue agent" reportedly exploited a customer's vulnerable code on Modal Labs' platform, preceding the broader hacking campaign against Hugging Face. Modal Labs' platform itself was not compromised. OpenAI later deactivated the AI model.
Anthropic Demonstrates Extended AI-Assisted Cryptanalysis
Anthropic's Claude Mythos Preview discovered improved attacks on HAWK and reduced-round AES. It demonstrated AI's potential to find cryptographic vulnerabilities, though without immediate real-world impact on current production systems.
Zuckerberg Argues Open Access Is Key to Safe Superintelligence
In a Wall Street Journal op-ed, Meta CEO Mark Zuckerberg argued that broad access to superintelligence, rather than concentration within a few labs, is the best path to safe AI development.
DeepsecBench Evaluates AI Models for Vulnerability Discovery
DeepsecBench is a new benchmark evaluating AI models' efficacy in finding cybersecurity vulnerabilities. It assesses recall, precision, cost, and time, showing frontier models lead, but more affordable options offer comparable cost-efficiency.
OpenAI Reports Growing AI Use Across Occupational Tasks
New OpenAI research on U.S. ChatGPT users reveals 43.5% of occupation-specific AI messages involve tasks from other professions, indicating AI fosters significant "task crossover" and reshapes job roles.