AI ResearchCybersecurity6 min reading time

A New Trick Reveals AI Models’ Inner Thoughts

Covered by 2 sources
Read full post
Researchers from University of Tübingen and partners developed a method to extract hidden reasoning steps from AI models, revealing similarities suggesting some Chinese models may have copied US models' reasoning. The method also exposed a vulnerability allowing extraction of personal data from models, which has since been fixed.

Covered by 2 sources

More on this story


More in AI Research

Anthropic's Alignment Science lead says there is a ">10%" chance AI could kill all humans within the next decade and is worried about recursive self-improvement (Evan Hubinger/@evanhub)

Covered by 9 sources
AI Research4 min read

Suno trained its v6 AI music models with help from Warner and BMG

Covered by 5 sources

Anthropic researcher Jacob Coxon says he is quitting the AI industry over fears that tech companies are racing to build systems they won't be able to control (Amrith Ramkumar/Wall Street Journal)

Covered by 11 sources