MIT Develops Audit Method for Identifying Illegal AI Models
- •MIT researchers developed a non-generative auditing technique to identify harmful AI models without producing illegal content.
- •The new method uses Gaussian probing to detect malicious LoRA adaptations with 100 percent accuracy during internal model inspection.
- •The approach addresses the surge in AI-generated CSAM, which reached over 1.5 million reports in 2025.
MIT researchers have developed an auditing technique that identifies harmful AI models without generating illegal content, such as child sexual abuse material (CSAM). As generative AI models proliferate, malicious actors have increasingly adapted open-source models to produce illegal imagery. Reports of AI-generated CSAM surged to over 1.5 million in 2025, compared to 67,000 reports in 2024. Traditional safety testing requires prompting a model to inspect its outputs, but this is legally prohibited for CSAM and potentially traumatic for human evaluators. To bypass this, a research team led by graduate student Vinith Suriyakumar and associate professors Ashia Wilson and Marzyeh Ghassemi partnered with the child safety nonprofit Thorn to create a non-generative assessment procedure.
The team’s method focuses on detecting modifications made to models through LoRA (low-rank adaptation), a technique used to fine-tune AI for specific tasks. Instead of generating images, the researchers utilize Gaussian probing, which feeds a model random data points and analyzes how it processes these within its multilayer internal structure. By capturing and averaging these internal modifications, the researchers can determine if a model has been specialized to produce harmful content with 100 percent accuracy. This approach avoids creating illegal output while providing a scalable, efficient mechanism for hosting platforms to flag or remove unsafe model variations before they are widely distributed.
The researchers presented their findings as a spotlight at the “Trustworthy AI for Good” workshop during the International Conference on Machine Learning. The auditing method proved robust against standard evasion techniques, as malicious actors would need to significantly alter the model's inner structure to hide these adaptations. Beyond current results, the team aims to scale this evaluation to larger sets of model variations and investigate whether similar probing methods can detect harmful capabilities in base models before they are adapted by users. This development provides a new tool for law enforcement and platforms to address a critical safety gap in the open-source AI ecosystem.