Detecting inappropriate material used to train AI image generation models

Jeremy Budd; Gandhar Joshi; Lucia Noelle; Matthew Pickering; Siddharth Setlur; Irina Starikova; Charles Morehead; Ningyuan  Xu; Kairui Zhang; Maxim Zyskin

doi:10.33774/miir-2026-rx50v

Computing & Robotics

Search within Computing & Robotics

Detecting inappropriate material used to train AI image generation models

05 February 2026, Version 1

Working Paper

Show author details

Abstract

Generative AI models, for example diffusion models, have emerged as state-of-the-art methods for generating novel images described by a text prompt. Open-source AI models can furthermore be fine-tuned to produce images similar to a given dataset of images. However, bad actors may seek to use illegal images to fine-tune a model so that it produces inappropriate and harmful images. We investigate various methods for detecting whether such images have been used to fine-tune a given diffusion model. This task raises two key challenges: (1) Images from a suspicious model should not be produced. (2) Any prompts yielding inappropriate images may be obfuscated. We propose a multi-layered framework to overcome these challenges. We combine embedding analysis, trajectory classification, parameter inspections, and neural network encoding in a promising framework, and suggest that controlled experiments should be conducted to test this strategy in future work.

Content

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.