Researchers lack tools to systematically detect when LLM outputs are influenced by the model's own training data rather than external inputs. This creates risks of feedback loops.
Build an analysis tool that estimates the self-referential nature of LLM outputs across different prompting strategies. Include visualization of potential recursion patterns.
AI safety researchers and responsible deployment teams would pay for better introspection capabilities.
Start with a simple classifier trained on known self-referential examples.
Challenge is defining meaningful metrics for something inherently fuzzy.