There's no consensus on how to measure or predict AI self-improvement capabilities.
A dashboard that aggregates published benchmarks and runs standardized tests against model APIs.
Sell to AI labs and safety researchers who need comparative data.
MVP: Static reports on published models before building live testing.
Risk: Rapid pace of AI development may make metrics obsolete quickly.