Companies are deploying powerful AI with little standardization around safety testing. Each team reinvents evaluation methods.
A platform offering configurable test batteries for harmful outputs, bias, jailbreaks etc. Would provide scores comparable across models.
AI labs would pay for thorough testing. Regulators might adopt as standard.
MVP: Basic set of tests for text generation models.
Risk: Requires deep safety expertise to create valid tests.