AI researchers struggle to objectively compare model performance across different labs and architectures.
Build a suite of standardized tests that measure compute efficiency, training stability, and real-world generalization beyond just accuracy.
Sell enterprise licenses to AI labs and cloud providers who need credible benchmarking.
MVP could be a simple set of containerized evaluation scripts for common benchmarks.
Risk: Labs may prefer proprietary benchmarks that favor their own models.