Teams adopting LLMs lack objective benchmarks to compare model speed, cost, and quality across providers. An independent testing service could run standardized evaluations under controlled conditions.
Start with basic latency and throughput measurements on common tasks. Later add quality evaluations and price/performance metrics.
Initial version could test 2-3 models on simple tasks. Main risk is keeping up with rapid model changes.