Companies want to use open LLMs but lack trusted data on how they perform for specific use cases like support tickets or contract review.
Create a continuously updated service that runs new models through standardized business scenarios (email drafting, code review etc) with blind ratings.
Monetize through enterprise reports detailing which models work for which industries.
MVP: Compare 3 major open models on 5 task types.
Risk: Model providers may game the benchmarks.