Developers running LLMs locally lack consistent ways to compare models across hardware. A benchmark suite could measure tokens/sec, memory usage and quality for common tasks.
Offer as open source with paid detailed comparisons. Start with basic speed tests on common setups.
Risk: rapid hardware/model obsolescence.