As LLM API costs accumulate, teams seek smaller models but struggle with quality tradeoffs. Tools to analyze and optimize model size could help balance cost and performance.
Develop a profiling toolkit that suggests model reductions with minimal accuracy loss. Monetize through enterprise licenses for teams.
MVP could analyze popular open models like Mistral with basic pruning suggestions.
Fast-moving model architectures may make optimizations quickly obsolete.