Developers running local LLMs lack clear data on how different hardware handles model variations after framework changes.
Create a testing suite that benchmarks performance, memory usage, and accuracy across combinations of models, quantization methods, and hardware.
AI startups would pay for reliable compatibility data when choosing deployment targets.
Start with a simple script that tests a few common ggml models.
The risk is the rapid pace of change in quantization methods making results quickly outdated.