AI developers lack systematic ways to test for subtle edge cases that human experts instinctively recognize. This leads to overestimating system capabilities.
Build a platform where domain experts can describe challenging scenarios, which are then procedurally generated as test cases. Focus initially on games and strategy domains with clear evaluation criteria.
AI training teams would pay for high-quality evaluation datasets that reveal true limitations.
Start with Go position generators based on professional game archives.
Risk is creating test cases that are too esoteric to be practically useful.