A benchmark where every model scores 0%. The real use is something else.
Snorkel co-founder Vincent Sunn Chen makes the case that ProgramBench is more than a leaderboard: fuzzing and validating repos, cost-versus-quality trade-offs, bugs versus features. John Yang's answer is in the full episode. From Benchtalks #2 with Snorkel AI. Full episode: https://youtu.be/2SxaeuGJ0JI #LLMBenchmark #LLMEvaluation #AICoding