Does Your Coding Agent Get More Efficient Over Time?
Does your coding agent get faster at working in the same repo over time? Continual Learning Bench has a task for this. The model gets a sequence of tasks in the same repository. Success is measured by how many bash commands it takes to get to a correct solution — and whether that number goes down. Most models: it doesn't. #AIAgents #CodingAgents #ContinualLearning #Benchtalks