Can LLMs Actually Reason? Dan Roth on Agents, Retrieval & Evaluation | Oracle x Appen | Ep. 3
Can an LLM actually reason, or does reliable reasoning emerge from the system built around it? Dan Roth, Chief AI Scientist at Oracle and Distinguished Professor at the University of Pennsylvania, joins Jeanine Sinanan-Singh, Director of GenAI Research at Appen, and Brian Jenkins, VP of Sales & Marketing at Appen, for a technical discussion on what changes when we stop evaluating AI as a model and start evaluating it as a system. They get into a set of problems that are becoming increasingly important for researchers building and evaluating agentic AI systems: • Why an agent is increasingly a system, not a single model • Where LLM reasoning breaks down and when systems need specialized solvers • Why multi-agent architectures can be necessary for context, privacy, and governance • Why retrieval over enterprise data is still far from solved • How semantic enrichment can improve retrieval across structured and unstructured data • Why passing a reasoning benchmark does not prove a model is reasoning • How human-authored stress tests can expose failures that automated perturbations miss • Why realistic evaluation environments need the messiness of production data • How runtime monitoring changes the way we think about AI evaluation • Why “visibility of failure” may be one of the hardest problems for deployed AI systems • Why future agent evaluation may focus as much on the harness and routing system as the underlying models For researchers working on LLM reasoning, agentic AI, retrieval, RAG, model evaluation, AI reliability, and enterprise AI systems, this conversation gets into the gaps between benchmark performance and systems that can actually be trusted in production. CHAPTERS 00:00 What makes an AI system intelligent? 02:56 Are agents actually a new form of intelligence? 06:25 Evaluating agents as systems 10:03 Can LLMs genuinely reason? 12:56 Models, systems, and mathematical problem solving 15:41 How do you stress-test reasoning? 18:07 Parametric knowledge vs. external knowledge 21:48 Why retrieval is not solved 25:56 Tool use, agents, and offline data enrichment 32:59 Can evaluations predict production behavior? 35:52 Evaluation, routing, and visibility of failure 39:09 What changes over the next 3–5 years? 40:29 Why agent evaluation may move to the harness ABOUT THE DATA LAYER The Data Layer by Appen explores the research, systems, data, and evaluation methods shaping the next generation of AI. Subscribe for conversations with researchers and AI practitioners working on the problems behind increasingly capable and reliable AI systems. Learn more about Appen Research: https://x.com/AppenResearch #AgenticAI #LLM #AIEvaluation #AIResearch #Retrieval #RAG #AI Dan Roth, Oracle AI, agentic AI, LLM reasoning, AI reasoning, AI agents, agent evaluation, AI evaluation, LLM evaluation, retrieval augmented generation, RAG, AI retrieval, runtime monitoring, AI reliability, multi agent systems, AI agents research, enterprise AI, neurosymbolic AI, Appen Research