Session 5: Data in the Age of Generative AI

Stanford HAI
274 views October 30, 2025

Massive datasets are the cornerstone for developing large language models (LLMs) and other generative AI. However, these datasets have also sparked debates regarding generative AI, highlighted by several copyright disputes involving OpenAI. This talk explores critical aspects of data creation and attribution for generative AI. Throughout the project, the researchers aim to ground their research with real-world legal and policy considerations and high-impact applications in law and medicine. Learn more about other research supported by the Hoffman-Yee Research Grant program here: https://hai.stanford.edu/research/grant-programs/hoffman-yee-research-grants?section=2024-grant-recipients 00:00:00 Lecture 00:29:06 Q&A

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close