HAI Seminar: Addressing Challenges of Public Web Data

Stanford HAI
183 views October 31, 2025

This HAI seminar featured Common Crawl Foundation’s work on preserving humanity's knowledge and making it accessible through its free public web dataset, a vital resource since 2008. The Common Crawl team presented insights from a new data product that utilizes Common Crawl's metadata to explore concerns around robots.txt exclusions, legal demands, and "bot defenses," advocating for greater transparency and informed solutions for the future of public web data. This seminar was recorded on October 22, 2025 at Stanford University. 00:00:00 Introduction 00:01:01 Lecture 00:48:53 Q&A

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close