Web Scraping with CDP: What Actually Happens? | Zyte Podcast Ep. 12

Zyte
206 views September 9, 2026

When should you reach for a browser in web scraping—and when does managing one become a whole project of its own? In episode 12 of Extract Pod by Zyte, John and Ayan sit down with Paweł to explore browser-based scraping and Zyte’s Browser CDP. We unpack what actually happens when Playwright or Puppeteer connects to a remote browser, why browser fingerprints matter, and the trade-offs between running Chromium locally and using managed infrastructure. From JavaScript rendering and screenshots to AI agents navigating forms and multi-step workflows, this conversation covers where browsers earn their place in your scraping stack. We also dig into resource usage, session isolation, timeouts, concurrency, website-tier pricing, automatic proxy selection, and integrating existing Scrapy and Playwright projects. If you’re building scrapers, browser automations, or AI agents that need to interact with websites, this episode will help you understand the choices behind the connection. CHAPTERS 00:00 Introduction 00:35 Why use a browser for web scraping? 04:27 Browser overhead, ad blocking, and bandwidth 06:51 Local browsers vs managed infrastructure 13:03 How CDP connects your code to a remote browser 15:52 Use cases: screenshots, AI agents, and interactive flows 18:50 Browser fingerprints and stealth 21:34 Browser startup, reuse, and session isolation 23:19 Closing connections, timeouts, and failure handling 26:21 Browser contexts and resource management 28:28 Running concurrent browser sessions 30:52 Zyte API browser rendering vs CDP 33:59 Website-tier pricing and controlling costs 38:21 Automatic proxy selection 40:32 Geolocation and configuration trade-offs 43:05 Scrapy and scrapy-playwright integration 44:41 Puppeteer, raw CDP, and other libraries 46:35 Agent frameworks and authentication headers 47:56 Wrap-up ZYTE PRODUCTS & DOCUMENTATION Zyte Browser / Browser CDP: https://www.zyte.com/zyte-api/headless-browser/ Zyte API: https://www.zyte.com/zyte-api/ CDP setup, authentication, and Playwright / Puppeteer / scrapy-playwright examples: https://docs.zyte.com/zyte-api/usage/cdp.html Browser rendering, screenshots, and actions through Zyte API: https://docs.zyte.com/zyte-api/usage/browser.html CDP pricing and session duration: https://docs.zyte.com/zyte-api/usage/cdp.html#pricing CDP proxy, geolocation, and timeout settings: https://docs.zyte.com/zyte-api/usage/cdp.html#query-parameters Zyte API pricing guide: https://docs.zyte.com/zyte-api/pricing.html scrapy-playwright integration: https://github.com/scrapy-plugins/scrapy-playwright Product details reflect the discussion at recording time. For current capabilities, access requirements, and billing behaviour, use the documentation above. How are you using browsers in your scraping projects? Share your use case or a question in the comments, and subscribe for more conversations about web scraping, data extraction, and AI. #WebScraping #Playwright #Zyte

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close