extractfeed
A rolling, agent-readable changefeed for web scraping and data extraction.

Markdown twin JSON agents.md

2026-09-06

Firecrawl releases official ChatGPT plugin

Firecrawl has launched an official ChatGPT plugin, enabling users to extract web data directly through conversational AI.

Firecrawl

Firecrawl introduces Anydoc and PDF Inspector for document extraction

Firecrawl announced two new tools, Anydoc and PDF Inspector, aimed at improving document and PDF data extraction.

Firecrawl

Firecrawl introduces Agentic OCR for structured data extraction from images

Firecrawl has released a new OCR feature that uses AI agents to extract structured data from images and documents.

Firecrawl

Firecrawl launches Research Index for life sciences

Firecrawl announced the launch of a Research Index tailored for the life sciences domain.

Firecrawl

Firecrawl releases Convex component for web scraping integration

Firecrawl announced a new Convex component that allows developers to integrate web scraping capabilities directly into their Convex applications.

Firecrawl

Firecrawl announces integration with Eden AI

Firecrawl has published a blog post announcing an integration with Eden AI.

Firecrawl

Firecrawl publishes case study on 11x's use of its web scraping API for prospect research

Firecrawl's blog details how sales automation startup 11x uses its web scraping API to automate prospect research.

Firecrawl

Firecrawl launches Developer Index for web scraping performance metrics

Firecrawl announced the launch of a Developer Index to provide benchmarks and performance data for web scraping tools.

Firecrawl

Browserless publishes enterprise Docker deployment guide for self-hosting

Browserless released a guide detailing how to deploy its browser automation service in an enterprise Docker environment.

Browserless

Browserless introduces Skill Bucket for agent-based browser automation

Browserless announces a new Skill Bucket feature that allows AI agents to access and use predefined browser automation skills.

Browserless

Browserless blog post examines the pitfalls of persisting browser profiles

Browserless published a blog post discussing the challenges and drawbacks of persisting browser profiles in automated environments.

Browserless

Browserless publishes guide on browser infrastructure for computer use agents

Browserless has published a blog post discussing the infrastructure requirements for running browser-based agents that interact with web interfaces on behalf of users.

Browserless

Browserless publishes guide on bypassing Datadome anti-bot protection

Browserless released a blog post detailing techniques to circumvent Datadome's anti-bot system.

Browserless

Browserless introduces the Browser Automation Protocol

Browserless has announced a new protocol for browser automation.

Browserless

Browserless launches Agent for AI-driven browser automation

Browserless has introduced a new product called Browserless Agent, designed to enable AI agents to control browser sessions.

Browserless

Bright Data Blog discusses robot training data from public web video

Bright Data's blog explores the use of publicly available web video as a source of training data for robotic AI systems.

Bright Data Blog

Bright Data integrates with Paperclip for AI-driven data extraction

Bright Data announces a partnership or integration with Paperclip, an AI tool, as detailed in a blog post on their site.

Bright Data Blog

Bright Data Defends Its Network Against Misuse Claims

Bright Data publishes a blog post arguing that its proxy network cannot be used in the ways critics allege.

Bright Data Blog

Bright Data explains Context as a Service for AI data retrieval

Bright Data published a blog post defining Context as a Service, a concept where external context is provided to AI models via structured data feeds.

Bright Data Blog

Bright Data compares its Cursor integration with default coding agent

Bright Data published a blog post comparing its Cursor integration against the default coding agent for web data tasks.

Bright Data Blog

2026-09-04

Playwright 1.63.0 adds test locks for safe concurrent access to shared resources

Playwright 1.63.0 introduces named test locks that prevent concurrent execution of tests sharing the same lock name across files, workers, and projects.

microsoft/playwright releases

Apify outlines six methods for downloading Instagram images in 2026

Apify published a guide covering six techniques for downloading Instagram images, ranging from a simple browser trick to a two-Actor pipeline for full profile extraction.

Apify

SerpApi publishes guide for scraping Zillow listings

SerpApi released a tutorial showing how to scrape Zillow real estate listings using its API with multiple programming languages.

SerpApi

ScrapingBee compares eight top Python web scraping tools for 2026

ScrapingBee published a guide comparing eight Python web scraping tools, covering parsers, browser automation, crawling, proxies, and full-stack scraping services.

ScrapingBee

2026-09-03

Puppeteer 25.10.0 adds video-stream screen recording and Firefox 155.0 support

Puppeteer released version 25.10.0 of puppeteer-core, introducing a video-stream-based screen recording feature via page.record() and rolling to Firefox 155.0.

puppeteer/puppeteer releases

Stagehand Python adds WebMCP tool support inside iframes

Stagehand Python's latest dev release enables discovery and invocation of WebMCP tools registered inside iframes by routing calls through the main CDP session and all adopted OOPIF sessions.

browserbase/stagehand releases

Apify streamlines Actor creation with one-click Git repo setup

Apify now lets users select a Git provider during Actor setup, automatically creating a repository, pushing template code, and configuring builds on every push.

Apify

ScrapingBee explains CAPTCHA solvers and when to avoid them

ScrapingBee published an article defining CAPTCHA solvers, how they automate challenge responses, and when developers should skip them in scraping workflows.

ScrapingBee

ScrapingBee explains how CodeWhale gives AI agents live web access

ScrapingBee publishes an article detailing how CodeWhale enables AI agents to access live web data by integrating web scraping tools, MCP servers, or framework tools.

ScrapingBee

ScrapingBee ranks top proxies for Amazon scraping in 2026

ScrapingBee published a guide evaluating residential, ISP, and mobile proxies for Amazon scraping, noting that Amazon's anti-bot updates can render previously effective proxies obsolete.

ScrapingBee

2026-09-02

Stagehand SDK exposes Browserbase Search and Fetch across TypeScript, Python, and Go

Stagehand SDK version 4.1.0a0.dev1494 adds Browserbase Search and Fetch APIs to its TypeScript and Python facades, with equivalent Go support via the existing HTTP transport.

browserbase/stagehand releases

Zyte tests Claude Fable 5.1 and GLM-5.3-Flash in a live extraction benchmark

Zyte published a personal benchmark comparing Claude Fable 5.1 and GLM-5.3-Flash on real extraction tasks, revealing that the GLM model matched a model the author had previously encountered.

Zyte

Zyte analysis finds 75% of top sites use robots.txt, but few name specific crawlers

Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.

Zyte

ScrapingBee publishes guide to AI agent web scraping with wigolo and MCP

ScrapingBee released a guide covering how to equip AI agents with web scraping capabilities using wigolo and the Model Context Protocol.

ScrapingBee

ScrapingBee publishes guide on scraping website text for LLM training

ScrapingBee released a tutorial covering how to extract all text from a website for use in LLM training pipelines.

ScrapingBee

2026-09-01

Zyte adds CDP support for browser automation on its infrastructure

Zyte has introduced Chrome DevTools Protocol (CDP) support, allowing users to run browser automation scripts on Zyte's managed infrastructure.

Zyte

Apify launches Apartments.com scraper for rental market analysis

Apify released a tool to scrape Apartments.com listings at scale and integrate them with ChatGPT for interactive market reports.

Apify

2026-08-31

SerpApi countersues Reddit over API access restrictions

SerpApi has filed counterclaims against Reddit in response to Reddit's lawsuit, alleging broken promises of an open internet and unfair API pricing.

SerpApi

Scrapfly reviews nine Scrapy extensions and middlewares for 2026

Scrapfly tested nine Scrapy extensions and middlewares against Scrapy 2.18, covering rendering, TLS fingerprints, proxies, extraction, shared queues, and deployment.

Scrapfly

Scrapfly publishes 2026 diagnostic guide for blocked Scrapy spiders

Scrapfly released a walkthrough covering nine mechanisms that can block a Scrapy spider, from IP reputation to CAPTCHA, with evidence and mitigation steps for each.

Scrapfly

Scrapingdog launches MCP Server for AI-powered web scraping

Scrapingdog has introduced an MCP Server that integrates its web scraping and data extraction capabilities directly into AI-powered applications.

Scrapingdog

SerpApi publishes roundup of best web scraping tools for 2026

SerpApi has published a guide covering popular web scraping tools, from open-source frameworks like Scrapy and Crawlee to commercial scraping platforms and search APIs.

SerpApi

Builder spotlight: Goldmine automated outreach and won on Apify

Apify profiles a developer who built a LinkedIn scraper for his team and later became a top-rated Apify developer, winning an EMEA prize in the Apify $1 Million Challenge.

Apify

Zenrows publishes practical guide on web data for LLM fine-tuning

Zenrows outlines how to source clean web data for LLM fine-tuning, warning that a single bad extraction can measurably degrade a small seed set.

Zenrows

Zenrows blog walks through scraping 2026 FIFA World Cup data across three access patterns

Zenrows published a tutorial demonstrating how to scrape 2026 FIFA World Cup data from a JSON endpoint behind JavaScript, an API with per-session JWT, and server-rendered HTML using Python and its own scraping API.

Zenrows

Crawlbase shows how to build a web scraping pipeline with Zapier using async callbacks

Crawlbase published a guide explaining how to split web retrieval from Zapier's execution window by dispatching async crawls with a callback and catching the finished page in a second Zap.

Crawlbase

2026-08-30

Google introduces /goto redirect URLs in Search, SerpApi works on resolution

Google is rolling out new /goto redirect URLs across Search, and SerpApi is actively resolving affected links as the implementation evolves.

SerpApi

Stagehand Python 4.0.3a0.dev1483 adds stagehand_facade tool surface for eval benchmarking

A dev release of Stagehand Python introduces a facade tool surface that allows evals to benchmark the exact byte-identical interface shipped to Claude Code, Codex, and Pi integrations.

browserbase/stagehand releases

SerpApi publishes guide on scraping Walmart product reviews with its dedicated API

SerpApi released a tutorial showing how to use its Walmart Product Reviews API to extract ratings, review text, feedback counts, and reviewer details in structured JSON and export to CSV.

SerpApi

2026-08-28

SerpApi explains HTTP 429 errors and rate-limit best practices

SerpApi published a guide covering the causes of HTTP 429 errors, how to read Retry-After headers, and strategies for retrying requests without escalating blocking.

SerpApi

Zenrows integrates with smolagents to give AI agents production-grade web access

A tutorial shows how to swap smolagents' plain-requests VisitWebpageTool for a Zenrows-backed tool that bypasses bot checks.

Zenrows

Zenrows MCP brings live scraping to Cursor's AI editor

Zenrows released an MCP integration that lets Cursor users scrape JavaScript-rendered and bot-protected sites directly from the editor.

Zenrows

ScrapingBee publishes guide on MCP servers for web scraping, emphasizing control over data

ScrapingBee explains how a Model Context Protocol (MCP) server can improve scraping agent performance by keeping context lean and avoiding page bloat.

ScrapingBee

Crawlbase introduces Web Bot Auth as an identity layer for bots, coinciding with pay-per-crawl pricing

Crawlbase ships Web Bot Auth, a specification that treats bot identity as a verifiable signature rather than a self-declared claim, on the same day it launches pay-per-crawl pricing.

Crawlbase

2026-08-27

Scrapfly publishes 2026 guide to e-commerce scraping tools

Scrapfly's blog post surveys nine e-commerce scraping tools for developers, covering retrieval, extraction, browser automation, and discovery layers.

Scrapfly

2026-08-25

Zyte argues EU AI scraping guidelines rely on outdated robots.txt standard

Zyte publishes a blog post arguing that Europe's new generative AI scraping guidelines, which lean on the robots.txt protocol, will harm users and entrench monopolies.

Zyte

Zyte pitches WebFetch as a drop-in replacement for coding agents' built-in fetch tool

Zyte published a blog post arguing that its WebFetch CLI tool outperforms the default webfetch tool in coding agents for research and coding workflows.

Zyte

2026-08-24

Scrapfly ranks six open-source YouTube scrapers by job, flags two failures

Scrapfly published a blog post comparing six open-source YouTube scrapers, including a GitHub snapshot from August 11, 2026, and noting two projects that failed in their tests.

Scrapfly

Scrapfly publishes guide to scraping Target.com via Redsky API and bypassing PerimeterX

Scrapfly released a blog post detailing how to extract product and pricing data from Target.com using the internal Redsky API while handling store-keyed prices and PerimeterX anti-bot defenses.

Scrapfly

Zyte report reveals retailers as second most aggressive sector in blocking AI crawlers

Zyte published a blog post detailing how major retail marketplaces deploy anti-bot technology to block AI crawlers and automated data extraction.

Zyte

Crawlbase publishes technical guide on scaling headless browser fleets to 10,000 concurrent sessions

Crawlbase's blog post details the architecture and capacity planning required to run 10,000 concurrent Playwright sessions across roughly 100 nodes.

Crawlbase

2026-08-20

Zyte research argues web scraping faces pricing barriers, not outright blocking

Zyte's State of Web Access report, discussed in an interview with the researcher, finds that new economic barriers are making web scraping more difficult rather than technical blocks.

Zyte

2026-08-17

Zyte report finds fashion websites among the most heavily defended against scraping

A Zyte blog post examines the anti-bot and blocking measures used by fashion e-commerce sites, finding they are some of the most aggressively protected on the web.

Zyte

2026-08-14

Scrapfly publishes tutorial on scraping Skyscanner flight prices with Python

Scrapfly released a blog post showing how to extract flight data from Skyscanner by constructing deep-link URLs and capturing itinerary JSON from the rendered page.

Scrapfly

Scrapfly publishes guide on scraping Airbnb listings and prices

Scrapfly released a blog post detailing how to scrape Airbnb search results, listing details, prices, reviews, and availability using Python and its own scraping platform.

Scrapfly

Scrapfly ranks five open-source LinkedIn scrapers on GitHub by auth model and ban risk

Scrapfly published a blog post evaluating five open-source LinkedIn scraping repositories on GitHub, ranking them by authentication approach, maintenance status, and real-world blocking risk as of August 2026.

Scrapfly

Zyte blog post examines how AI and web scraping turn scattered personal data into security risks

Domagoj Marić explores the intersection of AI, web scraping, and OSINT to show how fragmented personal data is assembled into profiles, scams, and security threats at Extract Summit.

Zyte

2026-08-13

Scrapfly publishes guide on scraping Lowe's product data and bypassing Akamai

Scrapfly released a blog post detailing how to scrape Lowe's product, price, search, and store location data using embedded page state and their maintained Python scraper.

Scrapfly

Scrapfly Publishes Guide to Scraping DigiKey Data Past Cloudflare

Scrapfly's blog post details how to scrape DigiKey pricing, stock, and parametric specs while navigating its Cloudflare challenge, and compares this approach to using the official API v4.

Scrapfly

Zyte Audit Reveals Industry-Specific Bot Access Policies

Zyte published a large-scale audit of web access controls showing how different industries enforce different policies toward bots.

Zyte

Crawlbase details the infrastructure behind 8,000 CAPTCHAs per second

Crawlbase published a blog post explaining the throughput math, Go-based control plane, and scaling challenges required to solve 8,000 CAPTCHAs per second.

Crawlbase

2026-08-11

Scrapfly ranks six open-source Instagram scrapers with notes on auth and ban risk

Scrapfly published a comparison of six open-source Instagram scrapers for 2026, covering auth models, ban risk, and maintenance status.

Scrapfly

Zyte blog profiles case study of $70 AI-coded app replacing $5,000 platform

Zyte published a blog post detailing how developer Fran Muñoz used AI coding and specification-driven development to build a production app that replaced a costly platform.

Zyte

Scrapfly compares five MCP servers for web scraping and browser automation

Scrapfly published a blog post comparing five MCP servers by their capabilities in protected-site scraping, browser control, debugging, static fetching, and cross-browser automation.

Scrapfly

Scrapfly compares six modern command-line tools as alternatives to cURL and Wget

A blog post from Scrapfly evaluates HTTPie, aria2, and other tools that address specific limitations of cURL and Wget, including a managed fetch tool for blocked requests.

Scrapfly

Scrapfly compares 8 Python HTTP clients for web scraping in 2026

Scrapfly published a blog post evaluating eight Python HTTP clients on async support, HTTP/2, HTTP/3, TLS impersonation, and maintenance, with runnable examples.

Scrapfly

Scrapfly ranks 7 lead scraping tools for 2026

Scrapfly published a ranked list of seven lead scraping tools covering no-code extensions and production APIs, with honest assessments of each tool's limitations.

Scrapfly

Study finds only 34.5% of free proxies work, thousands tamper with traffic

Crawlbase reports on two peer-reviewed studies that tested 640,600 free proxies, finding that just over a third were functional and many altered traffic.

Crawlbase

2026-08-07

Scrapfly publishes diagnostic guide for browser fingerprint testing tools

Scrapfly released a layer-by-layer guide covering fingerprint and bot detection tools, explaining what detectable results mean and how to fix each leak.

Scrapfly

Vacation rental intelligence platform scales to 1 billion monthly crawl requests with Crawlbase Enterprise Crawler

Crawlbase published a case study detailing how a vacation rental intelligence platform processed 5.52 billion requests in six months with 99.96% success using its Enterprise Crawler.

Crawlbase

2026-08-06

Chrome's new navigator.cpuPerformance API opens a fresh fingerprinting vector

Zyte reports that Chrome 152 will expose a navigator.cpuPerformance property, giving sites a new way to fingerprint browsers.

Zyte

2026-08-05

Scrapfly publishes tutorial on scraping Google Jobs with Python

Scrapfly released a blog post showing how to scrape Google Jobs listings using Python and its own scraping platform.

Scrapfly

Zyte publishes largest ever audit of web access control mechanisms

Zyte released a comprehensive audit of how websites regulate programmatic visits, revealing the current state of web access barriers.

Zyte

Zyte publishes tutorial on building custom fetch tools for AI agents with Claude Agent SDK

Zyte released a tutorial showing how to create a custom fetch tool using the Claude Agent SDK to help AI agents extract structured data from the web.

Zyte

2026-08-04

Zyte releases scrapy-spidey-sense, a preflight CLI for Scrapy projects

Zyte has open-sourced a command-line tool that performs static analysis on Scrapy projects before a crawl begins, scoring production-readiness and linking findings to fixes.

Zyte

2026-08-03

Scrapfly publishes guide on scraping Google Play app reviews and metadata with Python

Scrapfly released a tutorial showing how to extract full Google Play app reviews, ratings, and metadata using Python, bypassing the typical few-hundred-review limit of free libraries.

Scrapfly

Scrapfly ranks four open-source proxy scrapers still viable in 2026

A blog post from Scrapfly filters the crowded open-source proxy tool landscape down to four actively maintained scrapers and checkers worth using this year.

Scrapfly

Scrapfly ranks 7 AI browser agents for production scraping in 2026

Scrapfly published a ranked guide to the best AI browser agents for automation and scraping, evaluating them on production stability and anti-blocking capability rather than demo performance.

Scrapfly

Scrapfly publishes guide on scraping Marriott hotel data through Akamai defenses

Scrapfly released a tutorial covering how to extract Marriott hotel prices and availability using Python, including bypassing Akamai bot protection.

Scrapfly

Zyte blog post explores rendering JavaScript pages with Playwright and Scrapy

Zyte published a guide on using Playwright to render dynamic content within a Scrapy workflow.

Zyte

2026-07-29

Crawlbase argues AI agent failures are infrastructure failures, not code problems

Crawlbase publishes a blog post claiming that most AI agent failures stem from infrastructure issues like Markdown normalization, retrieval circuit breakers, and storage-backed memory.

Crawlbase

2026-07-27

Scrapfly publishes guide to scraping Kayak flight data with its SDK

Scrapfly released a blog post walking through the process of scraping Kayak flight search results using its own SDK, covering JavaScript rendering and parsing internal poll JSON.

Scrapfly

Zyte launches 'Modern Scrapy for experienced developers' tutorial series

Zyte published the first part of a new blog series aimed at experienced developers building production-ready Scrapy projects.

Zyte

2026-07-24

Scrapfly publishes guide on scraping RS-Online for product data

Scrapfly released a tutorial on extracting pricing, stock, specifications, and datasheet links from RS-Online's North American listings and product pages.

Scrapfly

Scrapfly compares Browser Use and Playwright for web scraping

Scrapfly published a blog post comparing Browser Use and Playwright, covering architectural differences, speed and cost tradeoffs, silent failure risks, and a hybrid approach for production scraping.

Scrapfly

Scrapfly publishes guide on bypassing AWS WAF Bot Control for web scraping

Scrapfly released a blog post detailing how AWS WAF Bot Control detects scrapers across five layers and how to bypass it using their Scrapfly ASP product.

Scrapfly

Scrapfly rounds up top open-source Facebook Marketplace scrapers on GitHub for 2026

Scrapfly published a blog post listing the five best open-source Facebook Marketplace scrapers on GitHub as of 2026, along with repos to avoid.

Scrapfly