Tags: scraper*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Botasaurus is an all-in-one web scraping framework designed to help developers build undetectable scrapers that can bypass sophisticated bot detection systems like Cloudflare, Datadome, and BrowserScan. It simplifies the development process by providing high-level abstractions for browser automation, humane HTTP requests, and general data tasks. Key features include human-like mouse movements, browser-based fetch requests to significantly reduce proxy costs, and built-in utilities for caching, sitemap processing, and data cleaning.
    Main topics:
    - Bypassing Cloudflare WAF and Turnstile CAPTCHAs.
    - Creating UI-based scrapers for non-technical end-users.
    - Converting scrapers into standalone desktop applications.
    - Scaling scraping infrastructure using Docker and Kubernetes.
    - Cost-efficient proxy management and bandwidth reduction strategies.
  2. This post demonstrates how to use Cloudflare's Browser Rendering to easily crawl entire websites, even those with complex JavaScript. It simplifies web crawling by rendering pages with a single API call, bypassing the need for headless browsers and enabling efficient data extraction for tasks like SEO monitoring and content archiving.
  3. An open source project called Scrapling is gaining traction with AI agent users who want their bots to scrape sites without permission, and is being used to bypass anti-bot systems like Cloudflare Turnstile. Cloudflare is actively working to counter these efforts.
  4. Cloudflare converts HTML to Markdown on the fly when an AI agent requests it via the `Accept: text/markdown` header.
  5. Google Chrome is testing **WebMCP**, a new system to help AI agents interact with websites more efficiently. Currently, AI struggles with websites, often relying on slow and unreliable methods. WebMCP lets websites offer AI tools directly through a browser API, potentially lowering costs and speeding up development. It works alongside existing AI protocols like Anthropic’s MCP and focuses on improving how AI assists *with* human web use, not replacing backend systems. Essentially, it's aiming to be a standard way for AI to "talk" to websites.
  6. Fast, secure web scraping for Python.
    2025-12-28 Tags: , , , , by klotz
  7. The internet's new standard, RSL, is a clever fix for a complex problem, and it just might give human creators a fighting chance in the AI economy.
    2025-09-11 Tags: , , , , by klotz
  8. An open source web crawler that searches the internet. It's a minimal, real-time web search CLI that searches the internet for you. Enter a query and get search results as JSON (title, url, published_date), sorted by recency.
  9. Extensions load unknown sites into invisible Windows. What could go wrong?
  10. Extract data from websites in LLM ready JSON or CSV format. Crawl or Scrape entire website with Website Crawler
    2025-09-05 Tags: , , , , , , by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "scraper"

About - Propulsed by SemanticScuttle