martinsbalodis/web-scraper-chrome-extension

JavaScript

대체: OctoparseParseHub

Web data extraction tool implemented as chrome extension

web-scrapingchrome-extensiondata-extractionjavascript
스타 성장
스타
1.4k
포크
438
주간 성장
+0
이슈
0
5001k
2023년 1월2024년 3월2025년 6월2026년 9월
아티팩트npm
README

Web Scraper

Web Scraper is a chrome browser extension built for data extraction from web pages. Using this extension you can create a plan (sitemap) how a web site should be traversed and what should be extracted. Using these sitemaps the Web Scraper will navigate the site accordingly and extract all data. Scraped data later can be exported as CSV.

Install the extension from [Chrome store] chrome-store

Features

  1. Scrape multiple pages
  2. Sitemaps and scraped data are stored in browsers local storage or in CouchDB
  3. Multiple data selection types
  4. Extract data from dynamic pages (JavaScript+AJAX)
  5. Browse scraped data
  6. Export scraped data as CSV
  7. Import, Export sitemaps
  8. Depends only on Chrome browser

Help

Documentation and tutorials are available on webscraper.io webscraper.io

Ask for help, submit bugs, suggest features on [google groups] google-groups

Submit bugs and suggest features on [bug tracker] github-issues

Bugs

When submitting a bug please attach an exported sitemap if possible.

License

LGPLv3

Changelog

v0.2

  • Added Element click selector
  • Added Element scroll down selector
  • Added Link popup selector
  • Improved table selector to work with any html markup
  • Added Image download
  • Added keyboard shortcuts when selecting elements
  • Added configurable delay before using selector
  • Added configurable delay between page visiting
  • Added multiple start url configuration
  • Added form field validation
  • Fixed a lot of bugs

v0.1.3

  • Added Table selector
  • Added HTML selector
  • Added HTML attribute selector
  • Added data preview
  • Added ranged start urls
  • Fixed bug which made selector tree not to show on some operating systems
관련 저장소
firecrawl/firecrawl

The context API to search, scrape, and interact with the web at scale. 🔥

TypeScriptnpmGNU Affero General Public License v3.0aicrawler
firecrawl.dev
178.1k9.7k
D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

PythonPyPIlibraryBSD 3-Clause "New" or "Revised" Licensecrawlercrawling
scrapling.readthedocs.io/en/latest/
79.4k8k
scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

PythonPyPIlibraryBSD 3-Clause "New" or "Revised" Licensepythonscraping
scrapy.org
64.3k11.9k
JCodesMore/ai-website-cloner-template

Clone any website with one command using AI coding agents

JavaScriptnpmskillMIT Licenseaiai-agents
dsc.gg/jcodesmore
34.1k5k
dgtlmoon/changedetection.io

Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock alerts, and website defacement monitoring—all for free or enjoy our SaaS plan!

PythonPyPIApache License 2.0website-monitorwebsite-monitoring
changedetection.io
33.7k2k
CloakHQ/CloakBrowser

Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.

PythonPyPIMIT Licenseanti-detectbot-detection
cloakbrowser.dev
31.3k2.6k
ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

PythonPyPIMIT Licensescrapingscraping-python
scrapegraphai.com
30.7k3.1k
feder-cr/Jobs_Applier_AI_Agent_AIHawk

Open source AI job application toolkit in Python: generate a resume and cover letter tailored to each job posting, and drive a stealth browser from any AI client over MCP.

GNU Affero General Public License v3.0automationpython
30.3k4.6k
apify/crawlee

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

TypeScriptnpmlibraryApache License 2.0web-scrapingweb-crawling
crawlee.dev
25.7k1.7k
Evil0ctal/Douyin_TikTok_Download_API

🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。

PythonPyPIlibraryApache License 2.0pythonpywebio
douyin.wtf
20k2.8k
getmaxun/maxun

🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

TypeScriptnpmlibraryGNU Affero General Public License v3.0automationno-code
maxun.dev
17.4k1.5k
MODSetter/SurfSense

Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9

PythonPyPIOtheraifastapi
surfsense.com
16.1k1.5k