shengqiangzhang/examples-of-web-crawlers

HTML

一些非常有趣的python爬虫例子,对新手比较友好,主要爬取淘宝、天猫、微信、微信读书、豆瓣、QQ等网站。(Some interesting examples of python crawlers that are friendly to beginners. )

crawlerspidertaobaotmallexamplepythonseleniumpyquerystockfundmultithreadingagent-pool
Star 增长趋势
Star
14.7k
Forks
3.8k
周增长
+-1
Issues
19
5k10k
2019年3月2021年9月2024年3月2026年9月
相关仓库
firecrawl/firecrawl

The context API to search, scrape, and interact with the web at scale. 🔥

TypeScriptnpmGNU Affero General Public License v3.0aicrawler
firecrawl.dev
178.1k9.7k
D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

PythonPyPIlibraryBSD 3-Clause "New" or "Revised" Licensecrawlercrawling
scrapling.readthedocs.io/en/latest/
79.4k8k
scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

PythonPyPIlibraryBSD 3-Clause "New" or "Revised" Licensepythonscraping
scrapy.org
64.3k11.9k
NaiboWang/EasySpider

A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。

JavaScriptnpmappGNU Affero General Public License v3.0code-freecrawler
easyspider.net
44.5k5.4k
iawia002/lux

👾 Fast and simple video download library and CLI tool written in Go

GoGo ModulesMIT Licensedownloadergo
31.7k3.3k
ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

PythonPyPIMIT Licensescrapingscraping-python
scrapegraphai.com
30.7k3.1k
feder-cr/Jobs_Applier_AI_Agent_AIHawk

Open source AI job application toolkit in Python: generate a resume and cover letter tailored to each job posting, and drive a stealth browser from any AI client over MCP.

GNU Affero General Public License v3.0automationpython
30.3k4.6k
apify/crawlee

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

TypeScriptnpmlibraryApache License 2.0web-scrapingweb-crawling
crawlee.dev
25.7k1.7k
gocolly/colly

Elegant Scraper and Crawler Framework for Golang

GoGo ModuleslibraryApache License 2.0golangscraper
go-colly.org
25.5k1.9k
jhao104/proxy_pool

Python ProxyPool for web spider

PythonPyPIMIT Licensecrawlerproxy
jhao104.github.io/proxy_pool/
23.7k5.4k
Evil0ctal/Douyin_TikTok_Download_API

🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。

PythonPyPIlibraryApache License 2.0pythonpywebio
douyin.wtf
20k2.8k
projectdiscovery/katana

A next-generation crawling and spidering framework.

GoGo ModulescliMIT Licensecrawlerweb-spider
17.4k1.2k