Voltar ao ranking
elliotgao2/gain
PythonWeb crawling framework based on asyncio.
pythoncrawlerspiderasynciouvloopaiohttp
Métricas principais
Crescimento de estrelas
Estrelas
2k
Forks
205
Crescimento semanal
—
Issues
1
5001k1.5k2k
out. de 2023mar. de 2024set. de 2024fev. de 2025ago. de 2025jan. de 2026jul. de 2026
ArtefatosPyPI
pip install gainREADME
gain
Async web crawling framework for everyone.
Built on asyncio, aiohttp, and lxml/pyquery. Declare items and
parsers; gain handles the concurrency, retries, and persistence.
Install
pip install gain
Linux users can opt into uvloop for an extra speed bump:
pip install "gain[uvloop]"
Requires Python 3.10+.
Quickstart
import aiofiles
from gain import Css, Item, Parser, Spider
class Post(Item):
title = Css(".entry-title")
content = Css(".entry-content")
async def save(self):
async with aiofiles.open("scrapinghub.txt", "a+") as f:
await f.write(self.results["title"] + "\n")
class MySpider(Spider):
concurrency = 5
headers = {"User-Agent": "Google Spider"}
start_url = "https://blog.scrapinghub.com/"
parsers = [
Parser(r"https://blog.scrapinghub.com/page/\d+/"),
Parser(r"https://blog.scrapinghub.com/\d{4}/\d{2}/\d{2}/[a-z0-9\-]+/", Post),
]
MySpider.run()
Run it:
python spider.py
XPath parsers
from gain import Css, Item, Parser, Spider, XPathParser
class Post(Item):
title = Css(".breadcrumb_last")
async def save(self):
print(self.title)
class MySpider(Spider):
start_url = "https://mydramatime.com/europe-and-us-drama/"
concurrency = 5
headers = {"User-Agent": "Google Spider"}
parsers = [
XPathParser('//span[@class="category-name"]/a/@href'),
XPathParser('//div[contains(@class, "pagination")]/ul/li/a[contains(@href, "page")]/@href'),
XPathParser('//div[@class="mini-left"]//div[contains(@class, "mini-title")]/a/@href', Post),
]
proxy = "https://localhost:1234"
MySpider.run()
How it works
┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐
│ start_url │ ─▶ │ Parser │ ─▶ │ Item │ ─▶ │ save() │
│ │ │ (follow) │ │ (extract) │ │ (persist) │
└────────────┘ └────────────┘ └────────────┘ └────────────┘
▲ │
└──────────── new urls ────────────────┘
- Spider kicks off from
start_urlunder a concurrency budget. - Parsers either follow (one argument) — discovering more URLs to
queue — or extract (two arguments) — instantiating an
Itemfrom each matching page. - Items use
Css/Xpath/Regexselectors to pull fields out of HTML. save()is your async hook to persist results — write a file, push to a queue, insert into a database.
Examples
See the example/ directory for runnable scripts against
Scrapinghub, V2EX, and Sciencenet.
Development
git clone https://github.com/elliotgao2/gain.git
cd gain
uv sync # install deps into .venv
uv run pytest # run tests
uv run ruff check . # lint
We use uv for packaging and ruff for lint + format. Install the pre-commit hooks:
uv run pre-commit install
Contributing
Pull requests are welcome. For non-trivial changes, please open an issue
first to discuss. Make sure pytest and ruff check pass before
submitting.
License
MIT © Elliot Gao
Repositórios relacionados