返回排行榜

data-engineering-community/data-engineering-wiki

CSSdataengineering.wiki

The best place to learn data engineering. Built and maintained by the data engineering community.

datadatabasedata-engineeringdata-engineersqldata-modelingdata-pipelinesetl
Star 增长趋势
Star
2k
Forks
242
周增长
Issues
22
5001k1.5k2k
2023年1月2024年3月2025年5月2026年7月
制品库npmnpm install data-engineering-wiki
README

GitHub repo size Website

Data Engineering Wiki

The best place to learn data engineering. Built and maintained by the data engineering community.

Subreddit subscribers

What's inside?

A collection of notes that are connected organically but loosely organized into the following categories:

  1. Concepts: Concepts related to Data Engineering.
  2. FAQ: Frequently asked questions about Data Engineering.
  3. Guides: Understand how to make Data Engineering decisions.
  4. Tools: Commonly used tools for Data Engineering.
  5. Learning Resources: Learn Data Engineering with resources recommended by the Data Engineering community.

Sponsors

The Data Engineering Wiki is an CC0-1.0-licensed open source project with its ongoing development made possible entirely by the support of these awesome backers. If you'd like to join them, please consider sponsoring the Data Engineering Wiki's development.

Gold Sponsors

DataDriven — Data Engineer Interview Practice

How to run it locally

The wiki can be used offline and can be used as-is or incorporated into your own personal knowledge management system. It is built to be used with Obsidian (free, no affiliation) but is compatible with other tools as well such as Foam or Roam Research.

  1. Download this GitHub repository.
  2. Download the free Obsidian desktop app.
  3. Run the Obsidian app and choose Open folder as vault, click Open.
  4. In the file browser, choose the folder where you downloaded the GitHub repository, click Open.

See Obsidian help for questions on using Obsidian.

Contributing

There are many different ways to contribute to the wiki's development. If you're interested, check out our contributing guidelines to learn how you can get involved.

Thank you to all of our contributors who shared their data engineering knowledge!

相关仓库
D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

PythonPyPIBSD 3-Clause "New" or "Revised" Licensecrawlercrawling
scrapling.readthedocs.io/en/latest/
70.6k7k
Asabeneh/30-Days-Of-Python

The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

PythonPyPI30-days-of-pythonpython
68.8k12.8k
run-llama/llama_index

LlamaIndex is the leading document agent and OCR platform

PythonPyPIMIT Licenseagentsapplication
developers.llamaindex.ai
51k7.8k
TanStack/query

🤖 Powerful asynchronous state management, server-state utilities and data fetching for the web. TS/JS, React Query, Solid Query, Svelte Query and Vue Query.

TypeScriptnpmMIT Licensereacthooks
tanstack.com/query
50k4k
metabase/metabase

The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:

ClojureOtheranalyticsbusinessintelligence
metabase.com
48.3k6.7k
DataExpert-io/data-engineer-handbook

This is a repo with links to everything you'd ever want to learn about data engineering

Jupyter Notebookapachesparkawesome
42.3k7.9k
SheetJS/sheetjs

📗 SheetJS Spreadsheet Data Toolkit -- New home https://git.sheetjs.com/SheetJS/sheetjs

Apache License 2.0xlsxexcel
sheetjs.com
36.3k7.9k
vercel/swr

React Hooks for Data Fetching

TypeScriptnpmMIT Licensereacthook
swr.vercel.app
32.4k1.4k
mendableai/firecrawl

🔥 Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API.

TypeScriptnpmGNU Affero General Public License v3.0aicrawler
firecrawl.dev
29.5k2.5k
sinaptik-ai/pandas-ai

Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.

PythonPyPIOtherllmpandas
pandas-ai.com
23.7k2.3k
PrefectHQ/prefect

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

PythonPyPIApache License 2.0pythonworkflow
prefect.io
23.5k2.4k
airbytehq/airbyte

Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

PythonPyPIOtherdatapipeline
airbyte.com
21.7k5.3k