An orchestration platform for the development, production, and observation of data assets.
pip install dagsterpython_modules/dagster/README.md
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
Lean and mean distributed stream processing system written in rust and web assembly. Alternative to Kafka + Flink in one.
Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows, particularly visualizations, into single files, runnable completely in-browser, using Pyodide, DuckDB, Pandas, and Plotly, Matplotlib, etc. Build dashboards, reports, and notebooks that run offline, load fast, and share like a document.
Build data pipelines, the easy way 🛠️
A system for agentic LLM-powered data processing and ETL