Volver al ranking

bruin-data/ingestr

Gogetbruin.com/docs/ingestr/

ingestr is a CLI tool to copy data between any databases with a single command seamlessly.

bigquerycopy-databasedata-ingestiondata-integrationdata-pipelineduckdbingestion-pipelinemssqlpostgresqlsnowflake
Crecimiento de estrellas
Estrellas
3.8k
Forks
143
Crecimiento semanal
Issues
4
2k3k
feb 2024nov 2024sept 2025jul 2026
ArtefactosGo Modulesgo get github.com/bruin-data/ingestr
README

Copy data from any source to any destination without any code


ingestr is a command-line app that allows you to ingest data from any source into any destination using simple command-line flags, no code necessary.

  • ✨ copy data from your database into any destination
  • ➕ incremental loading: append, merge or delete+insert
  • 🐍 single-command installation

ingestr takes away the complexity of managing any backend or writing any code for ingesting data, simply run the command and watch the data land on its destination.

MongoDB to Postgres benchmark

Installation

You can install ingestr using the install script:

curl -LsSf https://getbruin.com/install/ingestr | sh

Alternatively, you can install it with pip:

pip install ingestr

The pip package can also be used from Python. Install the SDK extra for Python data ingestion:

pip install 'ingestr[sdk]'

Python rows, generators, and DataFrames are sent to the bundled ingestr binary as Arrow IPC streams by default:

import ingestr

ingestr.ingest(
    [{"id": 1, "name": "Ada"}, {"id": 2, "name": "Grace"}],
    dest_uri="duckdb:///tmp/warehouse.duckdb",
    dest_table="main.people",
)

DataFrames and yielded data use the same Arrow stream transport:

ingestr.ingest(df, dest_uri="duckdb:///tmp/warehouse.duckdb", dest_table="main.events")

def events():
    yield [{"id": 1, "event": "signup"}]
    yield [{"id": 2, "event": "purchase"}]

ingestr.ingest(events, dest_uri="postgresql://...", dest_table="public.events")

For push-style code, omit the data argument and use ingest as a context manager. The context value accepts the same shapes as ingestr.ingest(data, ...):

with ingestr.ingest(dest_uri="postgresql://...", dest_table="public.events") as ingest:
    for response in client.list_events():
        ingest(response["items"])

For very large already-materialized data, use the existing mmap Arrow IPC file transport:

ingestr.ingest(df, dest_uri="duckdb:///tmp/warehouse.duckdb", dest_table="main.events", transport="mmap")

For full CLI pass-through, use ingestr.run(["ingest", "--source-uri", "...", "--dest-uri", "...", "--source-table", "..."]), or ingestr.run_cli(...) for keyword arguments that map to CLI flags.

Quickstart

ingestr ingest \
    --source-uri 'postgresql://admin:admin@localhost:8837/web?sslmode=disable' \
    --source-table 'public.some_data' \
    --dest-uri 'bigquery://<your-project-name>?credentials_path=/path/to/service/account.json' \
    --dest-table 'ingestr.some_data'

That's it.

This command:

  • gets the table public.some_data from the Postgres instance.
  • uploads this data to your BigQuery warehouse under the schema ingestr and table some_data.

Documentation

You can see the full documentation here.

Community

Join our Slack community here.

Contributing

Pull requests are welcome. However, please open an issue first to discuss what you would like to change. We maybe able to offer you help and feedback regarding any changes you would like to make.

[!NOTE] After cloning ingestr make sure to run make setup to install githooks.

Supported sources & destinations

Source Destination CDC
Databases
AWS Athena -
Apache Iceberg - -
AWS Redshift -
Cassandra -
ClickHouse -
Couchbase - -
CrateDB -
Databricks -
DuckDB -
DynamoDB -
Elasticsearch -
Google BigQuery -
GCP Spanner - -
IBM Db2 - -
InfluxDB - -
Kafka - -
Local CSV file -
MaxCompute -
Microsoft Fabric -
Microsoft OneLake - -
Microsoft SQL Server
MongoDB
MotherDuck -
MySQL
Oracle - -
PlanetScale
Postgres
RabbitMQ - -
SAP Hana - -
Snowflake -
Socrata - -
SQLite -
StarRocks -
Synapse - -
Trino -
Vitess
Platforms
Adjust - -
Adapty - -
Airtable - -
Allium - -
Amazon Kinesis - -
Anthropic - -
API-Football - -
AppsFlyer - -
Apple Ads - -
Apple App Store - -
Applovin - -
Applovin Max - -
Asana - -
Attio - -
Azure Data Lake Storage Gen2 -
BallDontLie FIFA - -
Braze - -
Bruin - -
Chess.com - -
ClickUp - -
Cursor - -
Docebo - -
Dune - -
Facebook Ads - -
Fireflies - -
Fluxx - -
football-data.org - -
Frankfurter - -
Freshdesk - -
FundraiseUp - -
G2 - -
GitHub - -
GitLab - -
Google Ads - -
Google Analytics - -
Google Cloud Storage (GCS) -
Google Sheets - -
Gorgias - -
Granola - -
Hostaway - -
HubSpot - -
Indeed - -
Intercom - -
Internet Society Pulse - -
Jira - -
JobTread - -
Klaviyo - -
Linear - -
LinkedIn Ads - -
Mailchimp - -
Mixpanel - -
Monday - -
Notion - -
Paddle - -
Personio - -
PhantomBuster - -
Pinterest - -
Pipedrive - -
Plus Vibe AI - -
PostHog - -
Primer - -
QuickBooks - -
Reddit Ads - -
RevenueCat - -
S3 -
Salesforce - -
SFTP - -
SendGrid - -
Shopify - -
Slack - -
Smartsheet - -
Snapchat Ads - -
Solidgate - -
Square - -
Stripe - -
SurveyMonkey - -
TikTok Ads - -
Trello - -
Trustpilot - -
Twilio - -
Wise - -
Zendesk - -
Zoom - -

Feel free to create an issue if you'd like to see support for another source or destination.

License

ingestr is source-available under the Functional Source License 1.1, with Apache 2.0 as the future license. You can use ingestr freely for internal production use, development, testing, education, research, and professional services. You cannot use ingestr to offer a competing commercial ingestion, ELT, connector, or managed data pipeline product/service.

Each version becomes Apache 2.0 two years after release.

Repositorios relacionados
hasura/graphql-engine

Blazing fast, instant realtime GraphQL APIs on all your data with fine grained access control, also trigger webhooks on database events.

TypeScriptnpmApache License 2.0graphqlgraphql-server
hasura.io
32k2.9k
getredash/redash

Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard and share your data.

PythonPyPIBSD 2-Clause "Simplified" Licenseredashpython
redash.io
28.7k4.6k
beekeeper-studio/beekeeper-studio

Modern and easy to use SQL client for MySQL, Postgres, SQLite, SQL Server, and more. Linux, MacOS, and Windows.

TypeScriptnpmOtherdatabasesql
beekeeperstudio.io
23.2k1.6k
airbytehq/airbyte

Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

PythonPyPIOtherdatapipeline
airbyte.com
21.7k5.3k
cube-js/cube

📊 Cube Core is open-source semantic layer for AI, BI and embedded analytics

Rustcrates.ioOtheranalyticscube
cube.dev
20.5k2.1k
Canner/WrenAI

GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-language questions into trusted dashboards, charts, and SQL across 20+ data sources, such as BigQuery, Snowflake, PostgreSQL, ClickHouse, Amazon Redshift, Databricks and more.

PythonPyPIOtherbigqueryduckdb
docs.getwren.ai/oss/introduction
16.5k1.9k
googleapis/mcp-toolbox

MCP Toolbox for Databases is an open source MCP server for databases.

GoGo ModulesApache License 2.0genaimcp
mcp-toolbox.dev/documentation/introduction/
16k1.6k
apache/doris

Apache Doris is a real-time analytics and hybrid search database for AI agents.

JavaMavenApache License 2.0olapdatabase
doris.apache.org
15.7k3.9k
tobymao/sqlglot

Python SQL Parser and Transpiler

PythonPyPIMIT Licensetranspilersql
sqlglot.com
9.4k1.2k
growthbook/growthbook

Open Source Feature Flags, Experimentation, and Product Analytics

TypeScriptnpmOtherabtestingstatistics
growthbook.io
8k799
ibis-project/ibis

the portable Python dataframe library

PythonPyPIApache License 2.0pythonimpala
ibis-project.org
6.6k743
cloudquery/cloudquery

Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

GoGo ModulesMozilla Public License 2.0awsgcp
cloudquery.io
6.5k549