Marketplace

By Use Case Marketplace

Compare By Use Case products across curated subcategories and trusted providers

Products
15,460
Subcategories
699

Selected subcategory

Data Ingestion

Compare By Use Case products across curated subcategories and trusted providers

Explore By Use Case with structured category paths, practical filters, and independent product data

Showing 1–12 of 17

Active filtersOpen sourceClear all

MongoDB data stream pipeline tools by YouGov (adopted from MongoDB)

Data IngestionETL & Data IntegrationOpen Source

Bonobo

by Bonobo   —   ETL for Python 3.5+

Extract Transform Load for Python 3.5+

Data IngestionETL & Data IntegrationOpen Source

Gitingest

by coderamp

Replace 'hub' with 'ingest' in any GitHub URL to get a prompt-friendly extract of a codebase

Data IngestionETL & Data IntegrationOpen Source

Megatron's multi-modal data loader

Data IngestionETL & Data IntegrationOpen Source

Goavro

by Linkedin

Goavro is a library that encodes and decodes Avro data.

Data IngestionETL & Data IntegrationOpen Source

Sqlitedict

by Radim Řehůřek

Persistent dict, backed by sqlite3 and pickle, multithread-safe.

Data IngestionETL & Data IntegrationOpen Source

Bigtop

by Apache

Bigtop is an Apache Foundation project for Infrastructure Engineers and Data Scientists looking for comprehensive packaging, testing, and configuration of the leading open source big data components.

Data IngestionETL & Data IntegrationOpen Source

Fastest end-to-end CSV ingestion for Ruby (with C acceleration). SmarterCSV auto-detects formats, applies smart defaults, and returns Rails-ready hashes for seamless use with ActiveRecord, Sidekiq, parallel jobs, and S3 pipelines — even for messy user-uploaded real-world data.

Data IngestionETL & Data IntegrationOpen Source

FileHelpers

by Marcos Meli

The FileHelpers are a free and easy to use .NET library to read/write data from fixed length or delimited records in files, strings or streams

Data IngestionETL & Data IntegrationOpen Source

Go Stash

by Kevin Wan

go-stash is a high performance, free and open source server-side data processing pipeline that ingests data from Kafka, processes it, and then sends it to ElasticSearch.

Data IngestionETL & Data IntegrationOpen Source

Programmatically extract data and apply schemas to unstructured documents across text-based and multi-modal content using Azure AI Foundry, Azure OpenAI, Azure AI Content Understanding, and Cosmos DB.

Data IngestionETL & Data IntegrationOpen Source

Embulk

by Embulk

Embulk: Pluggable Bulk Data Loader.

Data IngestionETL & Data IntegrationOpen Source

Frequently asked questions