ORC
Apache ORC - the smallest, fastest columnar storage for Hadoop workloads
About
ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.
Open Source Health
- Stars
- 770
- Forks
- 517
- License
- Apache-2.0
- Last commit
- 1 months ago
Alternatives to ORC
Zuiby Brim DataZui is a powerful desktop application for exploring and working with data. The official front-end to the Zed lake.- Servicecomb Mesherby ApacheA high performance service mesh implementation written in go
Dubboby ApacheThe java implementation of Apache Dubbo. An RPC and microservice framework.- Incubator Kie Optaplanner Quickstartsby ApacheOptaPlanner quick starts for AI optimization: many use cases shown in many different technologies.
- Cordova Plugin Screen Orientationby ApacheCordova Screen Orientation plugin
Resources & Links
Related Categories
Using ORC?
Track its cost next to the rest of your stack and get a reminder before it renews.
Add to my stackVendor
Apache
Publisher of Apache Jena, Fuseki and Apache Kafka
