ORC
Apache ORC - the smallest, fastest columnar storage for Hadoop workloads
About
ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.
Open Source Health
- Stars
- 770
- Forks
- 517
- License
- Apache-2.0
- Last commit
- 1 months ago
Resources & Links
Related Categories
Vendor
Apache
Publisher of Apache Jena, Fuseki and Apache Kafka
Quick Links
Open Source
More by Apache
Related Products
Zui
Zui is a powerful desktop application for exploring and working with data. The official front-end to the Zed lake.
Top category match
Servicecomb Mesher
A high performance service mesh implementation written in go
More from this vendor
Dubbo
The java implementation of Apache Dubbo. An RPC and microservice framework.
More from this vendor
Incubator Kie Optaplanner Quickstarts
OptaPlanner quick starts for AI optimization: many use cases shown in many different technologies.
More from this vendor
Cordova Plugin Screen Orientation
Cordova Screen Orientation plugin
More from this vendor
