This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.

51 Branches 114 Tags

Name	Name	Last commit message	Last commit date
Latest commit andygrove update avro tests Sep 3, 2022 01f12d9 · Sep 3, 2022 History 4,272 Commits
.github	.github	MINOR: add github action trigger (#3323 )	Sep 1, 2022
benchmarks	benchmarks	Execute LogicalPlans after building for TPCH Benchmarks (#3290 )	Aug 30, 2022
ci/scripts	ci/scripts	Remove old ci directory (#3198 )	Aug 19, 2022
conbench	conbench	Fix typos (Datafusion -> DataFusion) (#1993 )	Mar 12, 2022
datafusion-cli	datafusion-cli	Bump lz4-sys from 1.9.3 to 1.9.4 in /datafusion-cli (#3335 )	Sep 3, 2022
datafusion-examples	datafusion-examples	Upgrade to arrow 21 (#3225 )	Aug 30, 2022
datafusion	datafusion	update avro tests	Sep 3, 2022
dev	dev	MINOR: Add notes on deleting old releases and release candidates from…	Aug 25, 2022
docs	docs	implement `drop view` (#3267 )	Sep 3, 2022
integration-tests	integration-tests	Use code points instead of grapheme clusters for string functions (#3054	Aug 8, 2022
parquet-testing @ ddd8989	parquet-testing @ ddd8989	add window expression stream, delegated window aggregation to aggrega…	May 26, 2021
python	python	docs: update the Python library repository (#3297 )	Aug 30, 2022
testing @ a8f7be3	testing @ a8f7be3	Avro Table Provider (#910 )	Sep 15, 2021
.asf.yaml	.asf.yaml	MINOR: Remove ballista files that are no longer needed (#2589 )	May 22, 2022
.dockerignore	.dockerignore	Move datafusion-cli to new crate (#231 )	May 3, 2021
.editorconfig	.editorconfig	minor: add editor config file (#2224 )	Apr 14, 2022
.github_changelog_generator	.github_changelog_generator	create datafusion 6.0.0, ballista 0.6.0 and python 0.4.0 releases (#1253	Nov 14, 2021
.gitignore	.gitignore	Build against `arrow-ballista` in CI & remove ballista code from this…	May 22, 2022
.gitmodules	.gitmodules	Fix CI (#10 )	Apr 19, 2021
.pre-commit-config.yaml	.pre-commit-config.yaml	ARROW-11180: [Developer] cmake-format pre-commit hook doesn't run	Mar 30, 2021
CHANGELOG.md	CHANGELOG.md	Prepare 9.0.0 release (#2714 )	Jun 10, 2022
CODE_OF_CONDUCT.md	CODE_OF_CONDUCT.md	use prettier to format md files (#367 )	May 24, 2021
CONTRIBUTING.md	CONTRIBUTING.md	separate contributors guide (#3128 )	Aug 13, 2022
Cargo.toml	Cargo.toml	Remove datafusion-data-access crate (#2904 )	Jul 14, 2022
LICENSE.txt	LICENSE.txt	Remove outdated license text left over from arrow repo (#3154 )	Aug 15, 2022
NOTICE.txt	NOTICE.txt	ARROW-5934: [Python] Bundle arrow's LICENSE with the wheels	Jul 15, 2019
README.md	README.md	[minor] add Coverage Status in readme (#3220 )	Aug 22, 2022
header	header	ARROW-259: Use Flatbuffer Field type instead of MaterializedField	Aug 18, 2016
pre-commit.sh	pre-commit.sh	[DataFusion] - Add show and show_limit function for DataFrame (#923 )	Aug 24, 2021
rustfmt.toml	rustfmt.toml	use 2021 edition (#1084 )	Oct 27, 2021

Repository files navigation

DataFusion

DataFusion is an extensible query planning, optimization, and execution framework, written in Rust, that uses Apache Arrow as its in-memory format.

Features

SQL query planner with support for multiple SQL dialects
DataFrame API
Parquet, CSV, JSON, and Avro file formats are supported natively. Custom file formats can be supported by implementing a TableProvider trait.
Supports popular object stores, including AWS S3, Azure Blob Storage, and Google Cloud Storage. There are extension points for implementing custom object stores.

Use Cases

DataFusion is modular in design with many extension points and can be used without modification as an embedded query engine and can also provide a foundation for building new systems. Here are some example use cases:

DataFusion can be used as a SQL query planner and query optimizer, providing optimized logical plans that can then be mapped to other execution engines.
DataFusion is used to create modern, fast and efficient data pipelines, ETL processes, and database systems, which need the performance of Rust and Apache Arrow and want to provide their users the convenience of an SQL interface or a DataFrame API.

Why DataFusion?

High Performance: Leveraging Rust and Arrow's memory model, DataFusion achieves very high performance
Easy to Connect: Being part of the Apache Arrow ecosystem (Arrow, Parquet and Flight), DataFusion works well with the rest of the big data ecosystem
Easy to Embed: Allowing extension at almost any point in its design, DataFusion can be tailored for your specific use case
High Quality: Extensively tested, both by itself and with the rest of the Arrow ecosystem, DataFusion can be used as the foundation for production systems.

DataFusion Community Extensions

There are a number of community projects that extend DataFusion or provide integrations with other systems.

Language Bindings

Integrations

Known Uses

Here are some of the projects known to use DataFusion:

Ballista Distributed SQL Query Engine
Blaze Spark accelerator with DataFusion at its core
CeresDB Distributed Time-Series Database
Cloudfuse Buzz
CnosDB Open Source Distributed Time Series Database
Cube Store
datafusion-tui Text UI for DataFusion
delta-rs Native Rust implementation of Delta Lake
Flock
InfluxDB IOx Time Series Database
qv Quickly view your data
ROAPI
Tensorbase
VegaFusion Server-side acceleration for the Vega visualization grammar

(if you know of another project, please submit a PR to add a link!)

Example Usage

Please see example usage to find how to use DataFusion.

Roadmap

Please see Roadmap for information of where the project is headed.

Architecture Overview

There is no formal document describing DataFusion's architecture yet, but the following presentations offer a good overview of its different components and how they interact together.

(July 2022): DataFusion and Arrow: Supercharge Your Data Analytical Tool with a Rusty Query Engine: recording and slides
(March 2021): The DataFusion architecture is described in Query Engine Design and the Rust-Based DataFusion in Apache Arrow: recording (DataFusion content starts ~ 15 minutes in) and slides
(February 2021): How DataFusion is used within the Ballista Project is described in *Ballista: Distributed Compute with Rust and Apache Arrow: recording

User Guide

Please see User Guide for more information about DataFusion.

Contributor Guide

Please see Contributor Guide for information about contributing to DataFusion.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

DataFusion

Features

Use Cases

Why DataFusion?

DataFusion Community Extensions

Language Bindings

Integrations

Known Uses

Example Usage

Roadmap

Architecture Overview

User Guide

Contributor Guide

About

Releases

Packages

Used by 2.8k

Contributors 741

Languages

License

apache/datafusion

Folders and files

Latest commit

History

Repository files navigation

DataFusion

Features

Use Cases

Why DataFusion?

DataFusion Community Extensions

Language Bindings

Integrations

Known Uses

Example Usage

Roadmap

Architecture Overview

User Guide

Contributor Guide

About

Topics

Resources

License

Code of conduct

Security policy

Stars

Watchers

Forks

Releases

Packages 0

Used by 2.8k

Contributors 741

Languages

Packages