Export to Apache Arrow
All results of a query can be exported to an Apache Arrow Table using the to_arrow_table function. Alternatively, results can be returned as a RecordBatchReader using the to_arrow_reader function and results can be read one batch at a time. In addition, relations built using DuckDB’s Relational API can also be exported.
Deprecated The
fetch_arrow_table,fetch_record_batch, andfetch_arrow_readerfunctions are deprecated. Useto_arrow_tableandto_arrow_readerinstead.
Export to an Arrow Table
import duckdbimport pyarrow as pa
my_arrow_table = pa.Table.from_pydict({'i': [1, 2, 3, 4], 'j': ["one", "two", "three", "four"]})
# query the Apache Arrow Table "my_arrow_table" and return as an Arrow Tableresults = duckdb.sql("SELECT * FROM my_arrow_table").to_arrow_table()Export as a RecordBatchReader
import duckdbimport pyarrow as pa
my_arrow_table = pa.Table.from_pydict({'i': [1, 2, 3, 4], 'j': ["one", "two", "three", "four"]})
# query the Apache Arrow Table "my_arrow_table" and return as an Arrow RecordBatchReaderchunk_size = 1_000_000result = duckdb.sql("SELECT * FROM my_arrow_table").to_arrow_reader(chunk_size)
# Loop through the results. A StopIteration exception is thrown when the RecordBatchReader is emptywhile (batch := result.read_next_batch()): # Process a single chunk here print(batch.to_pandas())Export from Relational API
Arrow objects can also be exported from the Relational API. A relation can be converted to an Arrow table using DuckDBPyRelation.to_arrow_table, and to an Arrow record batch reader using DuckDBPyRelation.to_arrow_reader.
import duckdb
# connect to an in-memory databasecon = duckdb.connect()
con.execute('CREATE TABLE integers (i integer)')con.execute('INSERT INTO integers VALUES (0), (1), (2), (3), (4), (5), (6), (7), (8), (9), (NULL)')
# Create a relation from the table and export the entire relation as Arrowrel = con.table("integers")relation_as_arrow = rel.to_arrow_table()
# Calculate a result using that relation and export that result to Arrowres = rel.aggregate("sum(i)").execute()arrow_table = res.to_arrow_table()
# You can also create an Arrow record batch reader from a relationarrow_batch_reader = res.to_arrow_reader()while (batch := arrow_batch_reader.read_next_batch()): # Process a single chunk here print(batch.to_pandas())