ClickHouse/azureBlobStorageCluster.md at bac948ec0ecd1520ee4f7618699fd29ba4e0a32d

mirror of https://github.com/ClickHouse/ClickHouse.git synced 2024-12-10 16:32:12 +00:00

This has been driving me crazy for a while 😄 The table functions are listed alphabetically except for this one - so it's a trivial fix.

2024-07-19 16:33:15 -06:00

3.0 KiB

Raw Blame History

slug	sidebar_position	sidebar_label	title
/en/sql-reference/table-functions/azureBlobStorageCluster	15	azureBlobStorageCluster	azureBlobStorageCluster Table Function

Allows processing files from Azure Blob Storage in parallel from many nodes in a specified cluster. On initiator it creates a connection to all nodes in the cluster, discloses asterisks in S3 file path, and dispatches each file dynamically. On the worker node it asks the initiator about the next task to process and processes it. This is repeated until all tasks are finished. This table function is similar to the s3Cluster function.

Syntax

azureBlobStorageCluster(cluster_name, connection_string|storage_account_url, container_name, blobpath, [account_name, account_key, format, compression, structure])

Arguments

cluster_name — Name of a cluster that is used to build a set of addresses and connection parameters to remote and local servers.
connection_string|storage_account_url — connection_string includes account name & key (Create connection string) or you could also provide the storage account url here and account name & account key as separate parameters (see parameters account_name & account_key)
container_name - Container name
blobpath - file path. Supports following wildcards in readonly mode: *, **, ?, {abc,def} and {N..M} where N, M — numbers, 'abc', 'def' — strings.
account_name - if storage_account_url is used, then account name can be specified here
account_key - if storage_account_url is used, then account key can be specified here
format — The format of the file.
compression — Supported values: none, gzip/gz, brotli/br, xz/LZMA, zstd/zst. By default, it will autodetect compression by file extension. (same as setting to auto).
structure — Structure of the table. Format 'column1_name column1_type, column2_name column2_type, ...'.

Returned value

A table with the specified structure for reading or writing data in the specified file.

Examples

Select the count for the file test_cluster_*.csv, using all the nodes in the cluster_simple cluster:

SELECT count(*) from azureBlobStorageCluster(
        'cluster_simple', 'http://azurite1:10000/devstoreaccount1', 'test_container', 'test_cluster_count.csv', 'devstoreaccount1',
        'Eby8vdM02xNOcqFlqUwJPLlmEtlCDXJ1OUzFT50uSRZ6IFsuFq2UVErCz4I6tq/K1SZFPTOtr/KBHBeksoGMGw==', 'CSV',
        'auto', 'key UInt64')

See Also

3.0 KiB Raw Blame History

3.0 KiB

Raw Blame History