ClickHouse

mirror of https://github.com/ClickHouse/ClickHouse.git synced 2024-12-15 02:41:59 +00:00

Author	SHA1	Message	Date
Azat Khuzhin	ff12f5102a	Avoid running LIMIT BY/DISTINCT step on the initiator for optimize_distributed_group_by_sharding_key Before the following queries was running LimitBy/Distinct step on the initator: select distinct sharding_key from dist order by k While this can be omitted.	2021-08-02 21:04:30 +03:00
Azat Khuzhin	2fb95d9ee0	Rework SELECT from Distributed query stages optimization Before this patch it wasn't possible to optimize simple SELECT * FROM dist ORDER BY (w/o GROUP BY and DISTINCT) to more optimal stage (QueryProcessingStage::WithMergeableStateAfterAggregationAndLimit), since that code was under allow_nondeterministic_optimize_skip_unused_shards, rework it and make it possible. Also now distributed_push_down_limit is respected for optimize_distributed_group_by_sharding_key. Next step will be to enable distributed_push_down_limit by default. v2: fix detection of aggregates	2021-08-02 21:04:29 +03:00
Azat Khuzhin	bb6d030fb8	Optimize distributed SELECT w/o GROUP BY	2021-08-02 21:04:29 +03:00
Nikolai Kochetov	61d8f880cd	Rename some files.	2021-07-26 19:48:25 +03:00
Nikolai Kochetov	2dc5c89b66	Update Storage::write	2021-07-23 17:25:35 +03:00
Amos Bird	dbfb699690	Asynchronously drain connections.	2021-07-19 21:53:29 +08:00
alexey-milovidov	bc907bd27c	Merge pull request #26336 from azat/dist-per-table-monitor-settings Add ability to set Distributed directory monitor settings via CREATE TABLE	2021-07-17 01:49:40 +03:00
alexey-milovidov	1701cc429d	Merge pull request #26353 from azat/optimize_distributed_group_by_sharding_key-fix Fix optimize_distributed_group_by_sharding_key for multiple columns	2021-07-17 01:45:10 +03:00
Azat Khuzhin	f3d3ec44a6	Add ability to set Distributed directory monitor settings via CREATE TABLE	2021-07-16 04:10:47 +03:00
Nikolai Kochetov	f36d14f68f	Add separate step to read from remote.	2021-07-15 19:15:16 +03:00
Azat Khuzhin	7b209694d5	Fix optimize_distributed_group_by_sharding_key for multiple columns Before we incorrectly check that columns from GROUP BY was a subset of columns from sharding key, while this is not right, consider the following example: select k1, any(k2), sum(v) from remote('127.{1,2}', view(select 1 k1, 2 k2, 3 v), cityHash64(k1, k2)) group by k1 Here the columns from GROUP BY is a subset of columns from sharding key, but the optimization cannot be applied, since there is no guarantee that particular shard contains distinct values of k1. So instead we should check that GROUP BY contains all columns that is required for calculating sharding key expression, i.e.: select k1, k2, sum(v) from remote('127.{1,2}', view(select 1 k1, 2 k2, 3 v), cityHash64(k1, k2)) group by k1, k2	2021-07-15 09:09:58 +03:00
Azat Khuzhin	533df9507f	Fix log message for optimize_skip_unused_shards_limit	2021-07-07 00:17:39 +03:00
Raúl Marín	bfc122df64	Fix some typos in Storage classes	2021-06-28 19:03:56 +02:00
alexey-milovidov	1b644b9a31	Merge pull request #25663 from azat/dist-startup Improve startup time of Distributed engine.	2021-06-27 18:22:45 +03:00
alexey-milovidov	f6e67d3dc1	Update StorageDistributed.cpp	2021-06-27 18:22:34 +03:00
Alexander Tokmakov	3a25b05765	fix rename Distributed table	2021-06-24 13:00:33 +03:00
Azat Khuzhin	a616ae8861	Improve startup time of Distributed engine. - create directory monitors in parallel (this also includes rmdir in case of directory is empty, since even if the directory is empty it may take some time to remove it, due to waiting for journal or if the directory is large, i.e. it had lots of files before, since remember ext4 does not truncate the directory size on each unlink [1]) - initialize increment in parallel too (since it does readdir()) [1]: https://lore.kernel.org/linux-ext4/930A5754-5CE6-4567-8CF0-62447C97825C@dilger.ca/	2021-06-24 10:27:51 +03:00
Anton Popov	d8b6f15ef4	Merge pull request #23027 from azat/distributed-push-down-limit Add ability to push down LIMIT for distributed queries	2021-06-20 23:08:50 +03:00
Maksim Kita	67e9b85951	Merge ext into common	2021-06-16 23:28:41 +03:00
alexey-milovidov	34d12063f8	Merge pull request #23349 from azat/dist-respect-insert_allow_materialized_columns Respect insert_allow_materialized_columns for INSERT into Distributed()	2021-06-14 07:23:00 +03:00
Nikita Mikhaylov	82b8d45cd7	Merge pull request #23518 from nikitamikhaylov/copier-stuck Bugfixes and improvements of `clickhouse-copier`	2021-06-09 11:36:42 +03:00
Azat Khuzhin	18e8f0eb5e	Add ability to push down LIMIT for distributed queries This way the remote nodes will not need to send all the rows, so this will decrease network io and also this will make queries w/ optimize_aggregation_in_order=1/LIMIT X and w/o ORDER BY faster since it initiator will not need to read all the rows, only first X (but note that for this you need to your data to be sharded correctly or you may get inaccurate results). Note, that having lots of processing stages will increase the complexity of interpreter (it is already not that clean and simple right now). Although using separate QueryProcessingStage looks pretty natural. Another option is to make WithMergeableStateAfterAggregation always, but in this case you will not be able to disable only this optimization, i.e. if there will be some issue with it. v2: fix OFFSET v3: convert 01814_distributed_push_down_limit test to .sh and add retries v4: add test with OFFSET v5: add new query stage into the bash completion v6/tests: use LIMIT O,L syntax over LIMIT L OFFSET O since it is broken in ANTLR parser https://clickhouse-test-reports.s3.yandex.net/23027/a18a06399b7aeacba7c50b5d1e981ada5df19745/functional_stateless_tests_(antlr_debug).html#fail1 v7/tests: set use_hedged_requests to 0, to avoid excessive log entries on retries https://clickhouse-test-reports.s3.yandex.net/23027/a18a06399b7aeacba7c50b5d1e981ada5df19745/functional_stateless_tests_flaky_check_(address).html#fail1	2021-06-09 02:29:50 +03:00
Amos Bird	78fca8f8fa	Fix possible race condition when getting cluster	2021-06-04 21:09:59 +08:00
Nikita Mikhaylov	312bb96eeb	Merge branch 'master' of github.com:ClickHouse/ClickHouse into copier-stuck	2021-06-02 01:04:47 +03:00
Nikita Mikhaylov	6d19dea761	better	2021-05-31 17:38:20 +03:00
Nikita Mikhaylov	90ab394769	better	2021-05-31 17:37:10 +03:00
kssenii	3dee003f9b	Merge branch 'master' of github.com:ClickHouse/ClickHouse into poco-file-to-std-fs	2021-05-20 19:20:09 +03:00
Azat Khuzhin	4d737a5481	Respect insert_allow_materialized_columns for INSERT into Distributed()	2021-05-20 07:40:46 +03:00
Alexander Kuzmenkov	e9b69bbd70	Merge pull request #23906 from azat/fix-distributed_group_by_no_merge distributed_group_by_no_merge fixes	2021-05-19 16:16:08 +03:00
Alexander Kuzmenkov	09cb467812	Update StorageDistributed.cpp	2021-05-19 16:14:33 +03:00
kssenii	9b8df78fdd	Merge branch 'master' of github.com:ClickHouse/ClickHouse into poco-file-to-std-fs	2021-05-17 17:42:05 +03:00
feng lv	c6f8ab9826	fix	2021-05-13 02:05:53 +00:00
kssenii	0527f0ea33	Merge branch 'master' of github.com:ClickHouse/ClickHouse into poco-file-to-std-fs	2021-05-12 16:54:18 +03:00
Amos Bird	cd6414639e	add metadata_snapshot to getQueryProcessingStage	2021-05-11 18:12:26 +08:00
Azat Khuzhin	eefd67fce5	Disable optimize_distributed_group_by_sharding_key with window functions	2021-05-06 00:44:22 +03:00
feng lv	39f68bf5ff	fix conflict	2021-05-02 16:33:45 +00:00
kssenii	ee06936596	Merge branch 'master' of github.com:ClickHouse/ClickHouse into poco-file-to-std-fs	2021-05-01 17:24:31 +03:00
feng lv	aed2f337e9	Fix CLEAR COLUMN does not work after #21303	2021-04-30 05:02:32 +00:00
kssenii	deb4903af8	Merge branch 'master' of github.com:ClickHouse/ClickHouse into poco-file-to-std-fs	2021-04-28 20:57:13 +03:00
kssenii	eeb71672a0	Change in Storages/*	2021-04-27 16:49:37 +03:00
feng lv	4ffe199d39	Implement table comments	2021-04-23 12:18:23 +00:00
Amos Bird	096d76627e	Skip unavaiable shards when writing to distributed tables	2021-04-21 10:30:40 +08:00
Maksim Kita	e361f5943f	Merge pull request #22999 from azat/no-optimize_skip_unused_shards-single-node Do not perform optimize_skip_unused_shards for cluster with one node	2021-04-15 14:36:56 +03:00
Nikita Mikhaylov	7a68820342	style	2021-04-13 22:39:42 +03:00
Nikita Mikhaylov	081ea84a41	save	2021-04-13 22:39:41 +03:00
tavplubix	1525e38a3c	Merge pull request #22990 from ClickHouse/tavplubix-patch-1 Fix excessive warning in StorageDistributed with cross-replication	2021-04-13 18:58:12 +03:00
Azat Khuzhin	a497d4d462	Do not perform optimize_skip_unused_shards for cluster with one node	2021-04-12 22:18:31 +03:00
tavplubix	a995962e6a	Update StorageDistributed.cpp	2021-04-12 14:58:24 +03:00
Azat Khuzhin	79bd8d4d3f	Respect optimize_skip_unused_shards_rewrite_in with optimize_skip_unused_shards_limit	2021-04-12 10:37:28 +03:00
Azat Khuzhin	e439914d38	Fix optimized cluster logic for optimize_skip_unused_shards	2021-04-12 10:37:28 +03:00

1 2 3 4 5

223 Commits