ClickHouse

mirror of https://github.com/ClickHouse/ClickHouse.git synced 2024-12-16 03:12:43 +00:00

Author	SHA1	Message	Date
Azat Khuzhin	c4b6342853	Improvements for `parallel_distributed_insert_select` (and related) (#34728 ) * Add a warning if parallel_distributed_insert_select was ignored Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Respect max_distributed_depth for parallel_distributed_insert_select Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Print warning for non applied parallel_distributed_insert_select only for initial query Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Remove Cluster::getHashOfAddresses() Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Forbid parallel_distributed_insert_select for remote()/cluster() with different addresses Before it uses empty cluster name (getClusterName()) which is not correct, compare all addresses instead. Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Fix max_distributed_depth check max_distributed_depth=1 must mean not more then one distributed query, not two, since max_distributed_depth=0 means no limit, and distribute_depth is 0 for the first query. Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Fix INSERT INTO remote()/cluster() with parallel_distributed_insert_select Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Add a test for parallel_distributed_insert_select with cluster()/remote() Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Return <remote> instead of empty cluster name in Distributed engine Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com> * Make user with sharding_key and w/o in remote()/cluster() identical Before with sharding_key the user was "default", while w/o it it was empty. Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com>	2022-03-08 15:24:39 +01:00
feng lv	6325d4d9b0	continue of #34317 fix fix	2022-02-06 08:59:17 +00:00
Azat Khuzhin	1637c41d42	Remove leftovers of old _shard_num via identifier implementation	2022-01-10 21:21:24 +03:00
Kruglov Pavel	2295a07066	Merge pull request #33534 from azat/fwd-decl RFC: Split headers, move SystemLog into module, more forward declarations	2022-01-18 17:22:49 +03:00
Azat Khuzhin	c341b3b237	Add current database to table names in JOIN section for distributed queries This should fix JOIN w/o explicit database. v2: rewrite only JOIN section, since there is old behavior that relies on default_database for IN section, see [1]: - 01487_distributed_in_not_default_db - 01152_cross_replication [1]: https://s3.amazonaws.com/clickhouse-test-reports/33611/d0ea3c76fa51131171b1825939680867eb1c04da/fast_test__actions_.html Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com>	2022-01-14 11:23:38 +03:00
Azat Khuzhin	0a9b1ee803	Remove RestoreQualifiedNamesMatcher::Data::rename (always true) Signed-off-by: Azat Khuzhin <a.khuzhin@semrush.com>	2022-01-14 11:18:52 +03:00
Azat Khuzhin	aee034a597	Use explicit template instantiation for SystemLog - Move some code into module part to avoid dependency from IStorage in SystemLog - Remove extra headers from SystemLog.h - Rewrite some code that was relying on headers that was included by SystemLog.h v2: rebase v3: squash move into module part with explicit template instantiation (to make each commit self compilable after rebase)	2022-01-10 22:01:41 +03:00
avogar	8112a71233	Implement schema inference for most input formats	2021-12-29 12:18:56 +03:00
Alexey Milovidov	5c90ed2ed9	Unambiguous formatting of distributed queries	2021-12-10 00:55:14 +03:00
Nikita Mikhaylov	dbf5091016	Parallel reading from replicas (#29279 )	2021-12-09 13:39:28 +03:00
Raúl Marín	7781fc12ed	Reduce dependencies on ASTSelectWithUnionQuery.h 521 -> 77 files requiring changes	2021-11-26 19:27:16 +01:00
Raúl Marín	b2cfa70541	Reduce dependencies on ASTFunction.h 481 -> 230	2021-11-26 18:21:54 +01:00
feng lv	6f12348282	enable modify table comment of some table	2021-10-29 12:31:18 +00:00
Alexander Tokmakov	2e7e195e77	change alter_lock to std::timed_mutex	2021-10-26 13:37:00 +03:00
Nikolai Kochetov	fd14faeae2	Remove DataStreams folder.	2021-10-15 23:18:20 +03:00
Nikolai Kochetov	2957971ee3	Remove some last streams.	2021-10-13 21:22:02 +03:00
Vitaly Baranov	1636ee24bb	Fix using materialized column as sharding key.	2021-10-04 10:56:42 +03:00
Nikolai Kochetov	341553febd	Fix build.	2021-09-16 20:40:42 +03:00
Nikolai Kochetov	b997214620	Rename QueryPipeline to QueryPipelineBuilder.	2021-09-14 20:48:18 +03:00
Nikolai Kochetov	0e267c50b4	Merge branch 'master' into rewrite-pushing-to-views	2021-09-14 16:13:54 +03:00
alexey-milovidov	ea13a8b562	Merge pull request #28659 from myrrc/improvement/tostring_to_magic_enum Improving CH type system with concepts	2021-09-12 15:26:29 +03:00
Nikolai Kochetov	f569a3e3f7	Merge branch 'master' into rewrite-pushing-to-views	2021-09-09 20:30:23 +03:00
Nikolai Kochetov	999a4fe831	Fix other tests.	2021-09-08 21:29:38 +03:00
ZhiYong Wang	978dd19fa2	Fix coredump in creating distributed table	2021-09-07 19:05:26 +08:00
Mike Kot	8e9aacadd1	Initial: replacing hardcoded toString for enums with magic_enum	2021-09-06 16:24:03 +02:00
Alexander Tokmakov	42378b5913	fix	2021-08-20 17:05:53 +03:00
Alexander Tokmakov	8c6dd18917	check cluster name before creating Distributed	2021-08-20 14:55:04 +03:00
Azat Khuzhin	702d9955c0	Fix distributed queries with zero shards and aggregation	2021-08-08 19:22:49 +03:00
Azat Khuzhin	3be3c503aa	Fix some comments	2021-08-08 09:58:07 +03:00
alexey-milovidov	c5207fc237	Merge pull request #26466 from azat/optimize-dist-select Rework SELECT from Distributed optimizations	2021-08-08 03:59:32 +03:00
mergify[bot]	dc57254982	Merge branch 'master' into improve_create_or_replace	2021-08-03 11:39:07 +00:00
Azat Khuzhin	97851bde08	Fix Distributed over Distributed for WithMergeableStateAfterAggregation* stages In case if one Distributed has multiple shards, and underlying Distributed has only one, there can be the case when the query will be tried to process from Complete to WithMergeableStateAfterAggregation, which is obviously wrong.	2021-08-03 10:10:08 +03:00
Azat Khuzhin	ff12f5102a	Avoid running LIMIT BY/DISTINCT step on the initiator for optimize_distributed_group_by_sharding_key Before the following queries was running LimitBy/Distinct step on the initator: select distinct sharding_key from dist order by k While this can be omitted.	2021-08-02 21:04:30 +03:00
Azat Khuzhin	2fb95d9ee0	Rework SELECT from Distributed query stages optimization Before this patch it wasn't possible to optimize simple SELECT * FROM dist ORDER BY (w/o GROUP BY and DISTINCT) to more optimal stage (QueryProcessingStage::WithMergeableStateAfterAggregationAndLimit), since that code was under allow_nondeterministic_optimize_skip_unused_shards, rework it and make it possible. Also now distributed_push_down_limit is respected for optimize_distributed_group_by_sharding_key. Next step will be to enable distributed_push_down_limit by default. v2: fix detection of aggregates	2021-08-02 21:04:29 +03:00
Azat Khuzhin	bb6d030fb8	Optimize distributed SELECT w/o GROUP BY	2021-08-02 21:04:29 +03:00
Nikolai Kochetov	61d8f880cd	Rename some files.	2021-07-26 19:48:25 +03:00
mergify[bot]	044be267d6	Merge branch 'master' into improve_create_or_replace	2021-07-26 08:38:48 +00:00
Nikolai Kochetov	2dc5c89b66	Update Storage::write	2021-07-23 17:25:35 +03:00
Amos Bird	dbfb699690	Asynchronously drain connections.	2021-07-19 21:53:29 +08:00
alexey-milovidov	bc907bd27c	Merge pull request #26336 from azat/dist-per-table-monitor-settings Add ability to set Distributed directory monitor settings via CREATE TABLE	2021-07-17 01:49:40 +03:00
alexey-milovidov	1701cc429d	Merge pull request #26353 from azat/optimize_distributed_group_by_sharding_key-fix Fix optimize_distributed_group_by_sharding_key for multiple columns	2021-07-17 01:45:10 +03:00
Azat Khuzhin	f3d3ec44a6	Add ability to set Distributed directory monitor settings via CREATE TABLE	2021-07-16 04:10:47 +03:00
Nikolai Kochetov	f36d14f68f	Add separate step to read from remote.	2021-07-15 19:15:16 +03:00
Azat Khuzhin	7b209694d5	Fix optimize_distributed_group_by_sharding_key for multiple columns Before we incorrectly check that columns from GROUP BY was a subset of columns from sharding key, while this is not right, consider the following example: select k1, any(k2), sum(v) from remote('127.{1,2}', view(select 1 k1, 2 k2, 3 v), cityHash64(k1, k2)) group by k1 Here the columns from GROUP BY is a subset of columns from sharding key, but the optimization cannot be applied, since there is no guarantee that particular shard contains distinct values of k1. So instead we should check that GROUP BY contains all columns that is required for calculating sharding key expression, i.e.: select k1, k2, sum(v) from remote('127.{1,2}', view(select 1 k1, 2 k2, 3 v), cityHash64(k1, k2)) group by k1, k2	2021-07-15 09:09:58 +03:00
Azat Khuzhin	533df9507f	Fix log message for optimize_skip_unused_shards_limit	2021-07-07 00:17:39 +03:00
Alexander Tokmakov	1b2416007e	fix	2021-07-01 19:43:59 +03:00
Alexander Tokmakov	d9a77e3a1a	improve CREATE OR REPLACE query	2021-07-01 16:21:38 +03:00
Raúl Marín	bfc122df64	Fix some typos in Storage classes	2021-06-28 19:03:56 +02:00
alexey-milovidov	1b644b9a31	Merge pull request #25663 from azat/dist-startup Improve startup time of Distributed engine.	2021-06-27 18:22:45 +03:00
alexey-milovidov	f6e67d3dc1	Update StorageDistributed.cpp	2021-06-27 18:22:34 +03:00

1 2 3 4 5 ...

258 Commits