ClickHouse/docs/en/engines/table-engines/integrations/kafka.md

---
slug: /en/engines/table-engines/integrations/kafka
sidebar_position: 110
sidebar_label: Kafka
---

# Kafka

This engine works with [Apache Kafka](http://kafka.apache.org/).

Kafka lets you:

- Publish or subscribe to data flows.
- Organize fault-tolerant storage.
- Process streams as they become available.

## Creating a Table {#creating-a-table}

``` sql
CREATE TABLE [IF NOT EXISTS] [db.]table_name [ON CLUSTER cluster]
(
    name1 [type1] [ALIAS expr1],
    name2 [type2] [ALIAS expr2],
    ...
) ENGINE = Kafka()
SETTINGS
    kafka_broker_list = 'host:port',
    kafka_topic_list = 'topic1,topic2,...',
    kafka_group_name = 'group_name',
    kafka_format = 'data_format'[,]
    [kafka_schema = '',]
    [kafka_num_consumers = N,]
    [kafka_max_block_size = 0,]
    [kafka_skip_broken_messages = N,]
    [kafka_commit_every_batch = 0,]
    [kafka_client_id = '',]
    [kafka_poll_timeout_ms = 0,]
    [kafka_poll_max_batch_size = 0,]
    [kafka_flush_interval_ms = 0,]
    [kafka_thread_per_consumer = 0,]
    [kafka_handle_error_mode = 'default',]
    [kafka_commit_on_select = false,]
    [kafka_max_rows_per_message = 1];
```

Required parameters:

- `kafka_broker_list` — A comma-separated list of brokers (for example, `localhost:9092`).
- `kafka_topic_list` — A list of Kafka topics.
- `kafka_group_name` — A group of Kafka consumers. Reading margins are tracked for each group separately. If you do not want messages to be duplicated in the cluster, use the same group name everywhere.
- `kafka_format` — Message format. Uses the same notation as the SQL `FORMAT` function, such as `JSONEachRow`. For more information, see the [Formats](../../../interfaces/formats.md) section.

Optional parameters:

- `kafka_schema` — Parameter that must be used if the format requires a schema definition. For example, [Cap’n Proto](https://capnproto.org/) requires the path to the schema file and the name of the root `schema.capnp:Message` object.
- `kafka_num_consumers` — The number of consumers per table. Specify more consumers if the throughput of one consumer is insufficient. The total number of consumers should not exceed the number of partitions in the topic, since only one consumer can be assigned per partition, and must not be greater than the number of physical cores on the server where ClickHouse is deployed. Default: `1`.
- `kafka_max_block_size` — The maximum batch size (in messages) for poll. Default: [max_insert_block_size](../../../operations/settings/settings.md#max_insert_block_size).
- `kafka_skip_broken_messages` — Kafka message parser tolerance to schema-incompatible messages per block. If `kafka_skip_broken_messages = N` then the engine skips *N* Kafka messages that cannot be parsed (a message equals a row of data). Default: `0`.
- `kafka_commit_every_batch` — Commit every consumed and handled batch instead of a single commit after writing a whole block. Default: `0`.
- `kafka_client_id` — Client identifier. Empty by default.
- `kafka_poll_timeout_ms` — Timeout for single poll from Kafka. Default: [stream_poll_timeout_ms](../../../operations/settings/settings.md#stream_poll_timeout_ms).
- `kafka_poll_max_batch_size` — Maximum amount of messages to be polled in a single Kafka poll. Default: [max_block_size](../../../operations/settings/settings.md#setting-max_block_size).
- `kafka_flush_interval_ms` — Timeout for flushing data from Kafka. Default: [stream_flush_interval_ms](../../../operations/settings/settings.md#stream-flush-interval-ms).
- `kafka_thread_per_consumer` — Provide independent thread for each consumer. When enabled, every consumer flush the data independently, in parallel (otherwise — rows from several consumers squashed to form one block). Default: `0`.
- `kafka_handle_error_mode` — How to handle errors for Kafka engine. Possible values: default (the exception will be thrown if we fail to parse a message), stream (the exception message and raw message will be saved in virtual columns `_error` and `_raw_message`).
- `kafka_commit_on_select` —  Commit messages when select query is made. Default: `false`.
- `kafka_max_rows_per_message` — The maximum number of rows written in one kafka message for row-based formats. Default : `1`.

Examples:

``` sql
  CREATE TABLE queue (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');

  SELECT * FROM queue LIMIT 5;

  CREATE TABLE queue2 (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka SETTINGS kafka_broker_list = 'localhost:9092',
                            kafka_topic_list = 'topic',
                            kafka_group_name = 'group1',
                            kafka_format = 'JSONEachRow',
                            kafka_num_consumers = 4;

  CREATE TABLE queue3 (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1')
              SETTINGS kafka_format = 'JSONEachRow',
                       kafka_num_consumers = 4;
```

<details markdown="1">

<summary>Deprecated Method for Creating a Table</summary>

:::note
Do not use this method in new projects. If possible, switch old projects to the method described above.
:::

``` sql
Kafka(kafka_broker_list, kafka_topic_list, kafka_group_name, kafka_format
      [, kafka_row_delimiter, kafka_schema, kafka_num_consumers, kafka_max_block_size,  kafka_skip_broken_messages, kafka_commit_every_batch, kafka_client_id, kafka_poll_timeout_ms, kafka_poll_max_batch_size, kafka_flush_interval_ms, kafka_thread_per_consumer, kafka_handle_error_mode, kafka_commit_on_select, kafka_max_rows_per_message]);
```

</details>

:::info
The Kafka table engine doesn't support columns with [default value](../../../sql-reference/statements/create/table.md#default_value). If you need columns with default value, you can add them at materialized view level (see below).
:::

## Description {#description}

The delivered messages are tracked automatically, so each message in a group is only counted once. If you want to get the data twice, then create a copy of the table with another group name.

Groups are flexible and synced on the cluster. For instance, if you have 10 topics and 5 copies of a table in a cluster, then each copy gets 2 topics. If the number of copies changes, the topics are redistributed across the copies automatically. Read more about this at http://kafka.apache.org/intro.

`SELECT` is not particularly useful for reading messages (except for debugging), because each message can be read only once. It is more practical to create real-time threads using materialized views. To do this:

1.  Use the engine to create a Kafka consumer and consider it a data stream.
2.  Create a table with the desired structure.
3.  Create a materialized view that converts data from the engine and puts it into a previously created table.

When the `MATERIALIZED VIEW` joins the engine, it starts collecting data in the background. This allows you to continually receive messages from Kafka and convert them to the required format using `SELECT`.
One kafka table can have as many materialized views as you like, they do not read data from the kafka table directly, but receive new records (in blocks), this way you can write to several tables with different detail level (with grouping - aggregation and without).

Example:

``` sql
  CREATE TABLE queue (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');

  CREATE TABLE daily (
    day Date,
    level String,
    total UInt64
  ) ENGINE = SummingMergeTree(day, (day, level), 8192);

  CREATE MATERIALIZED VIEW consumer TO daily
    AS SELECT toDate(toDateTime(timestamp)) AS day, level, count() as total
    FROM queue GROUP BY day, level;

  SELECT level, sum(total) FROM daily GROUP BY level;
```
To improve performance, received messages are grouped into blocks the size of [max_insert_block_size](../../../operations/settings/settings.md#max_insert_block_size). If the block wasn’t formed within [stream_flush_interval_ms](../../../operations/settings/settings.md/#stream-flush-interval-ms) milliseconds, the data will be flushed to the table regardless of the completeness of the block.

To stop receiving topic data or to change the conversion logic, detach the materialized view:

``` sql
  DETACH TABLE consumer;
  ATTACH TABLE consumer;
```

If you want to change the target table by using `ALTER`, we recommend disabling the material view to avoid discrepancies between the target table and the data from the view.

## Configuration {#configuration}

Similar to GraphiteMergeTree, the Kafka engine supports extended configuration using the ClickHouse config file. There are two configuration keys that you can use: global (below `<kafka>`) and topic-level (below `<kafka><kafka_topic>`). The global configuration is applied first, and then the topic-level configuration is applied (if it exists).

``` xml
  <kafka>
    <!-- Global configuration options for all tables of Kafka engine type -->
    <debug>cgrp</debug>
    <statistics_interval_ms>3000</statistics_interval_ms>

    <kafka_topic>
        <name>logs</name>
        <statistics_interval_ms>4000</statistics_interval_ms>
    </kafka_topic>

    <!-- Settings for consumer -->
    <consumer>
        <auto_offset_reset>smallest</auto_offset_reset>
        <kafka_topic>
            <name>logs</name>
            <fetch_min_bytes>100000</fetch_min_bytes>
        </kafka_topic>

        <kafka_topic>
            <name>stats</name>
            <fetch_min_bytes>50000</fetch_min_bytes>
        </kafka_topic>
    </consumer>

    <!-- Settings for producer -->
    <producer>
        <kafka_topic>
            <name>logs</name>
            <retry_backoff_ms>250</retry_backoff_ms>
        </kafka_topic>

        <kafka_topic>
            <name>stats</name>
            <retry_backoff_ms>400</retry_backoff_ms>
        </kafka_topic>
    </producer>
  </kafka>
```


For a list of possible configuration options, see the [librdkafka configuration reference](https://github.com/edenhill/librdkafka/blob/master/CONFIGURATION.md). Use the underscore (`_`) instead of a dot in the ClickHouse configuration. For example, `check.crcs=true` will be `<check_crcs>true</check_crcs>`.

### Kerberos support {#kafka-kerberos-support}

To deal with Kerberos-aware Kafka, add `security_protocol` child element with `sasl_plaintext` value. It is enough if Kerberos ticket-granting ticket is obtained and cached by OS facilities.
ClickHouse is able to maintain Kerberos credentials using a keytab file. Consider `sasl_kerberos_service_name`, `sasl_kerberos_keytab` and `sasl_kerberos_principal` child elements.

Example:

``` xml
  <!-- Kerberos-aware Kafka -->
  <kafka>
    <security_protocol>SASL_PLAINTEXT</security_protocol>
	<sasl_kerberos_keytab>/home/kafkauser/kafkauser.keytab</sasl_kerberos_keytab>
	<sasl_kerberos_principal>kafkauser/kafkahost@EXAMPLE.COM</sasl_kerberos_principal>
  </kafka>
```

## Virtual Columns {#virtual-columns}

- `_topic` — Kafka topic. Data type: `LowCardinality(String)`.
- `_key` — Key of the message. Data type: `String`.
- `_offset` — Offset of the message. Data type: `UInt64`.
- `_timestamp` — Timestamp of the message Data type: `Nullable(DateTime)`.
- `_timestamp_ms` — Timestamp in milliseconds of the message. Data type: `Nullable(DateTime64(3))`.
- `_partition` — Partition of Kafka topic. Data type: `UInt64`.
- `_headers.name` — Array of message's headers keys. Data type: `Array(String)`.
- `_headers.value` — Array of message's headers values. Data type: `Array(String)`.

Additional virtual columns when `kafka_handle_error_mode='stream'`:

- `_raw_message` - Raw message that couldn't be parsed successfully. Data type: `String`.
- `_error` - Exception message happened during failed parsing. Data type: `String`.

Note: `_raw_message` and `_error` virtual columns are filled only in case of exception during parsing, they are always empty when message was parsed successfully.

## Data formats support {#data-formats-support}

Kafka engine supports all [formats](../../../interfaces/formats.md) supported in ClickHouse.
The number of rows in one Kafka message depends on whether the format is row-based or block-based:

- For row-based formats the number of rows in one Kafka message can be controlled by setting `kafka_max_rows_per_message`.
- For block-based formats we cannot divide block into smaller parts, but the number of rows in one block can be controlled by general setting [max_block_size](../../../operations/settings/settings.md#setting-max_block_size).

## Experimental engine to store committed offsets in ClickHouse Keeper

If `allow_experimental_kafka_offsets_storage_in_keeper` is enabled, then two more settings can be specified to the Kafka table engine:
 - `kafka_keeper_path` specifies the path to the table in ClickHouse Keeper
 - `kafka_replica_name` specifies the replica name in ClickHouse Keeper

Either both of the settings must be specified or neither of them. When both of them are specified, then a new, experimental Kafka engine will be used. The new engine doesn't depend on storing the committed offsets in Kafka, but stores them in ClickHouse Keeper. It still tries to commit the offsets to Kafka, but it only depends on those offsets when the table is created. In any other circumstances (table is restarted, or recovered after some error) the offsets stored in ClickHouse Keeper will be used as an offset to continue consuming messages from. Apart from the committed offset, it also stores how many messages were consumed in the last batch, so if the insert fails, the same amount of messages will be consumed, thus enabling deduplication if necessary.

Example:

``` sql
CREATE TABLE experimental_kafka (key UInt64, value UInt64)
ENGINE = Kafka('localhost:19092', 'my-topic', 'my-consumer', 'JSONEachRow')
SETTINGS
  kafka_keeper_path = '/clickhouse/{database}/experimental_kafka',
  kafka_replica_name = 'r1'
SETTINGS allow_experimental_kafka_offsets_storage_in_keeper=1;
```

Or to utilize the `uuid` and `replica` macros similarly to ReplicatedMergeTree:

``` sql
CREATE TABLE experimental_kafka (key UInt64, value UInt64)
ENGINE = Kafka('localhost:19092', 'my-topic', 'my-consumer', 'JSONEachRow')
SETTINGS
  kafka_keeper_path = '/clickhouse/{database}/{uuid}',
  kafka_replica_name = '{replica}'
SETTINGS allow_experimental_kafka_offsets_storage_in_keeper=1;
```

### Known limitations

As the new engine is experimental, it is not production ready yet. There are few known limitations of the implementation:
 - The biggest limitation is the engine doesn't support direct reading. Reading from the engine using materialized views and writing to the engine work, but direct reading doesn't. As a result, all direct `SELECT` queries will fail.
 - Rapidly dropping and recreating the table or specifying the same ClickHouse Keeper path to different engines might cause issues. As best practice you can use the `{uuid}` in `kafka_keeper_path` to avoid clashing paths.
 - To make repeatable reads, messages cannot be consumed from multiple partitions on a single thread. On the other hand, the Kafka consumers have to be polled regularly to keep them alive. As a result of these two objectives, we decided to only allow creating multiple consumers if `kafka_thread_per_consumer` is enabled, otherwise it is too complicated to avoid issues regarding polling consumers regularly.
 - Consumers created by the new storage engine do not show up in [`system.kafka_consumers`](../../../operations/system-tables/kafka_consumers.md) table.

**See Also**

- [Virtual columns](../../../engines/table-engines/index.md#table_engines-virtual_columns)
- [background_message_broker_schedule_pool_size](../../../operations/server-configuration-parameters/settings.md#background_message_broker_schedule_pool_size)
- [system.kafka_consumers](../../../operations/system-tables/kafka_consumers.md)
-												Get rid of toc_en.yml (#10023)


											
										
										
											2020-04-03 13:23:32 +00:00
+								---
-												add slugs

											
										
										
											2022-08-28 14:53:34 +00:00
+								slug: /en/engines/table-engines/integrations/kafka
-												Alphabetize table functions and engines

											
										
										
											2023-06-23 13:16:22 +00:00
+								sidebar_position: 110
-												Removed /ja folder, cleaned up /ru markdown

											
										
										
											2022-04-09 13:29:05 +00:00
+								sidebar_label: Kafka
-												Get rid of toc_en.yml (#10023)


											
										
										
											2020-04-03 13:23:32 +00:00
+								---
-												Remove H1 anchor tags from docs

											
										
										
											2022-06-02 10:55:18 +00:00
+								# Kafka
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								This engine works with [Apache Kafka](http://kafka.apache.org/).
 								Kafka lets you:
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- Publish or subscribe to data flows.
 								- Organize fault-tolerant storage.
 								- Process streams as they become available.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Docs: remove anchor prefix

											
										
										
											2023-09-18 18:29:13 +00:00
+								## Creating a Table {#creating-a-table}
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								CREATE TABLE [IF NOT EXISTS] [db.]table_name [ON CLUSTER cluster]
 								(
-												Allow using Alias column type for KafkaEngine

```
create table kafka
(
 a UInt32,
 a_str String Alias toString(a)
) engine = Kafka;

create table data
(
  a UInt32;
  a_str String
) engine = MergeTree
order by tuple();

create materialized view data_mv to data
(
  a UInt32,
  a_str String
) as
select a, a_str from kafka;
```
Alias type works as expected in comparison with MATERIALIZED/EPHEMERAL
or column with default expression.

Ref: https://github.com/ClickHouse/ClickHouse/pull/47138

Co-authored-by: Azat Khuzhin <a3at.mail@gmail.com>

											
										
										
											2023-05-12 10:47:14 +00:00
+								    name1 [type1] [ALIAS expr1],
 								    name2 [type2] [ALIAS expr2],
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								    ...
 								) ENGINE = Kafka()
 								SETTINGS
 								    kafka_broker_list = 'host:port',
 								    kafka_topic_list = 'topic1,topic2,...',
 								    kafka_group_name = 'group_name',
 								    kafka_format = 'data_format'[,]
 								    [kafka_schema = '',]
 								    [kafka_num_consumers = N,]
-												Add missing kafka settings into docs

											
										
										
											2020-04-27 05:02:45 +00:00
+								    [kafka_max_block_size = 0,]
 								    [kafka_skip_broken_messages = N,]
-												Fix code style, and update docs for Kafka engine

											
										
										
											2020-09-01 08:37:12 +00:00
+								    [kafka_commit_every_batch = 0,]
-												Improve and refactor Kafka/StorageMQ/NATS and data formats

											
										
										
											2022-10-28 16:41:10 +00:00
+								    [kafka_client_id = '',]
 								    [kafka_poll_timeout_ms = 0,]
 								    [kafka_poll_max_batch_size = 0,]
 								    [kafka_flush_interval_ms = 0,]
 								    [kafka_thread_per_consumer = 0,]
 								    [kafka_handle_error_mode = 'default',]
 								    [kafka_commit_on_select = false,]
 								    [kafka_max_rows_per_message = 1];
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								Required parameters:
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- `kafka_broker_list` — A comma-separated list of brokers (for example, `localhost:9092`).
 								- `kafka_topic_list` — A list of Kafka topics.
 								- `kafka_group_name` — A group of Kafka consumers. Reading margins are tracked for each group separately. If you do not want messages to be duplicated in the cluster, use the same group name everywhere.
 								- `kafka_format` — Message format. Uses the same notation as the SQL `FORMAT` function, such as `JSONEachRow`. For more information, see the [Formats](../../../interfaces/formats.md) section.
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								Optional parameters:
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- `kafka_schema` — Parameter that must be used if the format requires a schema definition. For example, [Cap’n Proto](https://capnproto.org/) requires the path to the schema file and the name of the root `schema.capnp:Message` object.
 								- `kafka_num_consumers` — The number of consumers per table. Specify more consumers if the throughput of one consumer is insufficient. The total number of consumers should not exceed the number of partitions in the topic, since only one consumer can be assigned per partition, and must not be greater than the number of physical cores on the server where ClickHouse is deployed. Default: `1`.
-												Fix anchors to settings.md

											
										
										
											2023-12-20 18:26:36 +00:00
+								- `kafka_max_block_size` — The maximum batch size (in messages) for poll. Default: [max_insert_block_size](../../../operations/settings/settings.md#max_insert_block_size).
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- `kafka_skip_broken_messages` — Kafka message parser tolerance to schema-incompatible messages per block. If `kafka_skip_broken_messages = N` then the engine skips *N* Kafka messages that cannot be parsed (a message equals a row of data). Default: `0`.
 								- `kafka_commit_every_batch` — Commit every consumed and handled batch instead of a single commit after writing a whole block. Default: `0`.
 								- `kafka_client_id` — Client identifier. Empty by default.
 								- `kafka_poll_timeout_ms` — Timeout for single poll from Kafka. Default: [stream_poll_timeout_ms](../../../operations/settings/settings.md#stream_poll_timeout_ms).
 								- `kafka_poll_max_batch_size` — Maximum amount of messages to be polled in a single Kafka poll. Default: [max_block_size](../../../operations/settings/settings.md#setting-max_block_size).
 								- `kafka_flush_interval_ms` — Timeout for flushing data from Kafka. Default: [stream_flush_interval_ms](../../../operations/settings/settings.md#stream-flush-interval-ms).
 								- `kafka_thread_per_consumer` — Provide independent thread for each consumer. When enabled, every consumer flush the data independently, in parallel (otherwise — rows from several consumers squashed to form one block). Default: `0`.
-												Add documentation

											
										
										
											2023-10-11 17:35:18 +00:00
+								- `kafka_handle_error_mode` — How to handle errors for Kafka engine. Possible values: default (the exception will be thrown if we fail to parse a message), stream (the exception message and raw message will be saved in virtual columns `_error` and `_raw_message`).
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- `kafka_commit_on_select` —  Commit messages when select query is made. Default: `false`.
 								- `kafka_max_rows_per_message` — The maximum number of rows written in one kafka message for row-based formats. Default : `1`.
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								Examples:
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE TABLE queue (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');
 								  SELECT * FROM queue LIMIT 5;
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								  CREATE TABLE queue2 (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka SETTINGS kafka_broker_list = 'localhost:9092',
 								                            kafka_topic_list = 'topic',
 								                            kafka_group_name = 'group1',
 								                            kafka_format = 'JSONEachRow',
 								                            kafka_num_consumers = 4;
-												table name typo fix (#11100)

According to the context, the table name of queue2 should be queue3.
											
										
										
											2020-05-21 12:14:39 +00:00
+								  CREATE TABLE queue3 (
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1')
 								              SETTINGS kafka_format = 'JSONEachRow',
 								                       kafka_num_consumers = 4;
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								<details markdown="1">
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								<summary>Deprecated Method for Creating a Table</summary>
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												standardize admonitions

											
										
										
											2023-03-27 18:54:05 +00:00
+								:::note
-												Removed /ja folder, cleaned up /ru markdown

											
										
										
											2022-04-09 13:29:05 +00:00
+								Do not use this method in new projects. If possible, switch old projects to the method described above.
 								:::
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								Kafka(kafka_broker_list, kafka_topic_list, kafka_group_name, kafka_format
-												Improve and refactor Kafka/StorageMQ/NATS and data formats

											
										
										
											2022-10-28 16:41:10 +00:00
+								      [, kafka_row_delimiter, kafka_schema, kafka_num_consumers, kafka_max_block_size,  kafka_skip_broken_messages, kafka_commit_every_batch, kafka_client_id, kafka_poll_timeout_ms, kafka_poll_max_batch_size, kafka_flush_interval_ms, kafka_thread_per_consumer, kafka_handle_error_mode, kafka_commit_on_select, kafka_max_rows_per_message]);
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								```
 								</details>
-												Docs: Small cleanups after Kafka fix #47138

											
										
										
											2023-03-07 19:50:42 +00:00
+								:::info
 								The Kafka table engine doesn't support columns with [default value](../../../sql-reference/statements/create/table.md#default_value). If you need columns with default value, you can add them at materialized view level (see below).
 								:::
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Description {#description}
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								The delivered messages are tracked automatically, so each message in a group is only counted once. If you want to get the data twice, then create a copy of the table with another group name.
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								Groups are flexible and synced on the cluster. For instance, if you have 10 topics and 5 copies of a table in a cluster, then each copy gets 2 topics. If the number of copies changes, the topics are redistributed across the copies automatically. Read more about this at http://kafka.apache.org/intro.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								`SELECT` is not particularly useful for reading messages (except for debugging), because each message can be read only once. It is more practical to create real-time threads using materialized views. To do this:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+.  Use the engine to create a Kafka consumer and consider it a data stream.
 .  Create a table with the desired structure.
 .  Create a materialized view that converts data from the engine and puts it into a previously created table.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												DOCAPI-5758: EN review and RU translation of the Kafka engine descrip… (#4660)

* DOCAPI-5758: EN review and RU translation of the Kafka engine description.

* DOCAPI-5758: Markup fix.

* DOCAPI-5758: Markup fixes in the Kafka topic.

* DOCAPI-5758: Link fix.

											
										
										
											2019-03-15 16:39:59 +00:00
+								When the `MATERIALIZED VIEW` joins the engine, it starts collecting data in the background. This allows you to continually receive messages from Kafka and convert them to the required format using `SELECT`.
-												note about several MV to one kafka table
											
										
										
											2019-06-29 17:09:39 +00:00
+								One kafka table can have as many materialized views as you like, they do not read data from the kafka table directly, but receive new records (in blocks), this way you can write to several tables with different detail level (with grouping - aggregation and without).
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								Example:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE TABLE queue (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');
 								  CREATE TABLE daily (
 								    day Date,
 								    level String,
 								    total UInt64
 								  ) ENGINE = SummingMergeTree(day, (day, level), 8192);
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE MATERIALIZED VIEW consumer TO daily
 								    AS SELECT toDate(toDateTime(timestamp)) AS day, level, count() as total
 								    FROM queue GROUP BY day, level;
 								  SELECT level, sum(total) FROM daily GROUP BY level;
 								```
-												Fix anchors to settings.md

											
										
										
											2023-12-20 18:26:36 +00:00
+								To improve performance, received messages are grouped into blocks the size of [max_insert_block_size](../../../operations/settings/settings.md#max_insert_block_size). If the block wasn’t formed within [stream_flush_interval_ms](../../../operations/settings/settings.md/#stream-flush-interval-ms) milliseconds, the data will be flushed to the table regardless of the completeness of the block.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								To stop receiving topic data or to change the conversion logic, detach the materialized view:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  DETACH TABLE consumer;
-												docs/kafka: use ATTACH TABLE over ATTACH MATERIALIZED VIEW (all langs)

Since later requires full specification (engine and so on).

											
										
										
											2020-04-25 00:08:00 +00:00
+								  ATTACH TABLE consumer;
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
-												Update of english documentation (#2918)

* Updating of english translation.

* Some bugs are fixed.

											
										
										
											2018-09-04 11:18:59 +00:00
+								If you want to change the target table by using `ALTER`, we recommend disabling the material view to avoid discrepancies between the target table and the data from the view.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Configuration {#configuration}
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Change to S3 cfg syntax

											
										
										
											2023-02-23 20:04:41 +00:00
+								Similar to GraphiteMergeTree, the Kafka engine supports extended configuration using the ClickHouse config file. There are two configuration keys that you can use: global (below `<kafka>`) and topic-level (below `<kafka><kafka_topic>`). The global configuration is applied first, and then the topic-level configuration is applied (if it exists).
-												Allow configuration of Kafka topics with periods

The Kafka table engine allows global configuration and per-Kafka-topic
configuration. The latter uses syntax <kafka_TOPIC>, e.g. for topic
"football":

  <kafka_football>
      <retry_backoff_ms>250</retry_backoff_ms>
      <fetch_min_bytes>100000</fetch_min_bytes>
  </kafka_football>

Some users had to find out the hard way that such configuration doesn't
take effect if the topic name contains a period, e.g. "sports.football".
The reason is that ClickHouse configuration framework already uses
periods as level separators to descend the configuration hierarchy.
(Besides that, per-topic configuration at the same level as global
configuration could be considered ugly.)

Note that Kafka topics may contain characters "a-zA-Z0-9._-" (*) and
a tree-like topic organization using periods is quite common in
practice.

This PR deprecates the existing per-topic configuration syntax (but
continues to support it for backward compat) and introduces a new
per-topic configuration syntax below the global Kafka configuration of
the form:

<kafka>
   <topic name="football">
       <retry_backoff_ms>250</retry_backoff_ms>
       <fetch_min_bytes>100000</fetch_min_bytes>
   </topic>
</kafka>

The period restriction doesn't apply to XML attributes, so <topic
name="sports.football"> will work. Also, everything Kafka-related is
below <kafka>.

Considered but rejected alternatives:
- Extending Poco ConfigurationView with custom separators (e.g."/"
  instead of "."). Won't work easily because ConfigurationView only
  builds a path but defers descending the configuration tree to the
  normal configuration classes.
- Reloading the configuration file in StorageKafka (instead of reading
  the loaded file) but with a custom separator. This mode is supported
  by XML configuration. Too ugly and error-prone since the true
  configuration is composed from multiple configuration files.

(*) https://stackoverflow.com/a/37067544

											
										
										
											2023-02-22 19:58:48 +00:00
 								``` xml
 								  <kafka>
-												Slighly improved example

											
										
										
											2023-02-23 20:07:06 +00:00
+								    <!-- Global configuration options for all tables of Kafka engine type -->
-												Allow configuration of Kafka topics with periods

The Kafka table engine allows global configuration and per-Kafka-topic
configuration. The latter uses syntax <kafka_TOPIC>, e.g. for topic
"football":

  <kafka_football>
      <retry_backoff_ms>250</retry_backoff_ms>
      <fetch_min_bytes>100000</fetch_min_bytes>
  </kafka_football>

Some users had to find out the hard way that such configuration doesn't
take effect if the topic name contains a period, e.g. "sports.football".
The reason is that ClickHouse configuration framework already uses
periods as level separators to descend the configuration hierarchy.
(Besides that, per-topic configuration at the same level as global
configuration could be considered ugly.)

Note that Kafka topics may contain characters "a-zA-Z0-9._-" (*) and
a tree-like topic organization using periods is quite common in
practice.

This PR deprecates the existing per-topic configuration syntax (but
continues to support it for backward compat) and introduces a new
per-topic configuration syntax below the global Kafka configuration of
the form:

<kafka>
   <topic name="football">
       <retry_backoff_ms>250</retry_backoff_ms>
       <fetch_min_bytes>100000</fetch_min_bytes>
   </topic>
</kafka>

The period restriction doesn't apply to XML attributes, so <topic
name="sports.football"> will work. Also, everything Kafka-related is
below <kafka>.

Considered but rejected alternatives:
- Extending Poco ConfigurationView with custom separators (e.g."/"
  instead of "."). Won't work easily because ConfigurationView only
  builds a path but defers descending the configuration tree to the
  normal configuration classes.
- Reloading the configuration file in StorageKafka (instead of reading
  the loaded file) but with a custom separator. This mode is supported
  by XML configuration. Too ugly and error-prone since the true
  configuration is composed from multiple configuration files.

(*) https://stackoverflow.com/a/37067544

											
										
										
											2023-02-22 19:58:48 +00:00
+								    <debug>cgrp</debug>
-												refactore: improve reading several configurations for kafka

Simplify and do some refactoring for kafka client settings.

Allows to set up separate
settings for consumer and producer like:

```
<consumer>
    ...
</consumer>

<producer>
    <kafka_topic>
        <name>topic_name</name>
        ...
    </kafka_topic>
</producer>
```

Moreover, this fixes warnings from kafka client like:
`Configuration property session.timeout.ms is a consumer property and
will be ignored by this producer instance`

											
										
										
											2024-03-27 11:19:26 +00:00
+								    <statistics_interval_ms>3000</statistics_interval_ms>
-												docs: added kafka_topic directly under kafka tag

											
										
										
											2024-04-04 11:56:15 +00:00
+								    <kafka_topic>
 								        <name>logs</name>
 								        <statistics_interval_ms>4000</statistics_interval_ms>
 								    </kafka_topic>
-												refactore: improve reading several configurations for kafka

Simplify and do some refactoring for kafka client settings.

Allows to set up separate
settings for consumer and producer like:

```
<consumer>
    ...
</consumer>

<producer>
    <kafka_topic>
        <name>topic_name</name>
        ...
    </kafka_topic>
</producer>
```

Moreover, this fixes warnings from kafka client like:
`Configuration property session.timeout.ms is a consumer property and
will be ignored by this producer instance`

											
										
										
											2024-03-27 11:19:26 +00:00
+								    <!-- Settings for consumer -->
 								    <consumer>
 								        <auto_offset_reset>smallest</auto_offset_reset>
 								        <kafka_topic>
 								            <name>logs</name>
 								            <fetch_min_bytes>100000</fetch_min_bytes>
 								        </kafka_topic>
 								        <kafka_topic>
 								            <name>stats</name>
 								            <fetch_min_bytes>50000</fetch_min_bytes>
 								        </kafka_topic>
 								    </consumer>
 								    <!-- Settings for producer -->
 								    <producer>
 								        <kafka_topic>
 								            <name>logs</name>
 								            <retry_backoff_ms>250</retry_backoff_ms>
 								        </kafka_topic>
 								        <kafka_topic>
 								            <name>stats</name>
 								            <retry_backoff_ms>400</retry_backoff_ms>
 								        </kafka_topic>
 								    </producer>
-												Allow configuration of Kafka topics with periods

The Kafka table engine allows global configuration and per-Kafka-topic
configuration. The latter uses syntax <kafka_TOPIC>, e.g. for topic
"football":

  <kafka_football>
      <retry_backoff_ms>250</retry_backoff_ms>
      <fetch_min_bytes>100000</fetch_min_bytes>
  </kafka_football>

Some users had to find out the hard way that such configuration doesn't
take effect if the topic name contains a period, e.g. "sports.football".
The reason is that ClickHouse configuration framework already uses
periods as level separators to descend the configuration hierarchy.
(Besides that, per-topic configuration at the same level as global
configuration could be considered ugly.)

Note that Kafka topics may contain characters "a-zA-Z0-9._-" (*) and
a tree-like topic organization using periods is quite common in
practice.

This PR deprecates the existing per-topic configuration syntax (but
continues to support it for backward compat) and introduces a new
per-topic configuration syntax below the global Kafka configuration of
the form:

<kafka>
   <topic name="football">
       <retry_backoff_ms>250</retry_backoff_ms>
       <fetch_min_bytes>100000</fetch_min_bytes>
   </topic>
</kafka>

The period restriction doesn't apply to XML attributes, so <topic
name="sports.football"> will work. Also, everything Kafka-related is
below <kafka>.

Considered but rejected alternatives:
- Extending Poco ConfigurationView with custom separators (e.g."/"
  instead of "."). Won't work easily because ConfigurationView only
  builds a path but defers descending the configuration tree to the
  normal configuration classes.
- Reloading the configuration file in StorageKafka (instead of reading
  the loaded file) but with a custom separator. This mode is supported
  by XML configuration. Too ugly and error-prone since the true
  configuration is composed from multiple configuration files.

(*) https://stackoverflow.com/a/37067544

											
										
										
											2023-02-22 19:58:48 +00:00
+								  </kafka>
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
-												Allow configuration of Kafka topics with periods

The Kafka table engine allows global configuration and per-Kafka-topic
configuration. The latter uses syntax <kafka_TOPIC>, e.g. for topic
"football":

  <kafka_football>
      <retry_backoff_ms>250</retry_backoff_ms>
      <fetch_min_bytes>100000</fetch_min_bytes>
  </kafka_football>

Some users had to find out the hard way that such configuration doesn't
take effect if the topic name contains a period, e.g. "sports.football".
The reason is that ClickHouse configuration framework already uses
periods as level separators to descend the configuration hierarchy.
(Besides that, per-topic configuration at the same level as global
configuration could be considered ugly.)

Note that Kafka topics may contain characters "a-zA-Z0-9._-" (*) and
a tree-like topic organization using periods is quite common in
practice.

This PR deprecates the existing per-topic configuration syntax (but
continues to support it for backward compat) and introduces a new
per-topic configuration syntax below the global Kafka configuration of
the form:

<kafka>
   <topic name="football">
       <retry_backoff_ms>250</retry_backoff_ms>
       <fetch_min_bytes>100000</fetch_min_bytes>
   </topic>
</kafka>

The period restriction doesn't apply to XML attributes, so <topic
name="sports.football"> will work. Also, everything Kafka-related is
below <kafka>.

Considered but rejected alternatives:
- Extending Poco ConfigurationView with custom separators (e.g."/"
  instead of "."). Won't work easily because ConfigurationView only
  builds a path but defers descending the configuration tree to the
  normal configuration classes.
- Reloading the configuration file in StorageKafka (instead of reading
  the loaded file) but with a custom separator. This mode is supported
  by XML configuration. Too ugly and error-prone since the true
  configuration is composed from multiple configuration files.

(*) https://stackoverflow.com/a/37067544

											
										
										
											2023-02-22 19:58:48 +00:00
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								For a list of possible configuration options, see the [librdkafka configuration reference](https://github.com/edenhill/librdkafka/blob/master/CONFIGURATION.md). Use the underscore (`_`) instead of a dot in the ClickHouse configuration. For example, `check.crcs=true` will be `<check_crcs>true</check_crcs>`.
-												WIP on docs/website (#3383)

* CLICKHOUSE-4063: less manual html @ index.md

* CLICKHOUSE-4063: recommend markdown="1" in README.md

* CLICKHOUSE-4003: manually purge custom.css for now

* CLICKHOUSE-4064: expand <details> before any print (including to pdf)

* CLICKHOUSE-3927: rearrange interfaces/formats.md a bit

* CLICKHOUSE-3306: add few http headers

* Remove copy-paste introduced in #3392

* Hopefully better chinese fonts #3392

* get rid of tabs @ custom.css

* Apply comments and patch from #3384

* Add jdbc.md to ToC and some translation, though it still looks badly incomplete

* minor punctuation

* Add some backlinks to official website from mirrors that just blindly take markdown sources

* Do not make fonts extra light

* find . -name '*.md' -type f | xargs -I{} perl -pi -e 's//g' {}

* find . -name '*.md' -type f | xargs -I{} perl -pi -e 's/ sql/g' {}

* Remove outdated stuff from roadmap.md

* Not so light font on front page too

* Refactor Chinese formats.md to match recent changes in other languages

											
										
										
											2018-10-16 10:47:17 +00:00
-												Revert "Revert "Test and doc for PR12771 krb5 + cyrus-sasl + kerberized kafka""

This reverts commit c298c633a793e52543f38df78ef0e8098be6f0d6.

											
										
										
											2020-09-29 08:56:37 +00:00
+								### Kerberos support {#kafka-kerberos-support}
 								To deal with Kerberos-aware Kafka, add `security_protocol` child element with `sasl_plaintext` value. It is enough if Kerberos ticket-granting ticket is obtained and cached by OS facilities.
-												Cleanup code in KerberosInit, HDFSCommon and StorageKafka; update English and Russian documentation.

											
										
										
											2022-06-08 14:57:45 +00:00
+								ClickHouse is able to maintain Kerberos credentials using a keytab file. Consider `sasl_kerberos_service_name`, `sasl_kerberos_keytab` and `sasl_kerberos_principal` child elements.
-												Revert "Revert "Test and doc for PR12771 krb5 + cyrus-sasl + kerberized kafka""

This reverts commit c298c633a793e52543f38df78ef0e8098be6f0d6.

											
										
										
											2020-09-29 08:56:37 +00:00
 								Example:
 								``` xml
 								  <!-- Kerberos-aware Kafka -->
 								  <kafka>
 								    <security_protocol>SASL_PLAINTEXT</security_protocol>
 									<sasl_kerberos_keytab>/home/kafkauser/kafkauser.keytab</sasl_kerberos_keytab>
 									<sasl_kerberos_principal>kafkauser/kafkahost@EXAMPLE.COM</sasl_kerberos_principal>
 								  </kafka>
 								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Virtual Columns {#virtual-columns}
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
-												Better docs for virtual columns in Kafka/RabbitMQ/NATS/FileLog

											
										
										
											2023-11-14 21:15:30 +00:00
+								- `_topic` — Kafka topic. Data type: `LowCardinality(String)`.
 								- `_key` — Key of the message. Data type: `String`.
 								- `_offset` — Offset of the message. Data type: `UInt64`.
 								- `_timestamp` — Timestamp of the message Data type: `Nullable(DateTime)`.
 								- `_timestamp_ms` — Timestamp in milliseconds of the message. Data type: `Nullable(DateTime64(3))`.
 								- `_partition` — Partition of Kafka topic. Data type: `UInt64`.
 								- `_headers.name` — Array of message's headers keys. Data type: `Array(String)`.
 								- `_headers.value` — Array of message's headers values. Data type: `Array(String)`.
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
-												Add documentation

											
										
										
											2023-10-11 17:35:18 +00:00
+								Additional virtual columns when `kafka_handle_error_mode='stream'`:
-												Better docs for virtual columns in Kafka/RabbitMQ/NATS/FileLog

											
										
										
											2023-11-14 21:15:30 +00:00
+								- `_raw_message` - Raw message that couldn't be parsed successfully. Data type: `String`.
 								- `_error` - Exception message happened during failed parsing. Data type: `String`.
-												Add documentation

											
										
										
											2023-10-11 17:35:18 +00:00
-												Fix typos

											
										
										
											2023-10-11 17:38:17 +00:00
+								Note: `_raw_message` and `_error` virtual columns are filled only in case of exception during parsing, they are always empty when message was parsed successfully.
-												Add documentation

											
										
										
											2023-10-11 17:35:18 +00:00
-												Improve and refactor Kafka/StorageMQ/NATS and data formats

											
										
										
											2022-10-28 16:41:10 +00:00
+								## Data formats support {#data-formats-support}
 								Kafka engine supports all [formats](../../../interfaces/formats.md) supported in ClickHouse.
 								The number of rows in one Kafka message depends on whether the format is row-based or block-based:
 								- For row-based formats the number of rows in one Kafka message can be controlled by setting `kafka_max_rows_per_message`.
 								- For block-based formats we cannot divide block into smaller parts, but the number of rows in one block can be controlled by general setting [max_block_size](../../../operations/settings/settings.md#setting-max_block_size).
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
+								## Experimental engine to store committed offsets in ClickHouse Keeper
-												Rename experimental flag to `allow_experimental_kafka_offsets_storage_in_keeper`

											
										
										
											2024-07-15 09:03:05 +00:00
+								If `allow_experimental_kafka_offsets_storage_in_keeper` is enabled, then two more settings can be specified to the Kafka table engine:
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
+								 - `kafka_keeper_path` specifies the path to the table in ClickHouse Keeper
 								 - `kafka_replica_name` specifies the replica name in ClickHouse Keeper
-												Address small review comments

											
										
										
											2024-07-31 18:06:09 +00:00
+								Either both of the settings must be specified or neither of them. When both of them are specified, then a new, experimental Kafka engine will be used. The new engine doesn't depend on storing the committed offsets in Kafka, but stores them in ClickHouse Keeper. It still tries to commit the offsets to Kafka, but it only depends on those offsets when the table is created. In any other circumstances (table is restarted, or recovered after some error) the offsets stored in ClickHouse Keeper will be used as an offset to continue consuming messages from. Apart from the committed offset, it also stores how many messages were consumed in the last batch, so if the insert fails, the same amount of messages will be consumed, thus enabling deduplication if necessary.
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
 								Example:
 								``` sql
 								CREATE TABLE experimental_kafka (key UInt64, value UInt64)
 								ENGINE = Kafka('localhost:19092', 'my-topic', 'my-consumer', 'JSONEachRow')
 								SETTINGS
 								  kafka_keeper_path = '/clickhouse/{database}/experimental_kafka',
 								  kafka_replica_name = 'r1'
-												Rename experimental flag to `allow_experimental_kafka_offsets_storage_in_keeper`

											
										
										
											2024-07-15 09:03:05 +00:00
+								SETTINGS allow_experimental_kafka_offsets_storage_in_keeper=1;
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
+								```
 								Or to utilize the `uuid` and `replica` macros similarly to ReplicatedMergeTree:
 								``` sql
 								CREATE TABLE experimental_kafka (key UInt64, value UInt64)
 								ENGINE = Kafka('localhost:19092', 'my-topic', 'my-consumer', 'JSONEachRow')
 								SETTINGS
 								  kafka_keeper_path = '/clickhouse/{database}/{uuid}',
 								  kafka_replica_name = '{replica}'
-												Rename experimental flag to `allow_experimental_kafka_offsets_storage_in_keeper`

											
										
										
											2024-07-15 09:03:05 +00:00
+								SETTINGS allow_experimental_kafka_offsets_storage_in_keeper=1;
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
+								```
 								### Known limitations
 								As the new engine is experimental, it is not production ready yet. There are few known limitations of the implementation:
-												Apply suggestions from code review

Co-authored-by: Kseniia Sumarokova <54203879+kssenii@users.noreply.github.com>
											
										
										
											2024-07-31 17:30:58 +00:00
+								 - The biggest limitation is the engine doesn't support direct reading. Reading from the engine using materialized views and writing to the engine work, but direct reading doesn't. As a result, all direct `SELECT` queries will fail.
-												Improve wording of docs based on review comments

Co-authored-by: Kseniia Sumarokova <54203879+kssenii@users.noreply.github.com>
											
										
										
											2024-07-15 08:15:32 +00:00
+								 - Rapidly dropping and recreating the table or specifying the same ClickHouse Keeper path to different engines might cause issues. As best practice you can use the `{uuid}` in `kafka_keeper_path` to avoid clashing paths.
 								 - To make repeatable reads, messages cannot be consumed from multiple partitions on a single thread. On the other hand, the Kafka consumers have to be polled regularly to keep them alive. As a result of these two objectives, we decided to only allow creating multiple consumers if `kafka_thread_per_consumer` is enabled, otherwise it is too complicated to avoid issues regarding polling consumers regularly.
-												Extend known limitations

											
										
										
											2024-06-24 08:30:37 +00:00
+								 - Consumers created by the new storage engine do not show up in [`system.kafka_consumers`](../../../operations/system-tables/kafka_consumers.md) table.
-												Add minimal docs

											
										
										
											2024-06-18 20:17:32 +00:00
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
+								**See Also**
-												Docs: Replace annoying three spaces in enumerations by a single space

											
										
										
											2023-04-19 15:55:29 +00:00
+								- [Virtual columns](../../../engines/table-engines/index.md#table_engines-virtual_columns)
 								- [background_message_broker_schedule_pool_size](../../../operations/server-configuration-parameters/settings.md#background_message_broker_schedule_pool_size)
-												system_kafka_consumers: style check fixes

											
										
										
											2023-07-02 21:20:20 +00:00
+								- [system.kafka_consumers](../../../operations/system-tables/kafka_consumers.md)