ClickHouse/docs/en/engines/table-engines/integrations/kafka.md

---
slug: /en/engines/table-engines/integrations/kafka
sidebar_position: 8
sidebar_label: Kafka
---

# Kafka

This engine works with [Apache Kafka](http://kafka.apache.org/).

Kafka lets you:

-   Publish or subscribe to data flows.
-   Organize fault-tolerant storage.
-   Process streams as they become available.

## Creating a Table {#table_engine-kafka-creating-a-table}

``` sql
CREATE TABLE [IF NOT EXISTS] [db.]table_name [ON CLUSTER cluster]
(
    name1 [type1] [DEFAULT|MATERIALIZED|ALIAS expr1],
    name2 [type2] [DEFAULT|MATERIALIZED|ALIAS expr2],
    ...
) ENGINE = Kafka()
SETTINGS
    kafka_broker_list = 'host:port',
    kafka_topic_list = 'topic1,topic2,...',
    kafka_group_name = 'group_name',
    kafka_format = 'data_format'[,]
    [kafka_row_delimiter = 'delimiter_symbol',]
    [kafka_schema = '',]
    [kafka_num_consumers = N,]
    [kafka_max_block_size = 0,]
    [kafka_skip_broken_messages = N,]
    [kafka_commit_every_batch = 0,]
    [kafka_thread_per_consumer = 0]
```

Required parameters:

-   `kafka_broker_list` — A comma-separated list of brokers (for example, `localhost:9092`).
-   `kafka_topic_list` — A list of Kafka topics.
-   `kafka_group_name` — A group of Kafka consumers. Reading margins are tracked for each group separately. If you do not want messages to be duplicated in the cluster, use the same group name everywhere.
-   `kafka_format` — Message format. Uses the same notation as the SQL `FORMAT` function, such as `JSONEachRow`. For more information, see the [Formats](../../../interfaces/formats.md) section.

Optional parameters:

-   `kafka_row_delimiter` — Delimiter character, which ends the message.
-   `kafka_schema` — Parameter that must be used if the format requires a schema definition. For example, [Cap’n Proto](https://capnproto.org/) requires the path to the schema file and the name of the root `schema.capnp:Message` object.
-   `kafka_num_consumers` — The number of consumers per table. Default: `1`. Specify more consumers if the throughput of one consumer is insufficient. The total number of consumers should not exceed the number of partitions in the topic, since only one consumer can be assigned per partition, and must not be greater than the number of physical cores on the server where ClickHouse is deployed.
-   `kafka_max_block_size` — The maximum batch size (in messages) for poll (default: `max_block_size`).
-   `kafka_skip_broken_messages` — Kafka message parser tolerance to schema-incompatible messages per block. Default: `0`. If `kafka_skip_broken_messages = N` then the engine skips *N* Kafka messages that cannot be parsed (a message equals a row of data).
-   `kafka_commit_every_batch` — Commit every consumed and handled batch instead of a single commit after writing a whole block (default: `0`).
-   `kafka_thread_per_consumer` — Provide independent thread for each consumer (default: `0`). When enabled, every consumer flush the data independently, in parallel (otherwise — rows from several consumers squashed to form one block).

Examples:

``` sql
  CREATE TABLE queue (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');

  SELECT * FROM queue LIMIT 5;

  CREATE TABLE queue2 (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka SETTINGS kafka_broker_list = 'localhost:9092',
                            kafka_topic_list = 'topic',
                            kafka_group_name = 'group1',
                            kafka_format = 'JSONEachRow',
                            kafka_num_consumers = 4;

  CREATE TABLE queue3 (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1')
              SETTINGS kafka_format = 'JSONEachRow',
                       kafka_num_consumers = 4;
```

<details markdown="1">

<summary>Deprecated Method for Creating a Table</summary>

:::warning
Do not use this method in new projects. If possible, switch old projects to the method described above.
:::

``` sql
Kafka(kafka_broker_list, kafka_topic_list, kafka_group_name, kafka_format
      [, kafka_row_delimiter, kafka_schema, kafka_num_consumers, kafka_skip_broken_messages])
```

</details>

## Description {#description}

The delivered messages are tracked automatically, so each message in a group is only counted once. If you want to get the data twice, then create a copy of the table with another group name.

Groups are flexible and synced on the cluster. For instance, if you have 10 topics and 5 copies of a table in a cluster, then each copy gets 2 topics. If the number of copies changes, the topics are redistributed across the copies automatically. Read more about this at http://kafka.apache.org/intro.

`SELECT` is not particularly useful for reading messages (except for debugging), because each message can be read only once. It is more practical to create real-time threads using materialized views. To do this:

1.  Use the engine to create a Kafka consumer and consider it a data stream.
2.  Create a table with the desired structure.
3.  Create a materialized view that converts data from the engine and puts it into a previously created table.

When the `MATERIALIZED VIEW` joins the engine, it starts collecting data in the background. This allows you to continually receive messages from Kafka and convert them to the required format using `SELECT`.
One kafka table can have as many materialized views as you like, they do not read data from the kafka table directly, but receive new records (in blocks), this way you can write to several tables with different detail level (with grouping - aggregation and without).

Example:

``` sql
  CREATE TABLE queue (
    timestamp UInt64,
    level String,
    message String
  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');

  CREATE TABLE daily (
    day Date,
    level String,
    total UInt64
  ) ENGINE = SummingMergeTree(day, (day, level), 8192);

  CREATE MATERIALIZED VIEW consumer TO daily
    AS SELECT toDate(toDateTime(timestamp)) AS day, level, count() as total
    FROM queue GROUP BY day, level;

  SELECT level, sum(total) FROM daily GROUP BY level;
```
To improve performance, received messages are grouped into blocks the size of [max_insert_block_size](../../../operations/settings/settings.md#settings-max_insert_block_size). If the block wasn’t formed within [stream_flush_interval_ms](../../../operations/settings/settings.md/#stream-flush-interval-ms) milliseconds, the data will be flushed to the table regardless of the completeness of the block.

To stop receiving topic data or to change the conversion logic, detach the materialized view:

``` sql
  DETACH TABLE consumer;
  ATTACH TABLE consumer;
```

If you want to change the target table by using `ALTER`, we recommend disabling the material view to avoid discrepancies between the target table and the data from the view.

## Configuration {#configuration}

Similar to GraphiteMergeTree, the Kafka engine supports extended configuration using the ClickHouse config file. There are two configuration keys that you can use: global (`kafka`) and topic-level (`kafka_*`). The global configuration is applied first, and then the topic-level configuration is applied (if it exists).

``` xml
  <!-- Global configuration options for all tables of Kafka engine type -->
  <kafka>
    <debug>cgrp</debug>
    <auto_offset_reset>smallest</auto_offset_reset>
  </kafka>

  <!-- Configuration specific for topic "logs" -->
  <kafka_logs>
    <retry_backoff_ms>250</retry_backoff_ms>
    <fetch_min_bytes>100000</fetch_min_bytes>
  </kafka_logs>
```

For a list of possible configuration options, see the [librdkafka configuration reference](https://github.com/edenhill/librdkafka/blob/master/CONFIGURATION.md). Use the underscore (`_`) instead of a dot in the ClickHouse configuration. For example, `check.crcs=true` will be `<check_crcs>true</check_crcs>`.

### Kerberos support {#kafka-kerberos-support}

To deal with Kerberos-aware Kafka, add `security_protocol` child element with `sasl_plaintext` value. It is enough if Kerberos ticket-granting ticket is obtained and cached by OS facilities.
ClickHouse is able to maintain Kerberos credentials using a keytab file. Consider `sasl_kerberos_service_name`, `sasl_kerberos_keytab` and `sasl_kerberos_principal` child elements.

Example:

``` xml
  <!-- Kerberos-aware Kafka -->
  <kafka>
    <security_protocol>SASL_PLAINTEXT</security_protocol>
	<sasl_kerberos_keytab>/home/kafkauser/kafkauser.keytab</sasl_kerberos_keytab>
	<sasl_kerberos_principal>kafkauser/kafkahost@EXAMPLE.COM</sasl_kerberos_principal>
  </kafka>
```

## Virtual Columns {#virtual-columns}

-   `_topic` — Kafka topic.
-   `_key` — Key of the message.
-   `_offset` — Offset of the message.
-   `_timestamp` — Timestamp of the message.
-   `_timestamp_ms` — Timestamp in milliseconds of the message.
-   `_partition` — Partition of Kafka topic.
-   `_headers.name` — Array of message's headers keys.
-   `_headers.value` — Array of message's headers values.

**See Also**

-   [Virtual columns](../../../engines/table-engines/index.md#table_engines-virtual_columns)
-   [background_message_broker_schedule_pool_size](../../../operations/settings/settings.md#background_message_broker_schedule_pool_size)

[Original article](https://clickhouse.com/docs/en/engines/table-engines/integrations/kafka/) <!--hide-->
-												Get rid of toc_en.yml (#10023)


											
										
										
											2020-04-03 13:23:32 +00:00
+								---
-												add slugs

											
										
										
											2022-08-28 14:53:34 +00:00
+								slug: /en/engines/table-engines/integrations/kafka
-												Removed /ja folder, cleaned up /ru markdown

											
										
										
											2022-04-09 13:29:05 +00:00
+								sidebar_position: 8
 								sidebar_label: Kafka
-												Get rid of toc_en.yml (#10023)


											
										
										
											2020-04-03 13:23:32 +00:00
+								---
-												Remove H1 anchor tags from docs

											
										
										
											2022-06-02 10:55:18 +00:00
+								# Kafka
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								This engine works with [Apache Kafka](http://kafka.apache.org/).
 								Kafka lets you:
-												[experimental] add "es" docs language as machine translated draft (#9787)

* replace exit with assert in test_single_page

* improve save_raw_single_page docs option

* More grammar fixes

* "Built from" link in new tab

* fix mistype

* Example of include in docs

* add anchor to meeting form

* Draft of translation helper

* WIP on translation helper

* Replace some fa docs content with machine translation

* add normalize-en-markdown.sh

* normalize some en markdown

* normalize some en markdown

* admonition support

* normalize

* normalize

* normalize

* support wide tables

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* lightly edited machine translation of introdpection.md

* lightly edited machhine translation of lazy.md

* WIP on translation utils

* Normalize ru docs

* Normalize other languages

* some fixes

* WIP on normalize/translate tools

* add requirements.txt

* [experimental] add es docs language as machine translated draft

* remove duplicate script

* Back to wider tab-stop (narrow renders not so well)
											
										
										
											2020-03-21 04:11:51 +00:00
+								-   Publish or subscribe to data flows.
 								-   Organize fault-tolerant storage.
 								-   Process streams as they become available.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Restore some old manual anchors in docs (#9803)

* Simplify 404 page

* add es array_functions.md

* restore some old manual anchors

* update sitemaps

* trigger checks

* restore more old manual anchors

* refactor test.md + temporary disable failure again

* fix mistype
											
										
										
											2020-03-22 09:14:59 +00:00
+								## Creating a Table {#table_engine-kafka-creating-a-table}
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								CREATE TABLE [IF NOT EXISTS] [db.]table_name [ON CLUSTER cluster]
 								(
 								    name1 [type1] [DEFAULT|MATERIALIZED|ALIAS expr1],
 								    name2 [type2] [DEFAULT|MATERIALIZED|ALIAS expr2],
 								    ...
 								) ENGINE = Kafka()
 								SETTINGS
 								    kafka_broker_list = 'host:port',
 								    kafka_topic_list = 'topic1,topic2,...',
 								    kafka_group_name = 'group_name',
 								    kafka_format = 'data_format'[,]
 								    [kafka_row_delimiter = 'delimiter_symbol',]
 								    [kafka_schema = '',]
 								    [kafka_num_consumers = N,]
-												Add missing kafka settings into docs

											
										
										
											2020-04-27 05:02:45 +00:00
+								    [kafka_max_block_size = 0,]
 								    [kafka_skip_broken_messages = N,]
-												Fix code style, and update docs for Kafka engine

											
										
										
											2020-09-01 08:37:12 +00:00
+								    [kafka_commit_every_batch = 0,]
 								    [kafka_thread_per_consumer = 0]
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								Required parameters:
-												Edit and translated Kafka

											
										
										
											2021-02-11 18:07:38 +00:00
+								-   `kafka_broker_list` — A comma-separated list of brokers (for example, `localhost:9092`).
 								-   `kafka_topic_list` — A list of Kafka topics.
-												Avoid short syntax

											
										
										
											2021-05-27 19:44:11 +00:00
+								-   `kafka_group_name` — A group of Kafka consumers. Reading margins are tracked for each group separately. If you do not want messages to be duplicated in the cluster, use the same group name everywhere.
-												Edit and translated Kafka

											
										
										
											2021-02-11 18:07:38 +00:00
+								-   `kafka_format` — Message format. Uses the same notation as the SQL `FORMAT` function, such as `JSONEachRow`. For more information, see the [Formats](../../../interfaces/formats.md) section.
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								Optional parameters:
-												Edit and translated Kafka

											
										
										
											2021-02-11 18:07:38 +00:00
+								-   `kafka_row_delimiter` — Delimiter character, which ends the message.
 								-   `kafka_schema` — Parameter that must be used if the format requires a schema definition. For example, [Cap’n Proto](https://capnproto.org/) requires the path to the schema file and the name of the root `schema.capnp:Message` object.
-												`kafka_num_consumers` prop upper bound doc update

Add clarification about the upper bound of `kafka_num_consumers` property that was added in [#26640](https://github.com/ClickHouse/ClickHouse/pull/26642)
											
										
										
											2022-03-12 08:09:12 +00:00
+								-   `kafka_num_consumers` — The number of consumers per table. Default: `1`. Specify more consumers if the throughput of one consumer is insufficient. The total number of consumers should not exceed the number of partitions in the topic, since only one consumer can be assigned per partition, and must not be greater than the number of physical cores on the server where ClickHouse is deployed.
-												Edit and translated Kafka

											
										
										
											2021-02-11 18:07:38 +00:00
+								-   `kafka_max_block_size` — The maximum batch size (in messages) for poll (default: `max_block_size`).
 								-   `kafka_skip_broken_messages` — Kafka message parser tolerance to schema-incompatible messages per block. Default: `0`. If `kafka_skip_broken_messages = N` then the engine skips *N* Kafka messages that cannot be parsed (a message equals a row of data).
 								-   `kafka_commit_every_batch` — Commit every consumed and handled batch instead of a single commit after writing a whole block (default: `0`).
 								-   `kafka_thread_per_consumer` — Provide independent thread for each consumer (default: `0`). When enabled, every consumer flush the data independently, in parallel (otherwise — rows from several consumers squashed to form one block).
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								Examples:
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE TABLE queue (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');
 								  SELECT * FROM queue LIMIT 5;
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
 								  CREATE TABLE queue2 (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka SETTINGS kafka_broker_list = 'localhost:9092',
 								                            kafka_topic_list = 'topic',
 								                            kafka_group_name = 'group1',
 								                            kafka_format = 'JSONEachRow',
 								                            kafka_num_consumers = 4;
-												table name typo fix (#11100)

According to the context, the table name of queue2 should be queue3.
											
										
										
											2020-05-21 12:14:39 +00:00
+								  CREATE TABLE queue3 (
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
+								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1')
 								              SETTINGS kafka_format = 'JSONEachRow',
 								                       kafka_num_consumers = 4;
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								<details markdown="1">
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								<summary>Deprecated Method for Creating a Table</summary>
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												Remove H1 anchor tags from docs

											
										
										
											2022-06-02 10:55:18 +00:00
+								:::warning
-												Removed /ja folder, cleaned up /ru markdown

											
										
										
											2022-04-09 13:29:05 +00:00
+								Do not use this method in new projects. If possible, switch old projects to the method described above.
 								:::
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
+								Kafka(kafka_broker_list, kafka_topic_list, kafka_group_name, kafka_format
 								      [, kafka_row_delimiter, kafka_schema, kafka_num_consumers, kafka_skip_broken_messages])
 								```
 								</details>
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Description {#description}
-												DOCAPI-5758: New parameter kafka_skip_broken_messages in the Kafka engine description.

											
										
										
											2019-03-04 16:27:00 +00:00
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								The delivered messages are tracked automatically, so each message in a group is only counted once. If you want to get the data twice, then create a copy of the table with another group name.
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								Groups are flexible and synced on the cluster. For instance, if you have 10 topics and 5 copies of a table in a cluster, then each copy gets 2 topics. If the number of copies changes, the topics are redistributed across the copies automatically. Read more about this at http://kafka.apache.org/intro.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								`SELECT` is not particularly useful for reading messages (except for debugging), because each message can be read only once. It is more practical to create real-time threads using materialized views. To do this:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+.  Use the engine to create a Kafka consumer and consider it a data stream.
 .  Create a table with the desired structure.
 .  Create a materialized view that converts data from the engine and puts it into a previously created table.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												DOCAPI-5758: EN review and RU translation of the Kafka engine descrip… (#4660)

* DOCAPI-5758: EN review and RU translation of the Kafka engine description.

* DOCAPI-5758: Markup fix.

* DOCAPI-5758: Markup fixes in the Kafka topic.

* DOCAPI-5758: Link fix.

											
										
										
											2019-03-15 16:39:59 +00:00
+								When the `MATERIALIZED VIEW` joins the engine, it starts collecting data in the background. This allows you to continually receive messages from Kafka and convert them to the required format using `SELECT`.
-												note about several MV to one kafka table
											
										
										
											2019-06-29 17:09:39 +00:00
+								One kafka table can have as many materialized views as you like, they do not read data from the kafka table directly, but receive new records (in blocks), this way you can write to several tables with different detail level (with grouping - aggregation and without).
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								Example:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE TABLE queue (
 								    timestamp UInt64,
 								    level String,
 								    message String
 								  ) ENGINE = Kafka('localhost:9092', 'topic', 'group1', 'JSONEachRow');
 								  CREATE TABLE daily (
 								    day Date,
 								    level String,
 								    total UInt64
 								  ) ENGINE = SummingMergeTree(day, (day, level), 8192);
-												Added SETTINGS clause for Kafka storage engine

											
										
										
											2018-08-01 17:23:50 +00:00
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  CREATE MATERIALIZED VIEW consumer TO daily
 								    AS SELECT toDate(toDateTime(timestamp)) AS day, level, count() as total
 								    FROM queue GROUP BY day, level;
 								  SELECT level, sum(total) FROM daily GROUP BY level;
 								```
-												Removed /ja folder, cleaned up /ru markdown

											
										
										
											2022-04-09 13:29:05 +00:00
+								To improve performance, received messages are grouped into blocks the size of [max_insert_block_size](../../../operations/settings/settings.md#settings-max_insert_block_size). If the block wasn’t formed within [stream_flush_interval_ms](../../../operations/settings/settings.md/#stream-flush-interval-ms) milliseconds, the data will be flushed to the table regardless of the completeness of the block.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
 								To stop receiving topic data or to change the conversion logic, detach the materialized view:
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` sql
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  DETACH TABLE consumer;
-												docs/kafka: use ATTACH TABLE over ATTACH MATERIALIZED VIEW (all langs)

Since later requires full specification (engine and so on).

											
										
										
											2020-04-25 00:08:00 +00:00
+								  ATTACH TABLE consumer;
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
-												Update of english documentation (#2918)

* Updating of english translation.

* Some bugs are fixed.

											
										
										
											2018-09-04 11:18:59 +00:00
+								If you want to change the target table by using `ALTER`, we recommend disabling the material view to avoid discrepancies between the target table and the data from the view.
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Configuration {#configuration}
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												fix documentation kafka per topic configuration

											
										
										
											2018-09-18 12:59:12 +00:00
+								Similar to GraphiteMergeTree, the Kafka engine supports extended configuration using the ClickHouse config file. There are two configuration keys that you can use: global (`kafka`) and topic-level (`kafka_*`). The global configuration is applied first, and then the topic-level configuration is applied (if it exists).
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								``` xml
-												Doc fixes: remove double placeholders; add them where missing. (#3923)

* Doc fix: add spaces where missing

* Doc fixes: rm double spaces

* Doc fixes: edit spaces

* Doc fixes: rm double spaces in /fa

* Revert "Doc fixes: rm double spaces in /fa"

This reverts commit bb879a62ef5fa965d989fea3b1b2a693d2016a2d.

* Doc fix: resolve all problems with double spaces in /fa

* Doc fix: add spaces for readability

* Doc fix: add spaces

* Fix spaces

											
										
										
											2018-12-25 15:25:43 +00:00
+								  <!-- Global configuration options for all tables of Kafka engine type -->
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								  <kafka>
 								    <debug>cgrp</debug>
 								    <auto_offset_reset>smallest</auto_offset_reset>
 								  </kafka>
 								  <!-- Configuration specific for topic "logs" -->
-												fix documentation kafka per topic configuration

											
										
										
											2018-09-18 12:59:12 +00:00
+								  <kafka_logs>
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								    <retry_backoff_ms>250</retry_backoff_ms>
 								    <fetch_min_bytes>100000</fetch_min_bytes>
-												fix documentation kafka per topic configuration

											
										
										
											2018-09-18 12:59:12 +00:00
+								  </kafka_logs>
-												English translation is updated.

											
										
										
											2018-04-23 06:20:21 +00:00
+								```
 								For a list of possible configuration options, see the [librdkafka configuration reference](https://github.com/edenhill/librdkafka/blob/master/CONFIGURATION.md). Use the underscore (`_`) instead of a dot in the ClickHouse configuration. For example, `check.crcs=true` will be `<check_crcs>true</check_crcs>`.
-												WIP on docs/website (#3383)

* CLICKHOUSE-4063: less manual html @ index.md

* CLICKHOUSE-4063: recommend markdown="1" in README.md

* CLICKHOUSE-4003: manually purge custom.css for now

* CLICKHOUSE-4064: expand <details> before any print (including to pdf)

* CLICKHOUSE-3927: rearrange interfaces/formats.md a bit

* CLICKHOUSE-3306: add few http headers

* Remove copy-paste introduced in #3392

* Hopefully better chinese fonts #3392

* get rid of tabs @ custom.css

* Apply comments and patch from #3384

* Add jdbc.md to ToC and some translation, though it still looks badly incomplete

* minor punctuation

* Add some backlinks to official website from mirrors that just blindly take markdown sources

* Do not make fonts extra light

* find . -name '*.md' -type f | xargs -I{} perl -pi -e 's//g' {}

* find . -name '*.md' -type f | xargs -I{} perl -pi -e 's/ sql/g' {}

* Remove outdated stuff from roadmap.md

* Not so light font on front page too

* Refactor Chinese formats.md to match recent changes in other languages

											
										
										
											2018-10-16 10:47:17 +00:00
-												Revert "Revert "Test and doc for PR12771 krb5 + cyrus-sasl + kerberized kafka""

This reverts commit c298c633a793e52543f38df78ef0e8098be6f0d6.

											
										
										
											2020-09-29 08:56:37 +00:00
+								### Kerberos support {#kafka-kerberos-support}
 								To deal with Kerberos-aware Kafka, add `security_protocol` child element with `sasl_plaintext` value. It is enough if Kerberos ticket-granting ticket is obtained and cached by OS facilities.
-												Cleanup code in KerberosInit, HDFSCommon and StorageKafka; update English and Russian documentation.

											
										
										
											2022-06-08 14:57:45 +00:00
+								ClickHouse is able to maintain Kerberos credentials using a keytab file. Consider `sasl_kerberos_service_name`, `sasl_kerberos_keytab` and `sasl_kerberos_principal` child elements.
-												Revert "Revert "Test and doc for PR12771 krb5 + cyrus-sasl + kerberized kafka""

This reverts commit c298c633a793e52543f38df78ef0e8098be6f0d6.

											
										
										
											2020-09-29 08:56:37 +00:00
 								Example:
 								``` xml
 								  <!-- Kerberos-aware Kafka -->
 								  <kafka>
 								    <security_protocol>SASL_PLAINTEXT</security_protocol>
 									<sasl_kerberos_keytab>/home/kafkauser/kafkauser.keytab</sasl_kerberos_keytab>
 									<sasl_kerberos_principal>kafkauser/kafkahost@EXAMPLE.COM</sasl_kerberos_principal>
 								  </kafka>
 								```
-												Normalization for en markdown (#9763)


											
										
										
											2020-03-20 10:10:48 +00:00
+								## Virtual Columns {#virtual-columns}
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
-												[experimental] add "es" docs language as machine translated draft (#9787)

* replace exit with assert in test_single_page

* improve save_raw_single_page docs option

* More grammar fixes

* "Built from" link in new tab

* fix mistype

* Example of include in docs

* add anchor to meeting form

* Draft of translation helper

* WIP on translation helper

* Replace some fa docs content with machine translation

* add normalize-en-markdown.sh

* normalize some en markdown

* normalize some en markdown

* admonition support

* normalize

* normalize

* normalize

* support wide tables

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* lightly edited machine translation of introdpection.md

* lightly edited machhine translation of lazy.md

* WIP on translation utils

* Normalize ru docs

* Normalize other languages

* some fixes

* WIP on normalize/translate tools

* add requirements.txt

* [experimental] add es docs language as machine translated draft

* remove duplicate script

* Back to wider tab-stop (narrow renders not so well)
											
										
										
											2020-03-21 04:11:51 +00:00
+								-   `_topic` — Kafka topic.
 								-   `_key` — Key of the message.
 								-   `_offset` — Offset of the message.
 								-   `_timestamp` — Timestamp of the message.
-												update document: add _timestamp_ms of kafka

											
										
										
											2022-02-19 15:49:37 +00:00
+								-   `_timestamp_ms` — Timestamp in milliseconds of the message.
-												[experimental] add "es" docs language as machine translated draft (#9787)

* replace exit with assert in test_single_page

* improve save_raw_single_page docs option

* More grammar fixes

* "Built from" link in new tab

* fix mistype

* Example of include in docs

* add anchor to meeting form

* Draft of translation helper

* WIP on translation helper

* Replace some fa docs content with machine translation

* add normalize-en-markdown.sh

* normalize some en markdown

* normalize some en markdown

* admonition support

* normalize

* normalize

* normalize

* support wide tables

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* normalize

* lightly edited machine translation of introdpection.md

* lightly edited machhine translation of lazy.md

* WIP on translation utils

* Normalize ru docs

* Normalize other languages

* some fixes

* WIP on normalize/translate tools

* add requirements.txt

* [experimental] add es docs language as machine translated draft

* remove duplicate script

* Back to wider tab-stop (narrow renders not so well)
											
										
										
											2020-03-21 04:11:51 +00:00
+								-   `_partition` — Partition of Kafka topic.
-												Add Kafka headers virtual columns

Add documentation for https://github.com/ClickHouse/ClickHouse/pull/11283
											
										
										
											2022-05-11 14:38:09 +00:00
+								-   `_headers.name` — Array of message's headers keys.
 								-   `_headers.value` — Array of message's headers values.
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
 								**See Also**
-												[docs] split aggregate function and system table references (#11742)

* prefer relative links from root

* wip

* split aggregate function reference

* split system tables
											
										
										
											2020-06-18 08:24:31 +00:00
+								-   [Virtual columns](../../../engines/table-engines/index.md#table_engines-virtual_columns)
-												Update kafka.md

Updated See Also section, since number of kafka consumers controlled by background_message_broker_schedule_pool_size setting
											
										
										
											2021-11-25 12:20:16 +00:00
+								-   [background_message_broker_schedule_pool_size](../../../operations/settings/settings.md#background_message_broker_schedule_pool_size)
-												DOCAPI-7443: Virtual columns docs update. (#6382)


											
										
										
											2019-08-14 06:45:24 +00:00
-												find . -type f -name '*.md'| xargs -I{} perl -pi -e 's|https://clickhouse.tech|https://clickhouse.com|g' {}

											
										
										
											2021-09-19 20:05:54 +00:00
+								[Original article](https://clickhouse.com/docs/en/engines/table-engines/integrations/kafka/) <!--hide-->