
You can move data from MySQL to ClickHouse using change data capture (CDC), scheduled batch extraction, ClickHouse’s native ClickPipes integration, a managed data integration platform, or a custom streaming stack.
For operational MySQL databases that change continuously, binlog-based CDC is usually the most appropriate approach because it captures inserts, updates, and deletes without repeatedly scanning entire source tables. Batch extraction can still be a better fit when data only needs to refresh periodically.
ClickHouse Cloud now provides native MySQL CDC through ClickPipes. Estuary is another managed option for this route, capturing changes from the MySQL binary log and materializing them directly into ClickHouse Cloud or self-hosted ClickHouse. This guide explains how the approaches differ, what MySQL and ClickHouse require, and the production considerations that matter when keeping the two systems in sync.
MySQL to ClickHouse: Which Approach Should You Use?
| Approach | Best fit | Data freshness | Main tradeoff |
|---|---|---|---|
| ClickHouse ClickPipes MySQL CDC | ClickHouse Cloud users wanting first-party managed CDC | Continuous | ClickHouse Cloud specific |
| Managed CDC platform such as Estuary | Teams that want managed CDC across ClickHouse Cloud, self-hosted ClickHouse, or multiple destinations | Continuous or scheduled | Adds a separate data integration layer |
| Debezium + Kafka | Teams already operating Kafka-based streaming infrastructure | Continuous | Higher infrastructure and operational responsibility |
| Batch ETL / ELT | Periodic reporting where hourly or daily freshness is enough | Scheduled | Destination is only as fresh as the latest batch |
| Custom pipeline | Specialized workloads with dedicated engineering ownership | Depends | Higher development and maintenance burden |
ClickHouse also provides a MySQL table function and table engine that can query remote MySQL data directly. These are useful for federation and ad hoc access, but they are different from maintaining a continuously replicated analytical copy inside ClickHouse.
Quick recommendation: Use CDC when ClickHouse needs continuously fresh MySQL changes, including updates and deletes. Use batch when hourly or daily refreshes are sufficient. Native ClickPipes is a strong fit for ClickHouse Cloud-only environments, while a managed integration platform can make more sense when MySQL data also needs to serve other destinations.
For details on ClickHouse’s remote MySQL access, see ClickHouse’s MySQL integration documentation.
How MySQL to ClickHouse CDC Works
A CDC pipeline typically reads changes from MySQL’s binary log rather than repeatedly querying entire tables.
A simplified architecture looks like this:
MySQL → Binary Log → CDC Capture → Processing / Checkpointing → ClickHouse
There are usually two phases.
1. Initial load
Before ongoing replication begins, you need to copy the existing state of the selected MySQL tables into ClickHouse.
Depending on the platform, this may be called an:
- initial snapshot
- backfill
- initial sync
- historical load
A reliable CDC system coordinates this initial copy with the source change stream so it doesn't silently miss changes during the backfill.
2. Ongoing change capture
Once the initial state is available, the CDC process reads new row-level events from the MySQL binary log.
That typically includes:
- INSERT
- UPDATE
- DELETE
The replication layer keeps track of its source position, processes new changes, maps them into the destination representation, and advances its checkpoint after successful processing.
The exact destination behavior matters because ClickHouse is an analytical columnar database, not a row-oriented transactional database. Updates and deletes are therefore often represented through new row versions, background merges, or deletion markers rather than immediate in-place row mutation.
MySQL Requirements for CDC
The exact prerequisites depend on the replication implementation, but several requirements are common across MySQL CDC systems.
Use row-based binary logging
CDC systems need row-level changes, not just the original SQL statement.
For Estuary’s MySQL CDC connector, binlog_format must be:
ROW
Estuary’s connector reads source changes directly from the MySQL binary log.
See the Estuary MySQL CDC connector documentation for current source requirements.
Give the CDC user the required permissions
The replication user needs enough access to read the source changes and selected tables.
For Estuary, this includes:
- REPLICATION CLIENT
- REPLICATION SLAVE
- SELECT access to captured tables
- information_schema access when automatic discovery is used
Avoid granting broader database permissions than the pipeline needs.
Keep enough binary log history
CDC systems can only resume from a previous position while the required binary log history still exists.
Estuary recommends retaining at least seven days of MySQL binary logs. A shorter window is technically possible through an advanced override, but the documentation discourages it because losing the required binlog segment can force a new backfill.
Managed services can impose their own limits. For example, Amazon RDS for MySQL documents 168 hours as the maximum retention used by the Estuary setup.
The practical rule is simple:
Your binlog retention window should be longer than the longest pipeline interruption you realistically need to recover from.
Handle MySQL DATETIME time zones explicitly
MySQL DATETIME values do not inherently identify their timezone.
If captured tables contain DATETIME columns, Estuary requires either:
- an explicit MySQL time_zone, or
- a timezone configured in the connector.
This avoids ambiguous timestamp interpretation downstream.
Consider full row metadata if schema tracking becomes unreliable
This is not a prerequisite for every deployment.
However, if the same table repeatedly encounters metadata inconsistencies, Estuary recommends:
binlog_row_metadata = FULL
Without full metadata, the connector can rely more heavily on DDL parsing to track column changes. Some DDL operations may not appear in the binlog in a way that provides enough metadata to reliably reconstruct the new table definition.
What to Consider on the ClickHouse Side
Moving MySQL data into ClickHouse is not simply a table-copy exercise.
Before selecting an ingestion method, understand how the destination will handle:
- source primary keys
- updates
- deletes
- duplicate row versions
- schema changes
- initial backfills
- analytical query semantics
This matters because ClickHouse commonly handles changing records through insert-based patterns and background merges rather than traditional OLTP-style in-place updates.
Ways to Move Data From MySQL to ClickHouse
Option 1: ClickHouse ClickPipes MySQL CDC
For teams already standardized on ClickHouse Cloud, native MySQL CDC through ClickPipes is now a significant option.
ClickHouse made its MySQL CDC connector generally available on September 2, 2026. The GA release added parallel snapshotting, improved observability and reliability, and programmatic configuration through the ClickHouse Cloud API and Terraform.
See the ClickHouse MySQL CDC ClickPipes GA announcement for the current native implementation.
New MySQL CDC ClickPipes validate GTID-based replication and require at least 72 hours of binlog retention. ClickHouse explains that GTIDs make replication positions less dependent on a specific server, improving recovery during failovers. Existing file-position-based ClickPipes can continue to coexist with newer GTID-based pipes.
ClickPipes is particularly attractive when:
- ClickHouse Cloud is your destination.
- You want the ingestion layer managed within ClickHouse.
- You do not need the same capture layer to feed several unrelated destinations.
Option 2: A Managed CDC Platform
A managed data integration platform sits between MySQL and ClickHouse and handles much of the operational work associated with capture, checkpointing, recovery, backfills, and destination delivery.
This approach can make sense when:
- ClickHouse is one of several downstream destinations.
- You use self-hosted ClickHouse and ClickHouse Cloud.
- You want one capture of MySQL to serve multiple downstream systems.
- You operate many database and SaaS integrations.
- Some pipelines require CDC while others need scheduled extraction.
Estuary is one implementation of this architecture.
Option 3: Debezium + Kafka
A common self-managed architecture is:
MySQL → Debezium → Kafka → ClickHouse
This approach gives engineering teams significant control over the change stream and works well when Kafka already plays a central role in the data platform.
The tradeoff is operational complexity.
Your team becomes responsible for components such as:
- Debezium connectors
- Kafka brokers
- topic configuration
- consumer lag
- schema handling
- connector recovery
- infrastructure monitoring
If Kafka is already a strategic platform, that tradeoff may be reasonable.
Introducing Kafka solely to connect a single MySQL database to ClickHouse may be harder to justify.
Option 4: Scheduled Batch Extraction
Not every MySQL-to-ClickHouse workload needs CDC.
Batch extraction can be simpler when:
- analytics refresh hourly or daily
- source changes are relatively infrequent
- immediate delete propagation is unnecessary
- freshness is less important than pipeline simplicity
The main tradeoff is freshness and source efficiency.
As tables grow, repeatedly querying the source for changed or historical records can also place more analytical workload on the MySQL database than log-based CDC.
Estuary provides a separate MySQL Batch Query connector for cases where CDC is not appropriate. For example, when binary-log replication is unavailable, views need to be captured, or custom SQL queries are required.
See the Estuary MySQL Batch Query connector documentation for those use cases.
Initial Loads, Inserts, Updates, and Deletes
Understanding runtime behavior is more important than simply getting the first records into ClickHouse.
Initial load
A production CDC pipeline usually needs to establish the existing source state before it can accurately represent future changes.
For very large databases, this initial backfill should be planned separately from steady-state CDC because it can generate substantially more source reads and destination writes.
Inserts
New MySQL rows are captured from the binary log and written downstream.
Updates
Updates result in a new representation of the changed record.
With ClickHouse, newer versions may coexist physically with older versions for some period depending on the table engine and ingestion method.
Deletes
Delete handling varies by implementation.
Some pipelines retain the existing row and attach deletion metadata.
Others write a tombstone or deletion marker that ClickHouse uses to exclude or eventually remove the record.
This means you should evaluate delete semantics explicitly when comparing MySQL-to-ClickHouse CDC options rather than simply checking whether a vendor says it “supports deletes.”
Schema Changes and DDL
MySQL schemas change over time, so a production replication pipeline needs defined behavior for DDL operations.
With Estuary’s MySQL CDC connector:
- ALTER TABLE operations that add, drop, rename, or change the type of columns are handled automatically when the change appears in the binlog.
- DROP TABLE, RENAME TABLE, and DROP DATABASE deactivate the affected capture binding.
- If a dropped or renamed table is later recreated, the connector automatically backfills it again and resumes capture.
- TRUNCATE TABLE is different: the truncate is ignored as a set of row deletes. If ClickHouse must reflect the truncated state, the binding needs to be backfilled again.
This last behavior is especially important for pipelines where TRUNCATE is part of regular application or maintenance workflows.
Production Considerations for MySQL to ClickHouse
Monitor replication lag
For a CDC pipeline, “running” is not the same as “current.”
Monitor how far the capture process is behind the MySQL source and how long changes take to become queryable in ClickHouse.
Keep enough binlog recovery headroom
Binlog retention should include enough margin for:
- database maintenance
- networking problems
- destination outages
- prolonged pipeline pauses
- backpressure
If the required binlog segment has already been removed, the safest recovery may require re-establishing the current source state through another backfill.
Test large backfills separately
The resource profile of an initial load can be very different from steady-state CDC.
Test large production-like tables before assuming that a pipeline that performs well during ongoing replication will behave the same way during the historical load.
Plan for MySQL failover
Failover behavior differs between CDC implementations.
ClickHouse’s newer MySQL ClickPipes use GTID-based replication for new pipes, helping make replication positions portable across MySQL hosts.
Estuary’s MySQL capture currently tracks server-specific binlog coordinates. If the writer changes, the previously stored position can become invalid and replication can fail with ERROR 1236.
For a planned failover where writes can be paused, Estuary documents a process for re-establishing capture against the new writer without necessarily requiring a full backfill.
This is an important architectural difference to consider when designing for database failover.
Verify replica behavior before capturing from a read replica
Estuary can capture from a MySQL read replica when binary logging is enabled and the normal CDC prerequisites are satisfied.
Do not assume every managed-service reader endpoint meets those conditions; verify binary logging and replication access on the actual endpoint you intend to capture from.
Using Estuary for MySQL to ClickHouse
Estuary provides a managed path that captures MySQL changes from the binary log and then materializes the resulting collections into ClickHouse.
The architecture is:
Estuary currently offers two ClickHouse delivery paths.
The direct ClickHouse materialization is recommended for most workloads. It writes directly to ClickHouse using its Native protocol and works with both ClickHouse Cloud and self-hosted ClickHouse.
A Dekaf + ClickPipes path is also available when an organization specifically requires ClickPipes or prefers ClickHouse Cloud to own the final ingestion infrastructure.
See the Estuary ClickHouse materialization documentation for the current recommended path.
Step 1: Prepare MySQL for CDC
Configure MySQL with:
- row-based binary logging
- adequate binlog retention
- the required replication permissions
- table read access
- an explicit timezone when necessary
Estuary supports self-hosted MySQL as well as common managed MySQL deployments including Amazon RDS, Aurora, Google Cloud SQL, and Azure Database for MySQL.
Step 2: Create the MySQL capture
In Estuary, configure the MySQL endpoint and credentials and select the tables that should become collections.
By default, the capture first backfills the existing table contents and then transitions into ongoing binlog capture. Backfills can be controlled at the table level when necessary.
Step 3: Create the ClickHouse materialization
Create a ClickHouse destination and select the captured collections that should be written to ClickHouse.
The direct connector uses the ClickHouse Native protocol:
- port 9440 with TLS, which is the default
- port 9000 without TLS
It does not use ClickHouse’s HTTP interface on port 8123.
This is worth checking carefully when using ClickHouse Cloud because copying an HTTP endpoint into a Native-protocol connector is a common configuration mistake.
Step 4: Map collections to ClickHouse tables
For each selected collection, configure the ClickHouse target table.
The connector can create the destination tables as part of the materialization configuration.
Step 5: Validate the current state
Once the historical backfill and CDC stream are running, validate both data arrival and current-state query behavior.
For tables using Estuary’s standard ClickHouse update mode, queries that require fully deduplicated logical rows should use:
plaintextSELECT *
FROM my_table FINAL;How Estuary Handles Updates and Deletes in ClickHouse
This behavior is important because it directly affects how downstream queries should be written.
Standard updates
In standard mode, Estuary creates ClickHouse tables using ReplacingMergeTree with flow_published_at as the version column.
When a record changes, it inserts a newer version.
ClickHouse removes superseded versions during background merges, but those merges are asynchronous. Until the merge has occurred, multiple physical versions of a logical record may exist.
Using FINAL applies the replacement logic at query time so the latest logical state is returned.
ClickHouse itself recommends SELECT ... FINAL when query-time correctness is required for ReplacingMergeTree, rather than relying on background merges to have already completed.
For more detail, see ClickHouse’s guidance on ReplacingMergeTree and FINAL.
FINAL is not free, however. Applying deduplication during query execution adds work, so query design and table partitioning still matter at scale.
Soft deletes are the default
By default, Estuary uses soft deletes for ClickHouse.
The row remains present, and deletion state is represented through the _meta/op field.
Hard deletes
If hardDelete is enabled, Estuary writes a new row with the original key and:
_is_deleted = 1
This acts as a tombstone.
ReplacingMergeTree excludes the tombstoned record from FINAL results, and ClickHouse eventually removes it from persistent storage through background cleanup.
Delta updates
Estuary also supports delta_updates on individual bindings.
When enabled, the table uses MergeTree rather than ReplacingMergeTree, and operations are appended as-is without destination-side deduplication.
This can be useful for append-oriented or event-style workloads where retaining the sequence of changes is preferable to maintaining one reduced current-state row.
When to Use Estuary’s Dekaf + ClickPipes Path
The direct ClickHouse connector should be the default Estuary path for most new MySQL-to-ClickHouse pipelines.
Dekaf remains useful when:
- organizational policy requires ClickPipes
- the team wants ClickHouse Cloud to own the final ingestion layer
- Kafka-compatible topic consumption is specifically useful to the architecture
In that design:
MySQL → Estuary Capture → Estuary Collection → Dekaf → ClickPipes → ClickHouse Cloud
Estuary exposes collections as Kafka-compatible messages that ClickPipes can consume.
See the Estuary Dekaf for ClickHouse documentation if that architecture is required.
What Is Different About the Estuary Approach?
The useful differences for this route are architectural rather than generic product claims.
One MySQL capture can support broader downstream data movement
The source is captured into Estuary collections before materialization.
That means ClickHouse does not have to be the only downstream consumer of the captured data.
This is useful when the same MySQL changes also need to serve another warehouse, lake, application, or operational workflow.
Direct support for ClickHouse Cloud and self-hosted ClickHouse
ClickPipes is a ClickHouse Cloud service.
Estuary’s direct ClickHouse materialization can target both ClickHouse Cloud and self-hosted ClickHouse.
CDC and batch are separate source choices
Estuary recommends the MySQL CDC connector when binary-log capture is available because it provides lower-latency change capture and includes update and delete events.
The batch connector remains available when CDC is unavailable or when custom queries and views need to be captured.
Delete behavior is explicit
The ClickHouse materialization exposes both soft-delete and hard-delete modes rather than treating deletion behavior as an implementation detail.
That makes the downstream semantics easier to evaluate before putting the pipeline into production.
When Another Approach May Make More Sense
Estuary is not required for every MySQL-to-ClickHouse architecture.
Use ClickPipes when
- ClickHouse Cloud is your destination.
- You want native MySQL CDC managed within ClickHouse.
- You do not need a separate integration layer to serve other destinations.
Use Debezium + Kafka when
- Kafka already forms the core of your streaming architecture.
- Your team wants direct control of the change-event stream.
- You are comfortable operating the extra infrastructure.
Use batch extraction when
- data only needs to refresh periodically
- source volume is manageable
- low latency is unnecessary
- operational simplicity matters more than continuous synchronization
Consider a managed integration platform such as Estuary when
- MySQL changes need to serve multiple downstream systems
- you need self-hosted ClickHouse support
- you want managed source capture and destination delivery
- your broader integration estate includes other databases, SaaS systems, warehouses, or operational destinations
The architecture should follow the workload—not the other way around.
Move MySQL Data to ClickHouse With Estuary
If you need continuous MySQL CDC into ClickHouse, Estuary can capture existing records, continue reading new changes from the MySQL binary log, and materialize the resulting collections directly into ClickHouse.
Start with the MySQL capture documentation and ClickHouse materialization documentation to review the current prerequisites and connector behavior before configuring the pipeline.
FAQs
Does ClickHouse support native MySQL CDC?
Can ClickHouse query MySQL without copying the data?
Do I need Kafka to move MySQL data to ClickHouse?
Is CDC always better than batch ETL?

About the author
Team Estuary is a group of engineers, product experts, and data strategists building the future of real-time and batch data integration. We write to share technical insights, industry trends, and practical guides.
















