Commit cbd6fc2
feat: CDC-based mixed index synchronization (#4873)
Keep mixed indexes (ElasticSearch/Solr/Lucene) eventually consistent with the graph by deriving
their updates from a Change-Data-Capture stream of the committed graph data, instead of a
synchronous second write during the transaction that can diverge on failure and leave a
permanently stale index.
Pipeline (Apache Cassandra): commit (graph data only) -> Cassandra edgestore(cdc=true) ->
Debezium -> Kafka -> CdcIndexUpdateWorker consumer group -> reindex-from-current-state -> ES bulk.
Key design points:
- Reindex-from-current-state: the worker reads each changed element's current graph state and
fully replaces its index document (reusing IndexSerializer, like transaction recovery). This is
idempotent and order-independent, so out-of-order or duplicate events still converge to the
current state -- no strict ordering required, and a stale event can never overwrite a fresh
value. Worker transactions use skipDBCacheRead(): the worker JVM's database-level cache is
never invalidated by remote writers, so reads must hit the live graph.
- Additions are never dual-written: the only synchronous write is to storage; documents are
created and refreshed downstream from the committed change stream, so they cannot diverge.
- Deletions are routed by event identifiability, decided per deleted relation at commit:
* A removed vertex document is keyed by the vertex id -- the partition key every event carries,
including a whole-row partition delete -- so it is always removed by the worker.
* A removed MULTI-multiplicity edge with a surviving endpoint leaves an ordinary column
tombstone on that endpoint's row (its mirror copy) whose column carries the full edge
identity: the worker removes the document. In particular, removing a super node with
storage.drop-whole-row-on-vertex-removal while its neighbors survive costs one partition
delete and ZERO synchronous index operations -- the per-edge work happens asynchronously in
the CDC pipeline, where eventual consistency is the contract.
* Deletions no event can identify are written synchronously by the deleting transaction, the
only place their identities exist: constrained-multiplicity edges and meta-properties keep
their relation id in the storage value region (absent from tombstones), and unidirected
edges on removed vertices, self-loops, and edges whose both endpoints die in one transaction
have no surviving mirror. These carry dual-mode-grade durability; the docs recommend
tx.log-tx (whose PRECOMMIT entry durably records the deleted identities) plus transaction
recovery to close the crash window, and note that a graph-scanning REINDEX cannot remove
documents of already-deleted elements.
* Updates (a deleted relation whose id the same transaction re-adds) never delete
synchronously: the worker rewrites the still-live document id from current state, and a
delayed synchronous whole-document delete could otherwise erase that rewrite with no later
event to restore it.
* The reverse race -- the worker reads a relation as live, a concurrent transaction deletes it
and issues its synchronous document removal, then the worker's write lands last and would
resurrect the document permanently -- is closed by post-write verification: the applier
re-reads every relation document it wrote on a fresh snapshot and removes those whose
element vanished, which decides every interleaving correctly.
- Element-keyed Kafka partitioning gives per-element ordering and horizontal scaling via a
consumer group; batches are de-duplicated and applied in one transaction spanning all backing
indexes (one ElasticSearch _bulk per index), with VERTEX changes batch-preloaded via
getVertices(...) + multiQuery().properties().
- At-least-once: offsets are committed only after a batch is durably applied; on failure the
batch is reprocessed (rewind) rather than skipped, so the index eventually catches up. The
standalone runner supervises worker liveness and exits when every worker died unexpectedly, so
a process supervisor can restart it instead of leaving a healthy-looking zombie.
Configuration (opt-in, disabled by default):
- storage.cql.cdc: emit the Cassandra cdc=true table option on the edgestore table. Opening a
store logs a warning when the option is enabled but the live table lacks cdc=true (the option
only applies at table creation; an existing table must be ALTERed) -- otherwise that
misconfiguration is a silent no-capture drift.
- index.[X].cdc.enabled / index.[X].cdc.synchronous: per-index dual mode (write synchronously AND
via CDC) or cdc-only mode (skip synchronous additions; deletions no event can identify remain
synchronous). Both are GLOBAL_OFFLINE: the commit-side filter and the worker's index discovery
are one cluster-wide contract that per-instance values could split. cdc.synchronous=false
without cdc.enabled=true logs a warning instead of being silently inert.
Components:
- janusgraph-core: per-index CDC config options, the commit-side filters in StandardJanusGraph
(cdc-only mixed-index additions are filtered out at generation time; deletions choose their
filter per relation via the event-identifiability rule above; composite indexes are never
filtered; has2iMods/WAL/lock semantics are unchanged -- with a startup warning when cdc-only
mode is configured), and MixedIndexUpdateApplier (the backend-agnostic
reindex-from-current-state engine, covering vertex, edge and property-element mixed indexes,
and issuing removals for constraint-mismatched live vertices on graphs with user-settable ids,
where id reuse could leave a stale document from a previous incarnation). The restore paths
(ElementCategory.retrieve, IndexSerializer.removeElement) now accept custom String vertex ids
in addition to Long, RelationIdentifierUtils.findRelation no longer NPEs when a relation's
adjacent vertex has been removed, and findEdgeRelations returns an empty Iterable instead of
null. CdcElementChange documents the id contract for alternative capture sources (canonical
vertex ids; raw partition-representative RelationIdentifier endpoints).
- janusgraph-cql: the storage.cql.cdc table option and the live-table drift warning (no Kafka
dependency in production code).
- janusgraph-cdc (new module; core + kafka-clients 3.9.1, which carries the fixes for
CVE-2025-27817/27818): the CdcEventDecoder SPI, DebeziumCassandraJsonDecoder (parses the
relation header directly, so IN-direction edge columns and value-less delete tombstones of
MULTI edges resolve to the correct edge identity; canonicalizes partitioned-vertex ids for
VERTEX changes; skips corrupt records -- invalid JSON/Base64, malformed keys, and string-id
keys on graphs whose id regime forbids them -- while rethrowing transient backend failures so
the batch is redelivered), the CdcWorkerConfiguration (fail-fast validation incl. rejecting a
shared group.instance.id across multiple worker threads, which would fence itself forever;
cdc.consumer.* passthrough; auto.offset.reset defaults to "earliest" and the docs explain why
"latest" risks skipping events on a rebalance before a partition's first commit; pass-through
settings colliding with managed consumer keys are warned about instead of silently ignored),
the CdcIndexUpdateWorker (two-phase shutdown, interrupt-aware retries that stop when a
shutdown arrives mid-backoff, jittered exponential backoff so multiple workers do not retry
in lockstep, per-partition batch rewind on failure that tolerates partitions revoked
mid-rewind, a paced error loop, no consumer leaks), and the standalone
CdcIndexUpdateWorkerMain runner with worker-liveness supervision.
Testing: 80 tests across the touched modules, including unit/component coverage (decoder vs real
serialized bytes incl. poison-pill skips for invalid-Base64/malformed-key/string-id-keys,
delete-envelope after=null/before fallback, payload-wrapped envelopes, IN-direction columns and
value-less edge-delete tombstones, with assertions comparing full RelationIdentifier identity
incl. endpoint ids and the real meta-property id; reindex engine over vertex/edge/property-element
indexes incl. document removal when an element loses all indexed fields, removed-endpoint edges,
custom String vertex ids, stale-document removal after id reuse under a different label, the
partitioned-vertex id contract driven through the applier to real documents, multi-backing apply
and unknown-id-in-batch; worker loop via Kafka MockConsumer incl. multi-partition rewind,
no-op-batch offset commits, run()-loop recovery after a failed batch and Error-death
observability; commit-side deletion routing: synchronous removal for both-endpoints-removed,
constrained-multiplicity and meta-property deletions, no synchronous removal for mirror-identified
deletions and same-id replacements; composite indexes under cdc-only mode; full-chain convergence
over Lucene incl. vertex/edge add/update/remove and out-of-order delivery) and two real-container
E2Es -- worker -> Kafka -> ElasticSearch (runs on Java 8, 11 and 17), and the full Cassandra-CDC
-> Debezium -> Kafka -> ElasticSearch pipeline covering the vertex AND edge lifecycle
(add/update/property-removal/delete) against real Debezium delete envelopes, including a
whole-row both-endpoints vertex drop whose edge document converges with no per-edge event. The
Debezium pipeline test is gated behind the cassandra-cdc-e2e Maven profile (auto-activated on
JDK 17-23; Debezium 3.x requires 17+ and cassandra-all 4.1.7 does not run on 24+); the default
Java 8/11 build excludes only the two Debezium-dependent sources and stays green.
CI: a dedicated workflow (.github/workflows/ci-cdc.yml) runs the cdc unit tests plus the real
Kafka+ElasticSearch worker test on Java 8 and 11 (Testcontainers' optional jna dependency is
re-added at test scope, matching janusgraph-cql/janusgraph-es), and the full real-container
suite -- including the Cassandra-CDC -> Debezium -> Kafka -> ElasticSearch pipeline -- on Java 17
with Docker, so the integration is exercised on every change and guards against regression.
Docs: advanced-topics/cdc-mixed-index.md operator guide (incl. the deletion-routing design and
its tx.log-tx recommendation, cdc_raw backpressure and the CDC disable/teardown path, systematic
RF>1 event duplication, tuning guidance for the retry budget vs Kafka's max.poll.interval.ms
during long index outages, the auto.offset.reset caveat, and document-TTL limits), a 1.2.0
changelog upgrade note, and the regenerated configuration reference.
Fixes #4873
Replaces #4874
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rNFck9BY9s58qW3XQ1YTK
Signed-off-by: Oleksandr Porunov <alexandr.porunov@gmail.com>1 parent ac0eb23 commit cbd6fc2
39 files changed
Lines changed: 5781 additions & 9 deletions
File tree
- .github/workflows
- docs
- advanced-topics
- configs
- janusgraph-cdc
- src
- main/java/org/janusgraph/cdc
- test
- java
- io/debezium/connector/cassandra
- org/janusgraph/cdc
- resources
- janusgraph-core/src
- main/java/org/janusgraph/graphdb
- configuration
- database
- index
- internal
- relations
- test/java/org/janusgraph/graphdb
- configuration
- database/index
- janusgraph-cql/src
- main/java/org/janusgraph/diskstorage/cql
- test/java/org/janusgraph/diskstorage/cql
- janusgraph-lucene
- src/test/java/org/janusgraph/diskstorage/lucene
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
22 | 22 | | |
23 | 23 | | |
24 | 24 | | |
| 25 | + | |
25 | 26 | | |
26 | 27 | | |
27 | 28 | | |
28 | 29 | | |
29 | 30 | | |
30 | 31 | | |
31 | 32 | | |
| 33 | + | |
32 | 34 | | |
33 | 35 | | |
34 | 36 | | |
| |||
46 | 48 | | |
47 | 49 | | |
48 | 50 | | |
| 51 | + | |
49 | 52 | | |
50 | 53 | | |
51 | 54 | | |
| |||
0 commit comments