High availability and automatic failover

Automatic takeover without split-brain apply.

Durable leases, fencing epochs, capability-aware scheduling, and transferable checkpoints let a healthy agent resume a failed pipeline.

Lease
Durable ownership

Record which agent currently owns the pipeline.

Fence
Epoch-guarded apply

Old owners cannot write after takeover.

Resume
Durable positions

Recover source and sink progress before running.

Control-plane HA

Schedule from durable shared state.

Ownership, desired state, agent capability, and failover history survive restarts. Work moves only through an expired lease or fenced switchover.

Capability-aware agent selection
Lease renewal and expiry telemetry
Planned switchover and automatic takeover
One merged status after reassignment

Data-plane safety

Resume after identity and timeline checks.

The new owner restores durable state, validates source identity, acquires the next epoch, and then resumes. Ambiguous timelines stop loudly.

Remote checkpoint and spool recovery
Source promotion continuity by engine
Sink-side stale-epoch rejection
Crash and partition recovery tests

Keep exploring

See the connected parts of the platform.

Keep one safe owner through every failure.

Tell us the agent fleet, database topology, and recovery objectives. We will map leases and failover.