
Disaggregated KV-Cache Fabric
A hardware-agnostic memory virtualization architecture for LLM agent swarms. Scale autonomous workflows across high-speed GPU VRAM and dense host DRAM with concurrent lock-coupled radix prefix caching and race-free PCIe DMA eviction.
A high-performance dual-runtime system stripping away VRAM memory bottlenecks through disaggregated host DRAM offloading and lock-coupled prefix caching.
The KV-Cache VRAM Bottleneck Formula
In multi-agent swarms (CrewAI, LangGraph, AutoGen), context ballooning quickly exhausts 80GB VRAM. Fabric continuously evaluates allocation pressure through the canonical formula:

“Embrace the memory fabric and unlock boundless agent context.”Somya Prasad Sethy — Architecture Note
Architecture & Benchmark Explainer
// Hand-over-hand reader-writer mutex traversal
func (r *RadixTree) MatchPrefix(tokens []int64) (*MatchResult, error) {
curr := r.root
curr.mu.RLock()
for _, tok := range tokens {
next, exists := curr.children[tok]
if !exists {
// Terminal node reached: release and return match
curr.mu.RUnlock()
break
}
// Lock child before releasing parent (coupling)
next.mu.RLock()
curr.mu.RUnlock()
curr = next
}
return &MatchResult{Hit: curr.BlockHandle, SuffixDelta: 0}, nil
}Concurrent Radix Tree with Lock Coupling
Autonomous multi-agent pipelines reuse long shared prompts (system prompts, role declarations, JSON schemas). Fabric maintains an in-memory prefix tree where agents traverse common prefixes concurrently without global lock contention.
Architectural Building Blocks
Go Conductor
Maintains hierarchical token prefix radix trees with hand-over-hand reader-writer lock coupling and handles agent swarm ingress.
Rust Memory Daemon
Native crate controlling low-level PCIe DMA migrations between GPU VRAM physical blocks and NUMA host DRAM ring buffers.
Race-Free Eviction
Atomic double-check verification guarantees zero stale dereferences when evicting blocks under 90% high-watermark pressure.
Hardware Agnostic
Runs seamlessly across NVIDIA CUDA 12.4+, AMD ROCm MI300X, Apple Silicon Metal unified memory, and CPU AVX-512.
Spin Up the Disaggregated Engine

Ready to optimize your LLM memory fabric?
Clone the repository, inspect the Go Conductor and Rust memory daemon, or integrate the zero-copy DMA eviction pool.