D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation
D$^2$-VLA is presented, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA, which uses block-wise causal KV caching to encode observations incrementally and constructs separate historical KV read views for the VLM and action expert.