CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
CIMERA is presented, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution.