Jul 2026· 2026 5th International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE)· pp. 1-6· 0 citations· 14 references
Abstract
Multiply-Accumulate (MAC) units are fundamental hardware blocks in Digital Signal Processing (DSP) systems, where dynamic power efficiency is a critical design constraint. Traditional high-speed MAC architectures frequently employ Square Root Carry Select Adders (SQRT CSLA) for the accumulation stage. However, regular SQRT CSLAs rely on redundant Ripple Carry Adders (RCAs) to compute parallel potential sums for both $\mathbf{C}_{\mathbf{i n}}=\mathbf{0}$ and $\mathbf{C}_{\mathbf{i n}}=\mathbf{1}$ conditions, leading to excessive dynamic switching activity. This paper proposes a highly power-efficient 16-bit MAC architecture utilizing a Carry Enable Binary to Excess-1 Converter (CEBEC) SQRT CSLA. The proposed design entirely eliminates the redundant $\mathbf{C}_{\text {in }} \boldsymbol{=} \mathbf{1}$ RCA blocks, replacing them with a streamlined combinational logic path. This path utilizes optimized NOT and XOR gates for lower-order bits, coupled with targeted OR-gate logic at the Most Significant Bit (MSB) for rapid carry evaluation. The baseline and proposed architectures were functionally verified via Cadence SimVision and synthesized to the gate level using the Cadence Genus Synthesis Solution. Post-synthesis power and area analysis demonstrates that the proposed CEBEC-based MAC unit achieves a significant 36.7% reduction in dynamic switching power compared to the baseline CSLA, dropping from 15.22 $\boldsymbol{\mu} \mathbf{W}$ to $\mathbf{9. 6 3} \boldsymbol{\mu} \mathbf{W}$. While this hardware optimization trades a marginal 5.3% increase in total standard cell area, the substantial mitigation of switching activity makes the proposed architecture highly viable for low-power DSP ASICs.
Multiply Accumulate (MAC) units play an important role in high-performance AI/ML and digital signal processing systems where both reliability and computational efficiency are critical for achieving optimal performance. In this work, an optimized MAC unit is designed and implemented by integrating a Wallace Tree Multiplier to enable high speed multiplication and a Kogge-Stone adder to achieve low latency accumulation. The proposed architecture which is implemented in 90 nm technology is analyzed and compared with existing MAC designs in terms of performance, power consumption and area. The results show that the proposed architecture has a delay of 3.99 ns, power consumption of 0.507404 mW and an area of $2967.805 \mu \mathrm{m}^{2}$. To further improve the reliability of the design, a Triple Modular Redundancy (TMR) mechanism is incorporated which provides protection against both transient and permanent hardware faults. The simulation result verifies the implemented TMR mechanism in the proposed MAC architecture by operating correctly even when faults are present. A balanced trade off is achieved between speed, power efficiency and fault tolerance using the proposed MAC architecture which makes it ideal for high performance AI/ML applications.
Abhiram K. M., B. B., R. Ratnakumar· 2026 International Conferenc...· 0 citations
Arithmetic operations are fundamental to digital signal processing systems, where multipliers often decide overall performance constraints. They are key components of many high-performance systems such as Microprocessors, FIR Filters, Digital Signal Processors etc. The most common way of performing signed multiplication in digital circuits is by using booth multipliers, but the existing algorithm has a drain on power consumption since it never coerces operations to the full precision. In this paper, a novel 32-bit pipelined multiplier is designed aimed at achieving high throughput and low power consumption for VLSI applications. Modified Booth Encoding (MBE) with Radix-8 and Wallace tree reduction for partial products reduction and along with a CLA adder for partial products addition is used. Furthermore, a linear Pipelining technique with flipflops is implemented to minimize critical path delay. The Register Transfer level (RTL) model was implemented using Verilog and synthesized using Xilinx Vivado. Performance analysis demonstrates the reduction of delay by 23% and power consumption by 61%.
C.S. Chakradhar, M. Sreedhar· ITEGAM- Journal of Engineeri...· 0 citations
High-speed and space-efficient arithmetic units are necessary for the hardware implementation of cryptographic
algorithms to guarantee security and performance. Efficient multi-operand addition, especially three-operand binary addition, is
crucial for modular operations like multiplication and exponentiation. A high-speed, low-area three-operand binary adder for
cryptography and pseudorandom bit generator (PRBG) applications is presented in this study. Kogge-Stone, Han-Carlson, and
Ladner-Fischer are examples of parallel prefix adders that are used to increase throughput and decrease propagation latency.
For VLSI systems, heterogeneous delay-insensitive coding is used to further optimize power, area, and performance. A Carry
Look-Ahead (CLA) adder, which lowers critical route delay through parallel carry generation, is added to the design to improve
carry calculation. When compared to traditional designs, the suggested hybrid architecture provides increased speed and
efficiency. Its usefulness for high-performance computing and digital signal processing applications is demonstrated via
implementation using Xilinx Vivado
S. Aparna, C. Padma· International Journal for Re...· 0 citations
Need of Digital Signal Processing (DSP) systems which is embedded and portable has been increasing as a result of the speed growth of semiconductor technology. Multiplier is a most crucial part in almost every DSP application. So, the low power, high speed multipliers is needed for high-speed DSP applications. Vedic multiplier is one of the fastest and efficient multipliers, Main algorithm of Vedic multiplication is Urdhva Triyakbhyam. It is a general multiplication formula applicable to all cases of multiplication. It literally means Vertically and Crosswise. The multiplication is depending upon the previous computations of partial sum to produce the final output, so we need to design efficient adders. In arithmetic operation, major issue corresponds to carry in binary number system. Higher radix number system like Quaternary Signed Digit (QSD) can be used for performing arithmetic operations without carry. Designing the adder using QSD number representation allows fast addition which is capable of carry free addition because the carry propagation chain is eliminated, hence it reduces the propagation time in comparison with radix 2 system. In this project, a Vedic multiplier using Quaternary Carry Look Ahead Adder (QCLA) is existed design and Vedic multiplier using Quaternary Carry Increment Adder (QCIA) is proposed design, it has less area compared with the Existed design. In this project Xilinx-ISE 14.7 tool is used for simulation, logical verification, and further synthesizing and the HDL language used for the project is VERILOG.
S. Anjum, M. Jyothi· International Journal of AI...· 0 citations
The Arithmetic Logic Unit (ALU) is a fundamental building block of all modern central processing units (CPUs). This papers presents the design and implementation of a 16-bit, 4-function ALU constructed entirely from basic logic gates, using a structural Verilog approach. The ALU supports four operations: addition, subtraction, bitwise AND, and bitwise OR, implemented via a modular, bit-sliced architecture. A single ripple-carry adder with 2's complement logic performs both arithmetic operations efficiently. Operation selection is achieved through a 2-bit control signal and a 4-to-1 multiplexer per bit slice. The system was verified using QuestaSim simulation with both normal and boundary test vectors. Unlike existing behavioral or high-level implementations, this work contributes a fully gate-level structural model that makes every interconnection explicit, enabling transparent analysis of carry propagation and delay. Performance metrics including propagation delay and hardware resource usage are analyzed. The results validate the design's functional correctness and demonstrate its potential as a scalable building block for processor architectures.
P. M. B., Yashaswini H. A.· International Journal of Eng...· 0 citations
Hermitian symmetry constrained optical OFDM in IM/DD transmitters incurs higher computational costs per transmitted bit than conventional complex-valued OFDM, as only a small subset of the IFFT subcarriers carry independent information. This work introduces a unified optimisation framework that combines juxtaposed IFFT architectures, arithmetic rearrangement for complex multiplication, and a novel deterministic radix-4 pruning strategy tailored for specific optical OFDM modulation formats. The proposed approach achieves higher numerical precision than comparable radix- 2 implementations, operates with up to a twofold reduction in clock cycles, and requires $\mathbf{2 5 \%-5 0 \%}$ fewer multiplications relative to the current state of art. The architecture is validated on an RFSoC $4 \times 2$ platform using a fully parallel, unrolled $N=64$ implementation. A maximum operating fabric clock of 153.6 MHz was achieved, corresponding to a throughput of 19.6608 GS/s and a latency of 19.5312 ns, while consuming only 7200 FPGA LUTs and 168 DSP slices. These characteristics make the proposed architecture well suited to low-power, low-latency, and low-complexity optical transceivers.
Michael Codd, Ciara McDonald, John Dooley· International Symposium on C...· 0 citations