Abstract
This paper presents a new design approach towards an arithmetic macrocell for computing the scalar product of two vectors. The design uniquely addresses competitive optimization goals of VLSI area, power dissipation and latency in the deep submicron regime. In comparison with the conventional architecture, the proposed floorplanning of a 16 bit scalar product macrocell on input vectors of 16 elements can achieve a saving of 38.6% of silicon area, up to 73% increase in area usage efficiency and 29.4% saving from the interconnect delay. Prelayout simulations based on a 0.35 /spl mu/m CMOS process indicate an average power dissipation of 315.71 mW and a latency of 11.39 ns, a sufficient performance for interfacing with many high speed 16 bit digital signal processors with no more than three clock cycles.