High-efficiency Compressor Trees for Latest AMD FPGAs

Konstantin J Hoßfeld,Michaela Blott,Hans Jakob Damsgaard,Jari Nurmi,Thomas B Preußer

doi:10.1145/3645097

Konstantin J Hoßfeld, Michaela Blott + Show 3 more

Open Access

https://doi.org/10.1145/3645097

Copy DOI

Abstract

High-fan-in dot product computations are ubiquitous in highly relevant application domains, such as signal processing and machine learning. Particularly, the diverse set of data formats used in machine learning poses a challenge for flexible efficient design solutions. Ideally, a dot product summation is composed from a carry-free compressor tree followed by a terminal carry-propagate addition. On FPGA, these compressor trees are constructed from generalized parallel counters whose architecture is closely tied to the underlying reconfigurable fabric. This work reviews known counter designs and proposes new ones in the context of the new AMD Versal™ fabric. On this basis, we develop a compressor generator featuring variable-sized counters, novel counter composition heuristics, explicit clustering strategies, and case-specific optimizations like logic gate absorption. In comparison to the Vivado™ default implementation, the combination of such a compressor with a novel, highly efficient quaternary adder reduces the LUT footprint across different bit matrix input shapes by 45% for a plain summation and by 46% for a terminal accumulation at a slight cost in critical path delay still allowing an operation well above 500 MHz. We demonstrate the aptness of our solution at examples of low-precision integer dot product accumulation units.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

High-efficiency Compressor Trees for Latest AMD FPGAs

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Reconfigurable Technology and Systems

Lead the way for us

Journal: ACM Transactions on Reconfigurable Technology and Systems	Publication Date: Apr 30, 2024
License type: other-oa

Similar Papers

Bolt
Davis W Blalock ... John V Guttag
-
Davis W Blalock, et. al.Davis W Blalock ... John V Guttag
13 Aug 2017
13 Aug 2017

Orthogonal Local Image Descriptors with Convolutional Autoencoders
Edgar Roman-Rangel ... Stephane Marchand-Maillet
-
Edgar Roman-Rangel, et. al.Edgar Roman-Rangel ... Stephane Marchand-Maillet
01 Jan 2020
01 Jan 2020

Very fast and exact accumulation of products
Ulrich Kulisch
Computing | VOL. 91
Ulrich KulischUlrich Kulisch
05 Dec 2010
Computing | VOL. 91

Approximate Adder Tree Synthesis for FPGAs
Sina Boroumand ... Philip Brisk
-
Sina Boroumand, et. al.Sina Boroumand ... Philip Brisk
01 Dec 2019
01 Dec 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

High-efficiency Compressor Trees for Latest AMD FPGAs

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Reconfigurable Technology and Systems