FONT SIZE : AAA
Considering signal processing over the years has been the most significant application of embedded multipliers, it is just natural that they evolved into more complex blocks, called DSP blocks, like the one in Figure 4.3, which includes all resources required to implement a MAC unit, eliminating the need for using distributed logic.
Different architectures exist for DSP blocks, but most of them share three main stages, namely pre-adder, multiplier, and ALU. Depending on the device, the ALU can just consist of an adder/subtractor or include
FIGURE 4.3 DSP block from Xilinx 7 Series.
additional resources (like in the case of Figure 4.3) aimed at giving the DSP block increased computation power (Altera 2011; Xilinx 2014; Lattice 2016; Microsemi 2016).
As in the case of multipliers, registers are placed at both the input and the output of the circuit in Figure 4.3, where interstage registers can also be identified. In this way, pipeline structures achieving very high operating frequencies can be implemented. In some DSP blocks from different FPGA families, double-registered inputs (consisting of two registers connected in a chain) are available, whereas other blocks include additional pipeline registers oriented to the implementation of systolic FIR filters. In this case, registers are placed at the input of the multiplier and at the output of the adder (which would be the input and output, respectively, of each stage of an FIR filter), to reduce interconnection delays. These registers are optional, so they can be bypassed if operation at the maximum achievable frequency is not required.
The significant amount of MUXes available provides the structure with many configuration possibilities. Thanks to them, it is possible to define different data paths and, in turn, different operating modes. In the case of Figure 4.3, the DSP block supports several independent functions, such as addition/subtraction, multiplication, MAC, multiplication and addition/ subtraction, shifting, magnitude comparation, pattern detection, and counting. The selection inputs of the MUXes are usually accessible from distributed logic, allowing operating modes to be configured dynamically (normally in a synchronous way).
The pre-adder/subtractor may be used as an independent computing resource or to generate one of the input operands of the multiplier. This sec- ond alternative is useful for the implementation of some functionalities, for instance, the symmetric filter shown in Figure 4.4.
FIGURE 4.4 Eight-stage symmetric FIR filter.
The multiplier in Figure 4.3 operates with different input values (direct or registered inputs, data from the pre-adder or from the chain-connected adja- cent block) and either stores the result in an intermediate register or sends it directly to the ALU, the “Result” output, or the “chainout” output.
Some DSP blocks in Altera FPGAs include a small “coefficient memory”* con- nected to one of the inputs of the multiplier, aimed at the optimized implemen- tation of digital filters. Addressing is controlled from distributed logic, allowing filter reconfiguration to be dynamically performed during system operation. Thanks to this memory, there is no need for using embedded or distributed FPGA memory to store coefficients, therefore optimizing resource usage and reducing the time required to access coefficient values.
Regarding the ALU, it can perform arithmetic operations (addition, sub- traction, and accumulation), logic functions, and pattern detection. When the accumulator is not used in conjunction with the multiplier, it can operate as an up/down synchronous counter. In some DSP blocks, the ALU can be divided into smaller units connected in chain and operating in parallel. For instance, a 48-bit ALU might operate as two 24-bit units, four 12-bit units, and so on. This feature is useful for the implementation of SIMD algorithms, so it is usually referred to as SIMD mode, an example of which is shown in Figure 4.5.
The pattern detection circuitry checks whether or not there is coincidence between one specific input of the DSP block (C in Figure 4.3) and the output of the ALU. Depending on the configuration, it is possible to implement other functions, such as convergent rounding, masked bit-wise operations, termi- nal count detection or autoreset of the counter, and detection of overflow/ underflow conditions in the accumulator.
It is also possible to perform some combined multiplication–addition/ subtraction operations with input data, for example, (A · B) ± C or (A + D) · B ± C.
FIGURE 4.5 SIMD operating mode.
* “Internal Coefficient” in Arria 10 devices, which can store up to eight coefficients.
This allows, for instance, the result of the multiplication to be symmetrically rounded off to zero or to infinity.
As it may be expected, DSP blocks have dedicated lines for chain connec- tion between adjacent blocks. In this way, the number of bits of the operands in arithmetic operations can be extended, and complex arithmetic functions or processing algorithms requiring multiple stages operating in parallel (e.g., digital filters) can be implemented. Same as embedded multipliers, DSP blocks are usually placed adjacent to embedded memory blocks.
The amount of DSP blocks available in a given FPGA depends on the target application profile. In devices oriented to signal processing, there may be some thousands of them,* achieving performances in the order of hundreds of GMAC/s. These very high computing speeds allow time multiplexing methods to be applied, in order for multiple operations of lower frequency to be carried out in a single DSP block. This results in semiparallel struc- tures achieving very efficient trade-offs between resource usage and power consumption.
Manufacturer:Xilinx
Product Categories:
Lifecycle:Active Active
RoHS:
Manufacturer:Xilinx
Product Categories: FPGAs (Field Programmable Gate Array)
Lifecycle:Active Active
RoHS:
Manufacturer:Xilinx
Product Categories: Aluminum Electrolytic Capacitors
Lifecycle:Active Active
RoHS:
Manufacturer:Xilinx
Product Categories: FPGAs (Field Programmable Gate Array)
Lifecycle:Active Active
RoHS: No RoHS
Manufacturer:Xilinx
Product Categories: FPGAs (Field Programmable Gate Array)
Lifecycle:Active Active
RoHS: No RoHS
Support