This website uses cookies. By using this site, you consent to the use of cookies. For more information, please take a look at our Privacy Policy.
Home > FPGA Technical Tutorials > FPGA-Based Prototyping Methodology > Partitioning and reconnecting > Design synchronization across multiple FPGAs

TABLE OF CONTENTS

Xilinx FPGA FPGA Forum

Design synchronization across multiple FPGAs

FONT SIZE : AAA

Our SoC design started as a single chip and will end up the same way in silicon but  for now, it is spread over a number of chips and the contiguity of the design suffers  as a result. We have discussed some of these ways we can compensate for the on chip/off-chip boundary to improve performance and later we shall discuss  multiplexing of signals.  

There are three particular aspects of having our design spread over multiple chips  that we need to concentrate upon. These are the clocks, the reset and the start-up  conditions. For complete fidelity between the prototype and the final SoC, then the  clock, reset and start-up should behave as if the hard inter-FPGA boundaries did not  exist. Let’s consider each of these in turn.

Multi-FPGA clock synchronization

We saw in chapter 5 how the FPGA platform can be created with clock distribution  networks, delay-matched traces and PLLs in order to be as flexible as possible in  implementing SoC clocks. Now is the time to make use of those features. 

There are two potential problems that occur when synthesizing a clock network on a  multi-FPGA design:  

• Clock skew and uncertainty: a common design methodology uses the  board’s PLLs to generate the required clocks and distribute them as  primary inputs to each FPGA. The on-board traces and buffers can  introduce some uncertainty and clock skew between related clocks due to  the different paths taken to arrive at the FPGAs. If ignored, this skew can  cause hold-time violations on short paths between these related clocks.  

• The back-end place & route tools can typically resolve hold-time violations  within the FPGA but cannot currently import skew and uncertainty  information that is forward-annotated from synthesis. To address this  problem, the FPGAs own PLLs (part of the MCMMs in a Virtex®-6 FPGA  ) can be used for clock generation in combination with, or instead of the  board’s PLLs,. The back-end tools understand the skew and uncertainty of  an MCMM and can account for them during layout. However this use of  distributed MCMMs can introduce the second of our potential problems  mentioned above. 

• Clock synchronization: When related clocks are regenerated locally on  each FPGA there is a potential clock-synchronization problem that does  affect the original SoC design, where clocks are generated and distrusted  from a common source. The problem becomes particularly apparent when  multiple copies of a divide-by-N clock are derived from a base clock but are not in sync because of reset or initial conditions and just unlucky  environmental glitches.

Partitioned designs have their clock networks distributed across the FPGAs and  other clock components on the board. Spanning a large number of different clocks  across all FPGAs can be made easier by successful clock-gate conversion (see  chapter 7), reducing the number and complexity of the clocks. However, there will  still be a number of clock drivers which need to be replicated in multiple FPGAs. It  is important that these replicated clocks remain in sync and that any divided clocks  are generated on the correct edge of the master clock. This is dependent upon the  application of reset on the same clock edge at each FPGA as we shall describe in the  next section.

We therefore have a generic approach to clock distribution as shown in Figure 114, in which we see the source clock generated on a board while the local MMCMs are  used in each FPGA to resynchronize and then redistribute the clocks via the BUFG  global buffers. In many cases this small FPGA-clock tree will need to be inserted  into the design manually. The latest partitioning tools are introducing features which  automatically insert common clocking circuitry into each FPGA.

Clock distribution across FPGAs.png

Our earlier recommendation to design the SoC with all clock management in a toplevel block will really help now. We only need to make changes at that one toplevel block and then use replication to partition the same structure into each FPGA.  Even if the clocking in the SoC RTL is distributed throughout the design then  replication will help us restrict the changes to fewer RTL files than might otherwise  be necessary.  

Replication of clock buffers might be avoided if we use a technique discussed in  section 8.2.11.1 above for partitioning by clock domain. Success of this approach  will depend on relative fanout of the different clocks and the number of paths  between domains. 

Whatever the partitioning strategy, each FPGA is a discrete entity and for clock  synchronization, each must have its own clock generation rather than relying on  clocks arriving from a generator in another FPGA. Therefore a clock generator must  be instantiated in each FPGA even if only a small part of the SoC design is  partitioned there.  

Let’s look a little more closely at the clock generator and how it helps us during  prototyping. The MMCM’s PLL has a minimum frequency at which it can lock and  so the input clock to the FPGA must drive at least at that rate. Virtex®-6 MCMMs  have a minimum lock frequency of 10MHz, compared to 30MHz or higher for  previous technologies, which makes them particularly useful for our purposes. They  are able to generate much slower clocks than they can accept as inputs. Our task is  therefore to assemble a clock tree across the prototype where we maintain a higher  frequency outside of the FPGAs and then divide internally while keeping the  internal clocks in sync. 

We achieve this as follows:  

• If the SoC has clock generator at the top-level as recommended then 

• Create a clock generator block in RTL to replace the equivalent part of the  SoC generator. We use a global base clock to drive MCMMs which  generate slower derivatives via its divide-by-n outputs.  

• Select any one global system clock generated on board; its frequency must  be above the minimum MMCM lock frequency.  

• Drive the input to the new RTL clock tree with the global clock. 

• In the new clock generator RTL, create a tree of MMCM and BUFG  instantiations to create all necessary sub-clocks.  

• During partitioning, replicate the MMCM and BUFGs into each FPGA as  required to drive logic assigned there.  

• During trace assignment, the global clock must be assigned to the global  clock inputs for the FPGA.

• If clock gating and generation is more distributed then we may need to  instantiate BUFG and MMCM components directly into different RTL  files, but we should always be on the       lookout for how replication can make  this process easier.

For multi-board prototypes, we may need an extra level of hierarchy in the clock  tree. We should use the on-board PLLs to drive the master clock to each board and  resynchronize using PLLs at each board, using the board’s local PLL output to drive  each FPGA locally as described above. Skew between boards is avoided by using  PLLs and matched-delay clock traces and cables as described in section 5.3.1.  Because multiple PLLs and MCMMs may be used, the local slow clocks must be  synchronized to the global clock. This is achieved by using the base clock as the  feedback clock input at each MMCM. 

We might ask why we do not generate all clocks using the on-board global clock  resources. After all, we might have placed specialist PLL devices on the board with,  for example, even lower minimum lock frequencies. The issue to be aware of here is  that, depending upon the board, there may be some skew between the arrival times  of the clocks at each FPGA, especially for less sophisticated boards not specifically  designed and laid out for this purpose. This effect will be magnified in a larger  system and could lead to issues of hold-time violations on signals passing between  FPGAs. 

Multi-FPGA reset synchronization 

There will be “global” synchronous resets in the SoC which will fan out to very  many sequential elements in the design and assert or release each of them  synchronously, on the same clock edge. As those sequential elements are partitioned  across multiple FPGAs it is obviously critical to ensure that each still receives reset  in the same way, and at the same clock edge. This is not a trivial problem but  distributing a reset signal across several FPGAs in a high-speed design can be  achieved with some additional lines of code in the RTL design and a partitioning  tool that allows easy replication of logic.  

The enabling factor in this approach is that it is unlikely that the global reset needs  to be asserted or released at a specific clock edge, as long as it is the SAME clock  edge for every FPGA. Therefore, we can add as many pipeline stages into the reset  signal path as we need and we shall use that to our advantage in a moment.  

Another factor to consider is that the number of board traces which connect to every  FPGA is often limited so the good news is that we do not need any for routing  global resets. Instead we create a reset tree structure which routes through the  FPGAs themselves and then uses ordinary point-to-point traces between the FPGAs.

Considering the design example in Figure 115 in which a pipelined reset drives  sequential elements in four different FPGAs, it is clear that the elements in FPGA 4  will receive the reset long after those in FPGA 1.

Pipelined synchronous reset driving elements in 4 FPGAs.png

Notice also that the input could be an asynchronous reset, so the first stage acts as a  synchronizer (with double clocking if necessary for avoiding metastability issues).  Routing through the FPGAs in this way is not acceptable except for very low clock  rates, so to overcome this we will replicate part of the pipeline in each of the FPGAs  as shown in Figure 117.

Tree pipeline created by logic replication.png

Readers using Certify will find further instructions on the correct order for  replication in the apnote listed in this book’s references. 

The first stage is not replicated because the synchronization of the incoming reset  signal has to be done in only one place. There is still a possibility that the pipeline  stage in each FPGA would introduce delay because if there is one FF in each stage,  then it might be placed near the input pad, the output pad or anywhere in between;  this might also be different in each FPGA. The answer is to use three FFs for each  pipeline stage, as shown in Figure 116.

A three-FF pipeline stage gives more freedom to place & route.png

This allows the first and third FF to be placed in an IO FF at the FPGA’s edge. Then  there is a whole clock period for the reset to propagate to the internal FF and on to  the output FF, greatly relaxing the timing constraint on place & route. Once again,  these pipe stages only introduce an insignificant delay compared to the effect of the  global reset signal itself.

Note: the synthesis might try to map the pipeline FFs into a single shift register  LUT (SRL) feature in the FPGA, which would defeat the object of the exercise so  the relevant synthesis directive may be required to control the FF mapping. In the  case of Synopsys FPGA synthesis, this would be syn_srlstyle and for good measure  syn_useioff would be used to force the two FFs into the IO FF blocks, although that  is the default.

Multi FPGA start-up synchronization

Using the above technique we can ensure that all FPGAs emerge from start-up  simultaneously. But this is of little use if the clocks within a given FPGA are not  running correctly at that time. We must also ensure that all FPGA primary clocks  are running before reset is released. This is important because, owing to analog  effects, not all clock generation modules (MMCMs, PLLs) may be locked and ready  at the same time. We must therefore build in a reset condition tree. This can be  accomplished by adding a small circuit like the one shown in Figure 118 which  would be distributed across the FPGAs.  

Here we see a NAND function in each FPGA to gate the LOCKED signals from  only those MCMMs which are active in our design. This would be a combinatorial  function or otherwise registered only by a free-running clock that is independent of  the MMCM’s outputs.

Example reset condition tree.png

Each FPGA then feeds its combined all_ LOCKED signal to the master FPGA in  which it is ORed to drive the reset distribution tree described in the previous  section. The locked signals of any on-board PLLs used in this prototype must also  be gated into the master reset and there may other conditions not related to clocks  which also have to be true before reset can be released, for example, a signal that  external instrumentation is ready and, of course, user’s “push-button” reset should  also be included. Our example shows that these are all active high but of course the  reset gate can handle any combination. The whole tree would be written in RTL  which is added into the FPGA’s version of the chip support block at the top level  and we can use replication to simplify its partitioning. 

The global reset is only released when all system-wide ready conditions are  satisfied. The reset will also then release all clock dividers in the various FPGAs on  the same clock edge so that all divide-by-n clocks will be in sync across all FPGAs.




  • XC3S50A-4VQ100I

    Manufacturer:Xilinx

  • FPGA Spartan-3A Family 50K Gates 1584 Cells 667MHz 90nm Technology 1.2V 100-Pin VTQFP
  • Product Categories: FPGAs

    Lifecycle:Active Active

    RoHS: No RoHS

  • XC3S50A-5VQ100C

    Manufacturer:Xilinx

  • FPGA Spartan-3A Family 50K Gates 1584 Cells 770MHz 90nm Technology 1.2V 100-Pin VTQFP
  • Product Categories: FPGAs (Field Programmable Gate Array)

    Lifecycle:Active Active

    RoHS: No RoHS

  • XCS20-5PQ208I

    Manufacturer:Xilinx

  • Spartan and Spartan-XL Families Field Programmable Gate Arrays
  • Product Categories:

    Lifecycle:Obsolete -

    RoHS: -

  • XC5215-6HQ240C

    Manufacturer:Xilinx

  • FPGA XC5200 Family 23K Gates 1936 Cells 83MHz 0.5um Technology 5V 240-Pin HSPQFP EP
  • Product Categories:

    Lifecycle:Obsolete -

    RoHS: No RoHS

  • XC2V1000-4BGG575C

    Manufacturer:Xilinx

  • FPGA Virtex-II Family 1M Gates 11520 Cells 650MHz 0.15um Technology 1.5V 575-Pin BGA
  • Product Categories: FPGAs

    Lifecycle:Obsolete -

    RoHS:

Need Help?

Support

If you have any questions about the product and related issues, Please contact us.