This website uses cookies. By using this site, you consent to the use of cookies. For more information, please take a look at our Privacy Policy.
Home > FPGA Technical Tutorials > FPGA-Based Prototyping Methodology > Bring up and debug: the prototype in the lab > Quick turn-around after fixing bugs

TABLE OF CONTENTS

Xilinx FPGA FPGA Forum

Quick turn-around after fixing bugs

FONT SIZE : AAA

Debugging can become an iterative process and given that the time taken to process  a large SoC design through synthesis and place & route might be many hours, if we  find ourselves wasting a lot of time just waiting for the latest build in order to test a  fix, then we are probably doing something wrong. 

Here are three approaches we recommended in order to avoid this kind of painful “debug-by-iteration.”  

• Stick to a debug plan 

• Use incremental tool flows 

• Both of the above

A debug plan for a prototype is much like a verification plan for the SoC design as a  whole. As a full rerun of the prototype tool chains for large designs might take more  than a whole day, we need to have clear objectives for the build, which we should  aim to meet before re-running the next build. A debug plan sets out the parts of the  design which will be exercised for each build and also a schedule of forthcoming  builds. This is particularly useful when multiple copies of the prototype are created  and used in parallel. Revision control for each build and documentation of included  bug fixes and other changes are also critical. This is all good engineering advice, of  course, but a little discipline when chasing bugs, especially when time is short, can  sometimes be a rare commodity. With a good debug plan in place and a firm  understanding of the aims and content of each build, a regime of daily builds can be  very productive. 

A day’s turn-around is an extreme example and in some cases it would be much less  than that. For example, a bug fix that requires a change to only one FPGA will not  require the entire flow to be rerun. The partitioner may work on the design top down, but for small changes, the previous partition scripts and commands will still  be valid. As mentioned in chapter 4, a pre-synthesis partitioning approach helps to decrease runtime in this case because synthesis and place & route take place only on  the FPGA that holds the bug fix. 

To achieve a faster turn-around on small changes we can use incremental flows and  we shall take a closer look at these now.

Incremental synthesis flow

As mentioned in chapter 3, both synthesis and place & route tools have the ability to  reuse previous results and only process new or changed parts of the design.  Considering synthesis first, the incremental flow in most tools was created to reduce  the overall runtime taken to re-synthesize a design when only a small portion of the  design is modified. The synthesis will compare parts of the design to the same parts  during the previous run. If they are identical then the tools will simply read in the  previous results rather than recreate them. The tool then maps any remaining parts of the design into the FPGA elements instead of the entire design, combining the  results with those saved from the previous runs of the unchanged parts in order to  complete the whole FPGA. This incremental approach can save a great deal of time  and takes relatively small effort to set up.  

In Synopsys FPGA tools, incremental synthesis is supported by the compile point synthesis flow. In the compile point flow, we manually divide the design into a  number of smaller sub-designs or compile points (CPs) that can be processed  separately. This does not require any RTL changes but is controlled by small changes in the project and constraint files only, and can even be driven from a GUI.  

A CP covers the subtree of the design from that point down although CPs can also  be nested as we can see in Figure 143.

possible arrangement for Compile Points during FPGA Synthesis.png

The design can have any number of compile points, and we do not have to place  everything into CPs because the tool makes the top-level a CP by default. Therefore  the simplest approach might be to define CPs on all the modules which are frozen or  are not expected to change, allowing the tools to simply reuse the previous results  for that CP.  

Since each CP will be synthesized separately, we must define a separate constraint  file for each one which will include many of the same clock and boundary  constraints that we would need if the CP block was being assembled in a bottom-up  flow, or as a standalone FPGA. Details are of constraints available in the references  but the aim is that for a small up-front effort in time budgeting or by limiting our CP boundaries to registered signals, we can dramatically reduce our turn-around time. 

The tool needs to be able to spot when the contents of a CP have changed it  achieves this by noticing when the CP’s RTL source code logic or its constraints  have changed. Changes to some implementation options (such as retiming) will also  trigger all CPs to be re-synthesized.  

There is no such thing as a free lunch, however, and in the case of incremental flows the downside is that there may be a negative impact on overall device performance and resource usage. This is because cross-boundary optimizations may be arrested  at CP boundaries and so some long paths may not be as efficiently mapped as they  would have been had the whole design been synthesized top-down. In Synopsys  FPGA synthesis this can be mitigated to some degree and the amount of boundary  optimizations across the CP boundary can be controlled by setting the CP’s type as  either “soft,” “hard” or “locked.”

Another reason that incremental flows might not yield the highest performance  results is that the individual timing constraints for each CP may not be accurate as  they rely on estimated timing budgeting or user intervention. Conversely, this might  be seen as an advantage since we might purposely choose to tighten or relax the  constraints for each CP in order to focus the tool’s effort on more difficult-to achieve results for certain parts of the design, while relaxing others to save area or  runtime. 

Readers might be struck by the similarity between CP incremental synthesis and  traditional bottom-up synthesis of a design block-by-block. However, the scripting,  partitioning and resource management involved in a traditional bottom-up flow is  sometimes seen as too much of an investment for a prototyping team. There is also the problem of tracking exactly which files are to be re-synthesized in the new  build. In a design of thousands of RTL files, automation is very desirable.

Automation and parallel synthesis

If the workstation upon which the FPGA synthesis is running has multiple  processors then the synthesis task can be spread across the available processors, as  shown in Figure 144. 

With such multiprocessing enabled, multiple CPs can be synthesized simultaneously  further reducing the overall runtime of synthesis even for the first run or completere-builds when all CPs are re-synthesized. By using the automatic CPs and  multiprocessing options together, we can reduce the overall runtime of the synthesis  during the first time as well as the subsequent incremental iterations.  

A typical SoC design would need a number of CPs to be defined with individual  constraint files in order to achieve a significant reduction in the synthesis runtime.  Again, if the up-front investment for an incremental flow is too great or requires too  much maintenance (e.g., in scripts) then the flow becomes less attractive. Indeed,  defining a large number of CPs and creating constraints files for each may even be  seen as too time consuming. To overcome this apparent hurdle, Synopsys FPGA  synthesis is able to create and use CPs automatically.  

When automatic CPs are used, the tool can analyze a design and identify modules  that can be defined as CPs. The CP’s timing constraints can also be automatically  budgeted from the top-level constraints. This eliminates the need for us to define a  separate constraint file for each defined CP.

This is a relatively new technology and is a useful weapon for trading off runtime  against overall performance.

Incremental place & route flow

Incremental flow in place & route can be achieved using design preservation flow in  the Xilinx® back-end tool, the flow for which is seen in Figure 145.  

Design preservation flow in Xilinx® ISE® tools.png

Design preservation is one of the hierarchical design flows supported in the Xilinx® back-end tool. In the design preservation flow, the designs are broken into blocks  referred to as partitions. Partitions are the building blocks of all the hierarchical  design flows supported in the Xilinx® back-end tool. Partitions create boundaries around the hierarchical module instances so that they are isolated from other parts  of the design. 

A partition can either be implemented (mapped, placed and routed) or its previous  preserved implementation can be retained, depending on the current state of the  Figure 145: Design preservation flow in Xilinx® ISE® tools FPGA-Based Prototyping Methodology Manual 351  partition. If a partition’s state is set as “implement,” then the partition will be  implemented by force. And if the partition’s state is set as “import,” then the  implementation of that partition from the previous preserved implementation will be  retained.

When a design with multiple partitions is implemented in the Xilinx® back-end tool  for the first time, the state of all the partitions must be set to “implement.” For the  subsequent iterations, the state of the partition should be set either to “import” if the  partition has not changed or “implement” if it has changed. This way, only the  changed modules are re-implemented and the implementations of all the partitions  that have not changed are retained. This saves significant time in the  implementation of the entire design.

Combined incremental synthesis and P&R flow

Combined synthesis and place & route incremental flow.png

By combining incremental flows for synthesis and for place & route we can gain  our maximum runtime reduction. In fact, since place & route runtime is typically  Figure 146: Combined synthesis and place & route incremental flow 352 Chapter 12: Breaking out of the lab: the prototype in the field much longer than synthesis runtime, setting up incremental synthesis without  incremental place & route would not really improve turn-around time that much.  

With this in mind, Synopsys FPGA synthesis has been created to interleave its  compile point flow automatically with the Xilinx® design preservation flow, as  shown in Figure 146. In this integrated flow, the type of CP in the synthesis flow  should be set to “locked, partition,” so that it will be marked as a partition during  the place & route stage. When the integrated flow is enabled, the Synplify® Premier  software automatically writes the states of the partitions as “implement” or ‘import”  as appropriate.  

By using this integrated flow, as summarized in Figure 146, the overall turnaround  time to synthesis, place & route and design can be significantly reduced.

  • XC4028XL-09HQ240C

    Manufacturer:Xilinx

  • FPGA XC4000X Family 28K Gates 2432 Cells 0.35um Technology 3.3V 240-Pin HSPQFP EP
  • Product Categories: Contrôleur logique

    Lifecycle:Obsolete -

    RoHS: No RoHS

  • XC2C64-7CP56C

    Manufacturer:Xilinx

  • This lends power savings to High-end Communication equipment and speed to battery operated devices.
  • Product Categories: Programmable logic array

    Lifecycle:Any -

    RoHS: -

  • XC2C64-7VQ100I

    Manufacturer:Xilinx

  • This lends power savings to High-end Communication equipment and speed to battery operated devices.
  • Product Categories: Programmable logic array

    Lifecycle:Any -

    RoHS: -

  • XC2V1500-4BG575C

    Manufacturer:Xilinx

  • FPGA Virtex-II Family 1.5M Gates 17280 Cells 650MHz 0.15um Technology 1.5V 575-Pin BGA
  • Product Categories:

    Lifecycle:Obsolete -

    RoHS: No RoHS

  • XC2V1500-4FFG896I

    Manufacturer:Xilinx

  • FPGA Virtex-II Family 1.5M Gates 17280 Cells 650MHz 0.15um Technology 1.5V 896-Pin FCBGA
  • Product Categories: FPGAs (Field Programmable Gate Array)

    Lifecycle:Obsolete -

    RoHS:

Need Help?

Support

If you have any questions about the product and related issues, Please contact us.