This website uses cookies. By using this site, you consent to the use of cookies. For more information, please take a look at our Privacy Policy.
Home > FPGA Technical Tutorials > FPGA-Based Prototyping Methodology > Bring up and debug: the prototype in the lab > Ready to go on board?

TABLE OF CONTENTS

Xilinx FPGA FPGA Forum

Ready to go on board?

FONT SIZE : AAA

At this stage, some teams will decide to introduce the SoC design onto the FPGA  board. It may have seemed like a long time to finally reach this stage but an  experienced prototyping team will perform these test steps discussed in section 10.1  in a few hours or days at most. As mentioned, a piecemeal approach to bring-up  may save us a great deal of false debug effort.  

There is one further step which is recommended for first-time prototypers or any  team using a new implementation flow and that is to inspect the output of the  implementation flow back in the original verification environment. Our SoC design  has probably undergone a number of changes in order to make it FPGA-ready, not  least, partitioning into multiple devices. There may be other more subtle changes  that may have crept in during the kind of tasks described in chapter 7 of this book. How can we check that the design is still functionally the same as the original SoC,  Figure 137 : Expected FIR behavior across four FPGAs FPGA-Based Prototyping Methodology Manual 321  or at least only includes intentional changes? The answer is to reuse the verification  environment already in use for the SoC design itself.

Reuse the SoC verification environment

As mentioned in chapter 4, the RTL should have been verified to an acceptable  degree using the simulation methods and signed off as ready for prototyping by the  RTL verification team. Their “Hello World” test will probably have already been  run upon the RTL and it is very helpful if this same test harness can be reused on  the FPGA-ready version of the design.  

We have seen that by reusing the existing SoC test bench we can identify  differences between the behavior of the design before and after the design is made  ready for the prototype. It is important that any functional differences between the  SoC RTL and the FPGA-ready version can be accounted for as either intentional or  expected. Unexpected differences should be investigated as they may have been  introduced while we were preparing the prototype, often in the implementation  process.

Common FPGA implementation issues

Despite all our best efforts, initial FPGA implementations can be different to the  intended implementation. The following list describes common issues with initial  FPGA implementations:  

• Timing violations: timing analysis built into synthesis and place & route  tools operates upon each FPGA in isolation. Timing violations across  multiple FPGAs, for example, on paths routed through FPGAs or between  clock domains in different FPGAs, would not be highlighted during normal  synthesis and place & route.  

• Unintended logic removal: it may not be obvious at first, but modules and  IO seem to “disappear” from the resulting FPGA implementation due to  minimization during synthesis. The common cause is improper  connectivity or improper modules and core instantiations that result in undriven logic, which is subsequently minimized. Early detection of  accidental logic removal can save valuable FPGA implementation time and  bench debug time. A review of warnings and the FPGA resource  utilization, especially IO, after design completion will indicate unintended  logic removal, for example if there is a sudden unexplained drop in IO.  

• Improper inter-FPGA connectivity: despite all efforts, due to improper  pin location constraints, the place & route process will assign IO to  unintended pins resulting in unintended and incorrect inter-FPGA  322 Chapter 12: Breaking out of the lab: the prototype in the field connectivity. This often happens the first time a design is implemented or  drastically modified. The remedy for this issue is to carefully examine the  pin location report generated by the place & route tools for each FPGA and  compare against the intended pin assignments.

Let’s consider each of these types of implementation issues in more detail.

Timing violations

The design may have timing violations within an FPGA or between FPGAs. It is  always advisable to run the static timing analysis with proper constraints on the post  place & route netlist and make sure that all timing constraints are met. Even though  FPGA timing reports consider worst case process, voltage and temperature (PVT)  conditions, if timing constraints as reported by the timing analysis are not met, there  is no guarantee the design will work on board.  

Here are some common timing issues we might face while analyzing timing on a  design and some tips on how to handle them:  

Internal hold time: if we are prototyping the design at a very slow clock  rate then we might be tempted to think that there cannot be timing  violations on today’s fast FPGAs. As a result we might not choose to run  full timing analysis. However even in such scenarios, there can be hold  time violations which will prevent the design from working on board. So,  even if the timing requirements are very relaxed, FPGA internal timing  should be analyzed with appropriate constraints and with the use of  minimum timing models. 

Inter-FPGA delays: careful analysis of inter-FPGA timing, taking in  account board delays, should be made to ensure inter-FPGA setup and hold  times are met. For this we need to know the appropriate board and cable  delays to account for these either within the timing model or by setting  appropriate constraints for whole-board timing analysis. If off-the-shelf  standard prototyping boards and their associated standard cables are used  for inter-FPGA connection then the expected board and cable delays  should be available from the vendors. For example, the Synopsys HAPS  series boards are designed to have track delays matched across the boards  and between boards and delays are specified in terms of two constants, X  and Y. We referred to these X and Y delay specs in the PLL discussion in  section 5.3.1 and for a HAPS-54 boards for example the nominal values  are X=0.44ns and Y=1.45ns. These values can be used to add delays into  calculations for inter-FPGA delays.  

• It is an advantage if the boards are pre-characterized for use in timing  analysis tools, as is the case for HAPS boards in the Certify® built-in  timing analyzer. If a custom-made board or manual partitioning is used then whole-board timing analysis would be more complex but could still  be done as long-delay data could be provided by the PCB development and  fabrication tools. 

Timing complexity of TDM: If the prototype uses pin multiplexing to  share one connection between FPGAs for carrying multiple signals, then the effect of this optimization on the inter-FPGA timing should be  analyzed carefully. See chapter 7 for more consideration of the timing  effects of signal multiplexing. 

Input and output timing: constraints on all the FPGA ports through  which the design interacts with the external world should not be neglected.  Without proper IO constraints, the interface with the external peripherals  may not work reliably. Timing is easier to meet if the local FF features of  the FPGA’s IO buffers are used to their full extent, thus removing a  potentially long routing delay from an internal FF to reach the IO Blocks (IOBs). As was shown in chapter 3, all the IOBs of the FPGAs have  dedicated FFs for data input, output and tri-state enables. These FFs should  be used by default by the synthesis and place & route tools, but may  require some intervention using the tool-specific attributes.  

If the FFs have not been used properly during the implementation flow  then timing problems may be introduced. For example, consider the  simplified view of a typical FPGA IOB shown in Figure 138 and its use as a tri-state output pin. The recommendation is to use the FFs available in  IOB for both the tristate control path and the output datapath.For tristate  signals, timing advantage of using the IOB FF can be realized by using  both the tri-state FF and output data FF. If neither FF is used then the extra  routing delay in that path will nullify the timing advantage obtained in the  other path.  

For example, if the output path uses its IOB FF and the tristate control path does not use its FF then the output data may indeed arrive sooner at the tristate buffer. The output data would only reach the PAD after the tri-state control reaches the tri-state buffer and after that buffer’s switch-on delay.  In effect, the timing advantage produced by using the IOB FF in the  datapath is offset by the non-optimal delay in the tri-state control path. 

Similarly, if the tri-state control path uses the IOB FF and output datapath  does not use its FF then the tri-state buffer will put the previous output FF  data at the output PAD when the tri-state control is asserted and after a  little while the new data will arrive at the output PAD. This skew in the  arrival time of tri-state control and data may introduce glitches at the PAD  output which may or may not be tolerable at the external destination.

Inter-clock timing: as is true for all logic design, we should take special  care with multi-clock systems when signals traverse between FFs running  in different clock domains. If the two clocks are asynchronous to each  other then there can be setup and hold issues leading to metastability  (check references for more background on metastability). Avoiding  metastability between domains is as much a problem in FPGA as it is in  SoC design so similar care should be taken. In fact, the measures taken in  the SoC RTL to avoid or tolerate metastability (e.g., double-latching using  two FFs in series on the receiving clock) can be transferred directly into  the FPGA but we should also apply all the timing constraints used for  ASIC to the FPGA.  

• We can ensure that the probability of meta-stable states is within reason or  otherwise take counter-measures. In either case we should reassure  ourselves that the problem is under control before going onto the board.  Timing analysis for each FPGA may indicate where metastability can  occur. For example, Synopsys FPGA synthesis tools generate specific  warning messages for signals traversing clock domain boundaries and  these messages can be checked manually or by a script. This becomes more  complex when analyzing multiple FPGAs. One suggestion is to avoid  setting partition boundaries so that the sending FPGA and receiving FPGA  are on different asynchronous clocks.  

Gated clock timing: if there are gated clocks in the design which are not  converted then there is potential for hold-time violations to occur inside the  FPGAs, caused by clock skew and possibly even glitches on poorly constructed clock gates (although the latter is a real bug and good to feed back to the SoC team). Preferably all gated clocks will have been  converted using techniques discussed in chapter 7, but it is possible that  some have been omitted or not converted correctly leading to timing issues  after place & route. Close inspection of the synthesis tools report files for  messages regarding the success or failure of gated-clock conversion can  often give clues to the cause of unexpected timing violations. 

Internal clock skew and glitches: we should be using the built-in global  and regional clock networks which are inherently glitch free. Check for  low-fanout clocks that have been accidentally mapped to non-global  resources.  

Timing on multiplexed interconnect: as we saw in chapter 8, the timing  of TDM connections between FPGAs can be crucial. We need to ensure  that the timing constraints on both the design clock and the transfer clock are correct and confirm that they are met by performing the post place &  route timing analysis. It is especially important to re-confirm that the time  delay between design and transfer clock is within limits, that the on-board  flight time is not longer than expected, and that, if asynchronous TDM is in  use, synchronization has been included between the design and transfer  clock domains.

We can see from these examples above that thorough timing analysis of the whole board is very useful and can discover many possible causes of non-operation of a  prototype before we get to the stage of introducing the design onto the board. Any  of these timing problems might manifest themselves obscurely on the board itself  and potentially take days to uncover in the lab.

Improper inter-FPGA connectivity

Another difficult-to-find problem is improper connectivity between the FPGAs on  the board. That is, signals between FPGAs which are misplaced so that the source  and destination pins are not on the same board trace. Keeping correct contiguity  between FPGA pins should be a matter of disciplined specification and automation.  However, as the excerpt from a top-level partition view in Figure 139 shows us (or  rather scares us), there will be thousands of signal-carrying pins on the FPGAs in a  typical sized SoC prototype. If any one of these pins is misplaced (i.e., the signal is  placed on the wrong pin) then the design will probably not behave as expected in  the lab and the reason might prove very hard to find.  

If this happens then it is most likely during a design’s first implementation pass or  after major RTL modification or addition of new blocks of RTL at the top level.

Certify® Partition View illustrating complexity of signals between.png

Pin location constraints are passed through different tools in the flow, usually from  the partition tool to the synthesis tool to the place & route tool. If any of the  different tools in the flow drops a pin location constraint for some reason then the place & route tool will randomly assign a pin location with only low probability that  it will be in the correct place. The idea of pin placements being “lost” might seem  unlikely but when a design stops working, especially after only a small design  change, then this is a good place to look first.

Possible reasons for pin misplacement include accidentally setting a tool control to  not forward annotate location constraints. Another mistake is to rely on defaults that  change from time-to-time or over tool revisions. A further example happens when  we add some new IO functionality but neglect to add all of the extra pin locations.  In every case, however, we will notice the missing pin constraints if we look at the  relevant reports.  

We should carefully examine the pin location report generated by the place & route  tools for each FPGA and compare against the intended pin assignments. In the case  of the Xilinx® place & route tools, this is called a PAD report. Human inspection of  such reports, which can be many thousands of lines long, might eventually lead to  an error so some scripting is recommended. The script would open and read the pin  location reports and look for messages of unplaced pins but while doing so, could  also be looking for other key messages, for example, missed timing constraints or  over-used resources. In the case of the pin locations, we might maintain a “golden”  pin location reference for all the FPGAs so that the script can automatically compare this against the created PAD reports after each design iteration. This would  add an extra layer of confidence on top of looking for the relevant “missing  constraint” message. In addition, we could set up other scripted checks in order to  compare any pin location information at intermediate steps in the flow, for example,  between the location constraints passed forward by partition/synthesis to the place & route, and those present in the final PAD files.

It is worth noting that if a commercial partitioning tool is used then inter-FPGA pin  location constraints should be generated automatically and consistently by the tool,  perhaps on top of any user-specific locations. In the case of Certify, if the connectivity in the raw board description files are correct and the assignments of all  the inter FPGA signals to the board traces are complete then there is a very minimal  risk that the pin locations will be lost of misplaced by the end of the flow.  

On the other hand, if we manually partition the design, then we need to carefully  assign the pin location constraints for all the individual inter-FPGA connections as  well as make sure that these are propagated correctly to the place & route stage,  which can become a tedious and therefore error-prone task if not scripted or  otherwise automated. 

If all this attention to setting and checking pin locations seems rather paranoid then  it is worth remembering that there might be thousands of signals crossing between  FPGAs. In the final SoC, these signals will be buried within the SoC design and the  verification team will go to great lengths to ensure continuity both in the netlist and  in the actual physical layout. A single pin location misplacement on one FPGA  would be as damaging as a single broken piece of metallization in an SoC device  and perhaps harder to find (albeit easier to fix). Therefore it pays to be confident of  our pin locations before we load the design into the FPGAs on the board.

Improper connectivity to the outside world

As well as between FPGAs, we need the proper connectivity between the FPGAs  and any external interfaces. Some of the simplest but most crucial external  interfaces are clock and reset sources, configuration ports and debug ports. The  more sophisticated connections would be interfaces to external components like  DDR SDRAM, FLASH and other external interfaces.  

Once again, the implementation tools rely on there being correct and complete  information about physical on-board traces between the FPGAs and from the  FPGAs to the external daughter cards which may house the DDR etc. On a modular  system, with different boards being connected together, the board description will  be a hierarchy of smaller descriptions of the sub-systems and with a little care, we  can easily ensure that the hierarchy is consistent and that the sub-boards are  themselves correctly described. It may be worthwhile to methodically step through  the board description and compare it to the physical connection of the boards, however, some systems will also allow this to be done automatically using scan  techniques to interrogate each device on the boundary scan chain and check that  they appear where they should, according to the board description.  

A commercial board vendor will be able to provide complete connection  information for the FPGA boards and for their connection to daughter boards for  this purpose. In such a board description, inter-FPGA connections on the board will  be labeled generically because they might conceivably carry any signal (or signals)  in the design.  

For connections to dedicated external interfaces, however, the signal on each  connection is probably fixed e.g., a design signal called “acknowledge” must go to  the “ack” pin on the daughter card’s test chip. In those cases, it can help to also give  the traces in the board description the same meaningful name to make it is easier to  verify the connectivity to external interfaces and match signals to traces. So for  example, we can easily check that a trace called “ack” on the board is connected to  a pin called “ack” on the daughter card and is carrying a signal called “ack” from  the design. This naming discipline is especially useful for designs with wide bus  signals.  

Some advanced partitioning tools will be able to recognize that signals and traces,  or signals and daughter-card pins, have the same name and therefore quickly make  an automatic assignment of the signal to the correct trace, which if the board  description is correct, will automatically also assign the signal to the correct FPGA  pin(s). All of our pin assignments would therefore be correct by construction.

Incorrect FPGA IO pad configuration

When connecting various components at the board level, care must be paid to the  logic levels of all interconnecting signals and be sure they are all compatible at the  interface points. As well as having the correct physical connection, the inter-FPGA  signals and those between the FPGAs and other components must be swinging  between the correct voltage levels. As we saw in chapter 3, FPGA cores run at a common internal voltage but their IO pins have great flexibility and can support  multiple voltage standards. The required IO voltage is configured at each FPGA pin  (or usually for banks of adjacent pins) to operate with required IO standards and is  controlled by applying the correct property or attribute during synthesis or place &  route.  

This degree of flexibility must be controlled because a pin driving to one voltage  standard may be misinterpreted if interfacing with a pin set to receive in a different standard. This may seem obvious but as with the physical connection of the pins,  the scale of dealing with voltage standards on thousands of signals adds its own  challenge. For all the inter-FPGA connections, the IO standards of the driving  FPGA pin and the driven FPGA pin should be the same.

The partition tools should automatically take care of assigning the same IO  standards for connected pins, or warning when they are not. The default IO standard  for the synthesis and place & route may also suffice for inter-FPGA connections but  it is not recommended to rely only on default operation. In addition, care should be  taken when the design is manually partitioned.  The board-level environment may also constrain the voltage requirement and we  will need to move away from the default IO voltage settings. Then there are special  considerations for differential signals compared with single-ended.  

Let’s look at some of these board-level voltage issues here:

The correct IO standard for differential signals: on a prototyping board,  it is likely that only a subset of traces can be used in differential mode. Not  only must pin location constraints be correct to bring signals to the correct  FPGA pins to use these traces but also the pin’s IO standard must be set  correctly. It is a subtle mistake to have a differential signal standard at one  end and single-ended standard at the other, so that even though the pins are  physically connected, the signal will not pass correctly between FPGAs. 

Voltage requirements for peripherals: for all the connections with  external chips and IOs, we should first identify the IO standard of the pins  of the external chips and IOs and then apply the same IO standards for the  corresponding FPGA pin connected. 

IO voltage supply to FPGA: while configuring a certain IO standard for a  certain pin in an FPGA, we should connect the necessary voltage source to  the VCCO pins and VREF pins of the FPGA’s corresponding IO bank. As discussed in chapters 5 and 6, the boards must have the flexibility to be  able to switch different voltage supplies to different banks. We must then  make sure that proper supply voltages are connected to all the VCCO pins  according to the chosen IO standards. This is usually a task of setting  jumpers or switches or, in some cases, of using a supervisor program to  control on-board programmable switches, as seen in chapter 5. 

Termination settings: some of the IO standards require appropriate  termination impedance at the transmitting and receiving ends. The FPGA’s  IO pads can be configured for different kinds of terminations and different  impedances, for example using the digitally controlled impedance (DCI)  feature in Xilinx® FPGAs. Once again, it is worth a quick check at the end  of the flow to make sure that these are configured as expected. 

Drive current: the ports in the FPGA which interact with external chipsets  should be able to source or sink the required amount of drive current as  specified in the data sheets of the external chipsets. The FPGA’s IOs can  be configured to source and sink different currents, for example on normal  LVCMOS pins on a Virtex®-6 FPGA, this can be programmed to be  between 2mA and 24mA. This would have already been considered during the early stages of the prototyping project but it is important that the FPGA  pins be configured to supply the current or else inconsistent performance  will result. This is one of those user-errors that can remain hidden for much  of the project as in lab conditions we might be lucky while the peripheral is  not running at full spec or otherwise does not demand full current from the  FPGA pin. However, at other times the behavior of the design might  become inconsistent for no apparent reason because the software might be  using a feature of the peripheral not previously enabled.  

Poor signal integrity: commercial prototyping boards typically have  acceptable point-to-point signal integrity at the board level. the same board  design will probably have been previously used for multiple designs  worldwide. However, first time we use a custom-built boards we may need  some careful analysis and debug for such issues. Even off-the-shelf  systems can exhibit poor signal quality with “exotic” connectivity, for  example, when creating buses shared by multiple FPGAs or cables that are  too long or not properly terminated. It is always best to identify and fix the  root causes when possible before programming the board with the real  DUT. For marginal noise or signal quality issues, we can sometimes  change the programmable slew rate and drive strength at the FPGA IO pin.

NOTE: run scripts to check IO consistency

Although most of the inter-FPGA IO considerations above will be managed by the  partitioning tools and therefore consistent by construction, we should maintain a  “golden” reference for IO standard, voltage, drive current and placement for the  critical pins on the FPGA and certainly between the FPGAs and external  peripherals. We can then compare the golden reference against the created PAD  report for each FPGA, however, it is wrong to have to make repetitive manual  checks on every iteration of the design, so it is worth taking the time to script these  kinds of checks.  

We have mentioned how scripts can be used to automate these post-implementation,  pre-board checks. Tools will have their own commands for generating reports on a  large number of details, including top-level ports. A script can make multiple  checks on such reports in the same pass, for example verifying that the IO standards  could be combined with the pin location constraints, checking that every top-level  port or partition boundary signal has an equivalent entry in the FPGA location  constraints. Automatically running these scripts in a larger makefile process will  prevent running on into long tool passes using data that is incomplete, and wasting a  lot of time as a result.

  • XC2C512-7FT256I

    Manufacturer:Xilinx

  • CPLD CoolRunner -II Family 12K Gates 512 Macro Cells 179MHz 0.18um Technology 1.8V 256-Pin FTBGA
  • Product Categories: CPLDs

    Lifecycle:Active Active

    RoHS: No RoHS

  • XC3SD3400A-4FG676C

    Manufacturer:Xilinx

  • FPGA Spartan-3A DSP Family 3.4M Gates 53712 Cells 667MHz 90nm Technology 1.2V 676-Pin FBGA
  • Product Categories: FPGAs

    Lifecycle:Active Active

    RoHS:

  • XC3SD3400A-5CSG484C

    Manufacturer:Xilinx

  • FPGA Spartan-3A DSP Family 3.4M Gates 53712 Cells 770MHz 90nm Technology 1.2V 484-Pin LCSBGA
  • Product Categories: FPGAs

    Lifecycle:Active Active

    RoHS:

  • XC4002A-5PC84C

    Manufacturer:Xilinx

  • FPGA XC4000A Family 2K Gates 64 Cells 125MHz 5V 84-Pin PLCC
  • Product Categories:

    Lifecycle:Obsolete -

    RoHS: No RoHS

  • XC2C512-7PQG208C

    Manufacturer:Xilinx

  • CPLD CoolRunner -II Family 12K Gates 512 Macro Cells 179MHz 0.18um Technology 1.8V 208-Pin PQFP
  • Product Categories: CPLDs

    Lifecycle:Active Active

    RoHS:

Need Help?

Support

If you have any questions about the product and related issues, Please contact us.