This website uses cookies. By using this site, you consent to the use of cookies. For more information, please take a look at our Privacy Policy.
Home > FPGA Technical Tutorials > FPGA-Based Prototyping Methodology > THE FUTURE OF PROTOTYPING > Prototyping future: networking

TABLE OF CONTENTS

Xilinx FPGA FPGA Forum

Prototyping future: networking

FONT SIZE : AAA

In contrast to the mobile wireless and consumer application spaces, networking is  not determined by end-user applications but completely driven by data rates and  per-packet processing requirements. Nevertheless, given the need for product  flexibility and configuration options, software is once again used to determine  processing functionality and the hardware is increasingly designed to prioritize the  efficient execution of the software.  

As already outlined in chapter 1, the networking application domain is moving  towards architectures using multiple CPU cores and underlying flexible  interconnect fabrics. As predicted by ITRS and shown in chapter 1, the die area will  remain constant, the number of processors per design is predicted to grow by 1.4x  every year, on average. This will have a profound impact on the future of  prototyping given that the application partitioning between the processors has to be  properly tested and verified prior to silicon production.  

Considering multimedia applications briefly, a fair amount of the tasks to be  distributed between processors can be pre-scheduled. Proper execution can be  verified using virtual prototypes as they allow very efficient control of execution as  well as the necessary debug visibility into hardware and software co-dependencies.  For a networking application, however, incoming traffic is distributed to packet  processors and on-chip accelerators. Given the inherently parallel nature of packet  processing, the choice of compute resources is done at runtime, which means that  less pre-scheduling is required. Nevertheless, the various options of runtime  scheduling will need to be prototyped prior to committing to silicon, at as realistic  speed as possible. FPGA-based prototypes and virtual prototypes will be enhanced  to collect appropriate performance and debug data to optimize scheduling  algorithms. How the packet rates increase for different transfer rates and how packet processing  times get shorter is illustrated in Figure 166.

Packet rates and processing times.png

As an example for 10Gb Ethernet, 14.9 million packets have to be processed every  67 nanoseconds. To allow any kind of prototyping, both capacity and execution  speed of prototypes will have to improve. For virtual prototypes it is likely that this can only be achieved using parallel, distributed simulation. FPGA prototypes will  have to overcome capacity limitations using improved scheduling and partitioning  algorithms as well as intelligent stacking of prototypes themselves. 

Figure 167 illustrates how the different functions performed in the networking  control and data plane have evolved over time as processors have become more  capable. It also shows that the complexity of processing demands has been  outpacing the development of new processors. Developers have therefore moved to  multiple processor cores and the most complex challenge has become to efficiently  partition tasks across them.

In the control plane, the “in-band control traffic” is handled with the actual  connection processing for routing/session establishment of the network protocol in  use.

Functionality distribution in Data and Control Plane.png

Taking into account the underlying multi-processing nature of the hardware, task  distribution is important. Each “task” within the control processing is complex and  offers limited internal parallelism. The tasks themselves are fairly independent and  can be assigned individually to dedicated CPUs.

Unfortunately, network applications have processing loads that come in bursts,  which leads to CPUs that are powerful enough to handle the peak load, but  otherwise underused. As a result, these very capable processors are not an optimally  efficient use of silicon. Given the dependency of the processing requirements on the  actual networking traffic it is difficult, if not impossible, to assign tasks at compile  time to different processors. 

The data plane functionality is focused on forwarding packets and translates  information from control traffic into device-specific data structures. The data plane's  code is packet-processing code and easy to run on multiple cores operating in  parallel. This code can be more easily partitioned across multiple cores, because  packets can be processed in parallel as multiple cores run identical instances of the  packet-processing datapath code. 

Again, much like for control code, the actual assignment of processing units to tasks  is highly dependent in the actual network traffic and should be done at runtime.  Assignment of tasks to processors at compile time is, again, difficult if not  impossible.

We predict that prototyping of networking systems will largely continue to be done  in a hierarchical fashion. The individual processing units will continue to be tested  against their packet-processing requirements but, given the increased complexity of Figure 167 : Functionality distribution in Data and Control Plane (Source: AMCC) FPGA-Based Prototyping Methodology Manual 397  those requirements, it will become too risky to commit to hardware without proper  prototyping.

As discussed earlier, in contrast to mobile wireless and consumer applications the  assignment of processing tasks to processing units is not done at compile time and is  also not separated as a set of user applications from the hardware through operating  systems. The type of software used in networking applications is much more “bare  metal” and tightly coupled to the processors upon which it runs. Hardware  dependency in software is difficult to model without the hardware being present in  some form. Hence software debug is posing different challenges, i.e., requires  debug at runtime and is also driving requirements for redundancy, all of which will  make prototyping of all types even more compelling than it is today already.



  • XC5VFX100T-2FF1738C

    Manufacturer:Xilinx

  • FPGA Virtex-5 FXT Family 65nm Technology 1V 1738-Pin FCBGA
  • Product Categories: Connecteurs

    Lifecycle:Active Active

    RoHS: No RoHS

  • XC5VFX100T-2FFG1738I

    Manufacturer:Xilinx

  • FPGA Virtex-5 FXT Family 65nm Technology 1V 1738-Pin FCBGA
  • Product Categories: Connecteurs&adapteurs

    Lifecycle:Active Active

    RoHS:

  • XC5VFX130T-1FFG1738C

    Manufacturer:Xilinx

  • FPGA Virtex-5 FXT Family 65nm Technology 1V 1738-Pin FCBGA
  • Product Categories: Industrial components

    Lifecycle:Active Active

    RoHS:

  • XC4028XL-1HQ304C

    Manufacturer:Xilinx

  • FPGA XC4000X Family 28K Gates 2432 Cells 0.35um Technology 3.3V 304-Pin HSPQFP EP
  • Product Categories: FPGAs (Field Programmable Gate Array)

    Lifecycle:Obsolete -

    RoHS: No RoHS

  • XCS20XL-5PQ208C

    Manufacturer:Xilinx

  • FPGA Spartan-XL Family 20K Gates 950 Cells 250MHz 3.3V 208-Pin HSPQFP EP
  • Product Categories: FPGAs (Field Programmable Gate Array)

    Lifecycle:Obsolete -

    RoHS: No RoHS

Need Help?

Support

If you have any questions about the product and related issues, Please contact us.