| V0 | V1 | |
|---|---|---|
| Times(s) | 28.8741 | 0.0729225 |
| GFLOPS | 0.0569426 | 22.5468 |
| LUT | 8947 | 25716 |
| LUTAsMem | 695 | 728 |
| REG | 9378 | 39736 |
| BRAM | 79 | 79 |
| URAM | 407 | 467 |
| DSP | 7 | 322 |
| Freq(MHz) | 200 | 200 |
| WNS(ns) | 0.035 | 0.057 |
With no optimization at all, from v++.log, we know II=7. Therefore, the estimated running time is 7*4.11 = 28.77s.
After changing the loop order of convolution into this,
for (int h = 0; h < kOutImSize; ++h) {
for(int j = 0; j < kNum; ++j){
for (int p = 0; p < kKernel; ++p) {
for (int q = 0; q < kKernel; ++q){
for (int w = 0; w < kOutImSize; ++w) {
#pragma HLS PIPELINE II=1
for (int i = 0; i < kNum; ++i) {
local_output[h][w][i] += local_input[h+p][w+q][j] * local_weight[p][q][i][j];
}
}
}
}
}
}The estimated speed-up is 64x. Therefore, the estimated running time is 4.11/64 = 0.064s.
To make sure the II=1, the following optimization methods are used.
- Change
local_outputfromRAM_1P_URAMtoRAM_2P_URAM, i.e., dual-port URAM. This change is crucial because, in the innermost loop, there are read-modify-write operations on the same data unit. With dual-port memory, one port is used for reading and the other for writing, greatly reducing read-write conflicts and scheduling difficulty, thus increasing the likelihood of achieving II=1. - Perform complete partitioning on the
idimension (channel dimension) oflocal_output(#pragma HLS ARRAY_PARTITION complete dim=3). This means that for every spatial position (h,w)(h, w), all elements in theidimension oflocal_output[h][w][i]are split into independent storage units, enabling simultaneous access to different channels in the same cycle. - Perform complete partitioning on the
idimension oflocal_weight(#pragma HLS ARRAY_PARTITION variable=local_weight complete dim=3). By completely partitioning the weights of the kernel across the channel dimension, multiple channel weight data can be accessed simultaneously in the innermost loop. This removes the limitation of a single memory port, accelerating access and reducing bandwidth constraints. - Use the
#pragma HLS PIPELINE II=1directive in the innermost loop to instruct the tool to initiate a new iteration every clock cycle, thereby achieving instruction-level parallelism (ILP). By combining array partitioning and dual-port memory, the tool has the opportunity to complete multiple data accesses and computations within a single clock cycle, striving to achieve the performance target of II=1.