رسائل ماجيستير
English
1997
Organization of parallel memories
Husam Saad Abu-Haimed , Mayez Al-Mouhamed
كلية علوم وهندسة الحاسب الآلي-جامعة الملك فهد للبترول والمعادن · السعودية
The use of parallel memories has been the most promising technique to bridge the gap between high performance processors and available memories. By having N parallel memories, we aim to have a total memory bandwidth of NB, where B is the bandwidth of NB, where B is the bandwidth of a single memory bank. However, using simple interleaving techniques, the effective bandwidth becomes much less than that and is about √NB [21]. This is because of serialization of memory access which happens when a number of elements that will be referenced at the same time are stored into the same memory bank. Having low memory throughput can severely affect the overall performance of vector machines and SIMD systems. To minimize memory conflicts, researchrs have considered storage schemes for conflict-free access of frequently used patterns like rows, columns, and power of 2 patterns and strides. In this thesis, we consider power of 2 patterns as well as arbitrary strides. A new approach for combining different patterns into one linear bitwise storage scheme is proposed. We use 5 different approaches to construct combined storage schemes: 3 coloring-based heuristics, a Neural Networks approach and a Genetic Algorithms approach. In the case of powr of 2 patterns, swe were able to hit the lower bound on access time for small problems. For large problems, we achieved small deviations from the lowr bound (5% - 29%). In the case of arbitrary strides, access time of our schemes were 10% less than the best known bitwise schemes [41]. In addition, our schemes work with any power of 2 number of memories, while some other schemes work with a fixed number of memories like [41] which works with 8 memories.