Dissertation > Excellent graduate degree dissertation topics show

The Design and Implementation of Vector Memory Unit of Multi-Width SIMD DSP

Author: HuangYuanGuang
Tutor: ChenHaiYan
School: National University of Defense Science and Technology
Course: Software Engineering
Keywords: SIMD VM Memory Conflict VARU VWRU Finite Share Shuffle Vector Conditional Access
CLC: TP333
Type: Master's thesis
Year: 2012
Downloads: 13
Quote: 0
Read: Download Dissertation

Abstract


Over the past decade, with the development of integrated circuits and computer technology,the performance growth rate of the CPU is nearly60%every year,but the performance ofmemory access has been only improved by7%[1].The “memory wall” problem caused bymemory bandwidth and latency has become the performance bottleneck that restricts themicroprocessor to further improve. The Digital Signal Processor(DSP) based on multi-widthSingle Instruction-stream Multiple Data-stream(SIMD) architecture works for high density dataprocessing. It integrates multiple Vector Processing Element(VPE) that need higher memoryaccess performance. How to provide sufficient memory bandwidth for the VPEs and how toreduce the data shuffle and other additional operations between the VPEs of mult-width SIMDDSP to improve the memory accessing efficiency and reduce power consumption have becomea important issue of designing a vector memory system.YHFT-Matrix DSP for Software Defined Radio(SDR) base station is independentlyresearched and developed by the Microelectronics and Microprocessor research institute ofNational University Defense of Technology. It adopts10issues Very Long InstructionWord(VLIW) and multi-width SIMD architecture. Its Vector Processing Unit(VPU) contains16vector process elements, each of which contains two multiply-add units and other ALU. So itrequires a higher data throughput and memory bandwidth in order to make full use of thecomputing power of VPU. So we design and implement a novel and large capacity on-chipvector memory(VM) to save a large amounts of data for VPU operating.The VM designs a special Vector Address Generation Unit(VAGU) which supports bothlinear and circular addressing. The total memory capacity of VM is1M bytes. Its memory isorganized by the Multiple-Bank interleaved-addressing of low-bit address(MIDB) with a DoubleBuffers architecture. It fulfils the demand of multiple vector data parallel memory accessing ofVPU with a small cost of lower area and power and reduces the parallel memory accessingconflicts. In order to accelerate the related communication algorithms, we also implement aVector Access Reorder Unit(VARU) and a Vector Write-back Reorder Unit(VWRU) in the VM.So the VM can support the non-aligned and conditional vector accessing memory. This methodcould make all the VPEs of VPU share the VM finitely and accessing VM conditionally. VM hasachieved the design target of supporting512Gbps vector data accessing,256Gbps DMA dataaccessing and32Gbps scalar data accessing performance. VM can sustain continuous vector byteand halfword accessing after later logic optimization.The YHFT-QMBase based four YHFT-Matrix DSP has been successfully tapeout now.After the logic verification and testing, the test result showed that the design function of theVM is correct. The processing frequency of the VM has reached up to500MHz or above. Thefrequency can reach700MHz after the logic optimization. The MIDB architecture of VM cansignificantly reduce the access conflicts. The way of finite sharing and vector conditionalaccessing memory can reduce or eliminate the shuffle operation of related algorithm,compress code density, and accelerate algorithm implementation.

Related Dissertations

  1. Research on WiMAX Channel Codec Technology Based on Virtual Radio,TN911.22
  2. Research of Multimedia Enhancement Unit in the Embedded Processor,TP332
  3. Architecture of Reconfigutable Deblocking Filter Based on SIMD Technology,TN919.81
  4. The Research and Implementation of High Performance Vector FMAC Unit for LTE,TN929.5
  5. Design and Implementation of the Instruction Fetch Unit and Multiple Instruction Flows Extension in the YHFT-Matrix DSP,TP368.1
  6. The Optimization of High Performance MapReduce FairScheduler and the Implementation on Simulator of Huge Scale Cluster,TP311.13
  7. Study of Optical Implementation and Integrated Technology of Free-Space Optical Interconnection Network,O438
  8. Key Technologies and Optimization for Dynamic Migration of Virtual Machines in Cloud Computing,TP302
  9. Study on high - performance low-power multi-core processors,TP332
  10. Malicious Code Detection and Auditing Based on Hardware Virtualization Technology,TP393.08
  11. Automatic Generation and Optimization of Data Permutation Instructions for SIMD Devices,TP332
  12. The Research of Key Techniques of Trusted -Application Execution Mechanism Based on Virtual Machine Technology,TP309
  13. Design and Implementation of High Performance Encryption/Decryption System Based on FPGA,TP309.7
  14. Thermodynamic Optimization for Natural Gas Driven Vuilleumier Cycle Heat Pump,TU831
  15. Research on Decoding Algorithms and Theoretical Analysis of LDPC Codes,TN911.2
  16. 65nm ALU Full-custom Design Technology and Methodology,TN402
  17. Web-based Vocational Examination System Design and Implementation,TP311.52
  18. Accelerators and Update Algorithms for Ray Tracing Dynamic Scenes,TP391.41
  19. Hardware Design of an Embedded System Based on Dual PowerPC 7447A Cpus,TP368.12
  20. ARM9 Based Design Optimization of MPEG-4 Video Decoder,TN919.81
  21. Research and Realization of H.264/AVC Software Decoder Optimization,TN919.81

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Electronic digital computer (not a continuous role in computer ) > Memory
© 2012 www.DissertationTopic.Net  Mobile