sep
Exjobbspresentation: High Throughput Matrix Inversion Implementation for Distributed MIMO Testbed on ZCU102
Yifan Liu och Wei-tang Wang presenterar sitt exjobb High Throughput Matrix Inversion Implementation for Distributed MIMO Testbed on ZCU102 fredag 4 September i E2347a.
Matrix inversion is a computationally demanding operation in the distributed multiple-input multiple-output (D-MIMO) systems. In zero-forcing detection, channel information from multiple access points is combined into a Gram matrix whose inverse is required to suppress multi-user interference. Computing this inverse accurately under strict latency and throughput constraints is challenging on an FPGA, where both arithmetic and memory resources are limited.
This thesis presents a high-throughput 32 × 32 matrix inversion implementation for the D-MIMO testbed at Lund University. Instead of directly inverting the full Gram matrix, a structured preconditioner is first constructed from its tri-diagonal region. To reduce the hardware cost of the preconditioner inversion, the full matrix is approximated by diagonal blocks. Each block is inverted using Cuppen’s divide-and-conquer eigenvalue decomposition after a complex-to-real similarity transformation. The resulting block-diagonal inverse is then refined using the original 32 × 32 Gram matrix through a factorized Neumann Series (NS).
The accelerator targets the AMD ZCU102 FPGA platform. The divide-and-conquer path is time-multiplexed eight-fold, while four secular solvers operate in parallel. The 32 × 32 NS engine uses two physical 16 × 16 processing element arrays that are reused across four logical matrix tiles. Banked memories, local ping-pong storage, and concurrent execution between the two arrays are used to reduce data movement and expose parallelism without instantiating a full 32 × 32 matrix-multiplication array.
Post-implementation utilization report illustrates the high hardware cost of the algorithm. The measured end-to-end computation latency is 2205 clock cycles for a Gram matrix inversion with three NS iterations, and a throughput of 1113 clock cycles per Gram inversion. RTL simulation over 600 distinct Gram matrices gives an average mean-square error of 4.35 × 10−3 and a 64-QAM bit error rate of approximately 1.22 × 10−2. The results confirm that the proposed implementation satisfies the throughput requirement of 1329 clock cycles per inversion while maintaining reasonable accuracy, thereby validating the practical viability of the design.
Handledare: Sijia Cheng
Examinator: Pietro Andreani
Om evenemanget
Plats:
E:2347a
Kontakt:
susanna [dot] lonnqvist [at] eit [dot] lth [dot] se