Release of Results: Integrated Memory-Compute Graph Computing Platform Based on Flash Memory and FPGA

01

Field of Application

Edge Computing, Edge Storage, Intelligent Graph Computing

02

Project Introduction

2.1 Pain Points

Graphs, as an important data type, are increasingly prevalent in large-scale data, such as biological information network data, social network data, search web data, and knowledge graphs. The scale of graph data is growing rapidly, with the number of nodes reaching billions and edges reaching tens or even hundreds of billions. Therefore, efficiently processing such large-scale graph data poses a significant challenge for graph computing. Compared to traditional structured data processing, graph model data is characterized by its enormous size and high sparsity, leading to irregular I/O access patterns. Sparse data results in non-contiguous memory access, causing a decline in I/O performance. Additionally, the substantial storage overhead means that the entire graph model data can only be partially cached in memory, necessitating frequent data movement between external storage and memory, which incurs significant I/O overhead and reduces the computational efficiency of graph computing.

2.2 Solution

(1) Overview of Core Technology: As shown in Figure 1, the integration of reconfigurable processors (FPGA) with high-performance, large-capacity flash storage can bring the graph computing processing unit closer to the data, migrating the core operators of graph computing nearer to the storage data (indicated by the green arrows in Figure 1). This leads to the proposal of an integrated memory-compute graph computing platform based on FPGA and flash memory, which significantly outperforms traditional computing architectures (as indicated by the red arrows in Figure 1).

Release of Results: Integrated Memory-Compute Graph Computing Platform Based on Flash Memory and FPGA

Figure 1 Basic Principle of Near-Data Memory-Compute Integration

(2) Advantages of System Hardware: As shown in Figure 2, the proposed near-data memory-compute integration unit for ultra-high-definition image recognition directly connects the reconfigurable FPGA chip with large-capacity non-volatile flash memory chips through PCB-level connections. The FPGA chip integrates sensor payload interfaces, high-performance intelligent computing accelerators, and SSD controllers. The SSD memory bandwidth can reach 2GB/s, with a storage capacity of 1TB that can be flexibly expanded. The high-performance intelligent computing accelerator supports high-performance and customized graph computing, intelligent computing, and scientific computing tasks. This system has a high level of integration and can perform real-time intelligent computing with high throughput.

Release of Results: Integrated Memory-Compute Graph Computing Platform Based on Flash Memory and FPGA

Figure 2 Hardware Architecture of Near-Data Memory-Compute Integration

(3) Advantages of System Software: As shown in Figure 3, the proposed memory-compute integration software system mainly consists of three parts: storage-compute integration request management technology, high-concurrency I/O management technology for flash storage devices, and near-data graph computing hardware acceleration technology. This software stack is built at the firmware layer between the file system and flash SSD, allowing users to initiate storage-compute integration requests directly in the user space of the file system, breaking down the programming and usage barriers of traditional memory-compute integration devices. Furthermore, the software system’s awareness of the hardware improves the parallelism of storage access and the efficiency of computation, facilitating the rapid construction of a related software ecosystem.

Release of Results: Integrated Memory-Compute Graph Computing Platform Based on Flash Memory and FPGA

Figure 3 Near-Data Memory-Compute Integration Software Architecture

2.3 Competitive Advantage Analysis

(1) Benchmark Companies: The computational storage products (Computational Storage Drive, CSD) from ScaleFlux, USA.

(2) Performance, Cost, and Maturity Analysis: The CSD products from ScaleFlux can support native data compression, deduplication, and other operations, with a theoretical peak bandwidth of up to 5GB/s and storage capacity reaching the 10TB level. The prototype system based on this technology, through hardware-software collaborative optimization, fully utilizes the parallelism at the solid-state storage channel level, greatly enhancing read/write bandwidth and reducing tail latency. By constructing a user-friendly file system, it provides a unified management framework for solid-state storage devices and intelligent computing units. In terms of application support, it not only supports native data compression and deduplication but also common computing tasks in user space such as machine learning and scientific computing. Currently, the bandwidth of the implemented prototype is 2GB/s, with a storage capacity of 1TB, which is currently below the performance metrics of the CSD series products. However, this technology has a highly integrated high-performance hardware platform and a user-friendly software interface and development environment, which has greater potential for improvement in specific hardware metrics and can more easily build a software ecosystem through engineering refinement, thus surpassing the CSD series competitors. Additionally, due to the domestic advantages of this project compared to foreign products, it can establish a presence in many vertical fields in China earlier, allowing the product to land first.

(3) Intellectual Property Layout: This project has obtained one Chinese invention patent, with another one currently under application. Competitors focus on providing native data computing acceleration and support, while this project aims for a broader range of applicability and market for user-space general computing, offering a wider range of application scenarios.

2.4 Market Application Scenarios

(1) Application Fields: Aerospace, meteorology, electricity, environmental monitoring, etc.

(2) Target Customers: Aerospace research institutes, provincial and municipal meteorological bureaus, electric power companies (State Grid, China Guodian, China Tower), etc.

(3) Market Size: Computational storage is considered the next generation of storage technology. According to Markets & Markets, the Chinese market for next-generation computational storage technology is expected to grow from 30 billion yuan in 2020 to 80 billion yuan by 2025, with the market covered by this project estimated to account for about 15%, predicting a total market volume of 12 billion yuan by 2025.

(4) Profit Model: Selling technical services or complete software and hardware products.

2.5 Development Plan

This project plans to transform through equity investment in a newly established company:

(1) Within one year of the company’s establishment: Utilize seed round financing to engineer the prototype system, improving performance to match that of foreign competitors like ScaleFlux; simultaneously promote the product prototype with aerospace, electricity, and meteorology departments to open up the market for product sales.

(2) In the second year after the company’s establishment: Plan to conduct angel round financing while expanding the technical and sales team to 20 people; complete 3-4 complete products of memory-compute integration systems.

(3) In the third year after the company’s establishment: Plan to conduct Series A financing, achieving sales of tens of millions and serving 50 customers, realizing 10 products of memory-compute integration systems, with performance metrics fully superior to the entire CSD product line of foreign competitors like ScaleFlux.

03

Cooperation Needs

1. Incubation Resources: Approximately 3-6 months of funding around 3 million yuan for engineering and productization needs, requiring a space for 10-15 people, about 30 square meters, located near Tsinghua University.

2. Application Scenarios: Edge applications for graph computing such as aerospace object recognition and knowledge graphs.

3. Resource Connections: Enterprises in aerospace, electricity, energy, etc.

04

Team Introduction

Wang Shuo, Ph.D. in Science from Peking University in 2020, published 8 top academic papers in electronic design automation and computer architecture during his doctoral studies, and joined Professor Shu Jiwu’s team at Tsinghua University for postdoctoral research in 2020, during which he designed and developed a memory-compute integration system based on flash memory and FPGA.

Shu Jiwu, Ph.D. in Computer Science from Nanjing University, tenured professor at Tsinghua University, dean of the School of Information at Xiamen University (dual appointment), doctoral supervisor, Fellow of the IEEE, Fellow of the CCF, recipient of the 2020 Huawei Olympus Award, the first prize of the 2020 CCF Science and Technology Award, and the first prize of the 2019 Ministry of Education Science and Technology Invention Award.

Lu Youyou, Ph.D. from the Department of Computer Science at Tsinghua University, associate professor at Tsinghua University, recipient of the CCF Excellent Doctoral Dissertation Award, Tsinghua University Excellent Postdoctoral Award, ACM China Operating Systems Committee Rising Star Award, with over 30 papers accepted at top computer systems conferences in the CCFA category, and winner of the Best Paper Award at NVMSA 2014 and Best Paper Nomination at MSST 2015.

05

Contact Information

Contact Person: Liu

E-mail: [email protected]

Result Number: 2021070

Note: Please indicate the source when reprinting.

Release of Results: Integrated Memory-Compute Graph Computing Platform Based on Flash Memory and FPGA

Tsinghua University Technology Transfer

Innovative Achievements

Leading Technology

Serving Society

Leave a Comment