xSPI IP: Flexible High-Performance Serial Interface Design, Implementation, and Validation of an In-House IP Core Supporting SPI, DSPI, QSPI, OSPI, and Mixed Operating Modes

xSPI IP: Flexible High-Performance Serial Interface Design, Implementation, and Validation of an In-House IP Core Supporting SPI, DSPI, QSPI, OSPI, and Mixed Operating Modes


Introduction

The Serial Peripheral Interface (SPI) is a widely used data transfer protocol. With the growing demand for higher transfer speeds, newer versions of the SPI protocol—such as Dual SPI (DSPI) and Quad SPI (QSPI)—have been introduced to enable significantly faster data transmission. To meet these demands, we have developed a fully custom xSPI IP solution compliant with the Expanded Serial Peripheral Interface (xSPI) for Nonvolatile Memory Devices Standard (JESD251C).

Challenges

Developing the xSPI IP involves addressing several key challenges. It must be resource-efficient, user-friendly, and easily configurable, while also supporting a wide range of operating modes. In addition to standard SPI, the design supports Dual, Quad, and Octal SPI (OSPI) interfaces. Although the JESD251C specification formally defines only SPI and OSPI as valid operating modes, real-world implementations often deviate from this standard, making Dual and Quad SPI support equally important. In some cases, there is also a need to support mixed operating modes, which adds another layer of complexity. Ultimately, the product is intended to serve as a direct replacement for existing SPI, DSPI, and QSPI solutions.

Possible Approaches

  • Memory-mapped implementation
  • Support for JESD251C (SPI and OSPI operating modes)
  • Support for SPI, DSPI, QSPI, and OSPI operating modes
  • Support for SPI, DSPI, QSPI, OSPI, and mixed operating modes

The following sections elaborate on each approach and provide a comparison, which is summarized in the table below.

 Memory MappedJESD251CJESD251C + DSPI and QSPI JESD251C + DSPI and QSPI + Mixed operating modes
Pros Direct memory access without requiring knowledge of SPI transactions, resulting in shorter development time and faster time to market. Allows the use of additional SPI commands defined by the JESD251C standard, but only in SPI and OSPI operating modes. Enables the use of additional SPI commands defined by the JESD251C standard across SPI, DSPI, QSPI, and OSPI operating modes. Enables the use of additional SPI commands defined by the JESD251C standard in SPI, DSPI, QSPI, and OSPI operating modes, as well as commands that utilize multiple operating modes within a single memory access.
Cons SPI devices are not always fully compatible with one another, which may require adapting the solution. Requires knowledge of SPI transactions. Does not support DSPI and QSPI operating modes. Does not support commands that utilize multiple operating modes within a single memory access. Requires knowledge of SPI transactions. Does not support commands that utilize multiple operating modes within a single memory access. Requires knowledge of SPI transactions. Significantly longer development time.

Table 1. Overview of advantages and limitations of the possible approaches

Solution

The chosen solution supports SPI, DSPI, QSPI, OSPI, and mixed operating modes. The primary challenge was implementing all these modes while minimizing resource consumption. An additional challenge was developing a complete IP solution entirely with in-house resources, without requiring any additional licenses. A mechanism was implemented to allow users to easily submit commands and data to the IP. The communication format between the IP and the user is kept simple, while preserving the structure of the command transactions. The solution diagram is shown in Figure 1.

Figure 1. xSPI IP block diagram

Figure 1 shows the xSPI IP block diagram with all its interfaces. The Advanced eXtensible Interface (AXI) Lite slave interface is used for accessing xSPI IP status and configuration registers. AXI Stream master and slave interfaces handle data transfer between the host and the xSPI IP. The xSPI master interface manages data transfer between the xSPI IP and SPI slave devices. Internal buffers are included on the AXI Stream master and slave interfaces, and their sizes can be adjusted by the user if needed.

Results and Future Improvements

The hardware implementation of the xSPI IP was performed on a CRUVI CR00107 Base Board equipped with an AMD Spartan-7 Field Programmable Gate Array (FPGA) and EVERSPIN EMxxLX Spin-Transfer Torque Magnetoresistive Random-Access Memory (STT-MRAM).

Figure 2. CRUVI CR00107 Base Board with the EVERSPIN MRAM Daughter board

MRAM leverages magnetic states rather than electrical charges to store data, enabling the production of non-volatile memory (NVM) chips that operate at speeds typically associated with RAM. This combines the persistent storage of flash memory with the speed and resilience of RAM. It also eliminates the need for backup power sources such as batteries or capacitors, making MRAM particularly suitable for devices that must function reliably in critical environments.

MRAM’s standard interfaces, both parallel and serial, allow for seamless integration, offering low-latency storage and retrieval. Everspin has recently introduced the xSPI PERSYST MRAM product family, based on the Expanded Serial Peripheral Interface (xSPI), the latest JEDEC standard for non-volatile memory devices. Built on Everspin's industrial STT-MRAM technology, these products offer high performance, multiple I/O support, SPI compatibility, and a high-speed, low-pin-count SPI-compatible bus interface with clock frequencies up to 200 MHz.

These persistent MRAM devices operate on a single 1.8 V power supply and deliver up to 400 MBps for both reads and writes via eight I/O signals. This advancement ushers in a new era of universal memory solutions, replacing products such as SRAM, BBSRAM, NVSRAM, and NOR devices, and targeting Industrial Automation, Process Control, Emulation, Automotive and Transportation, Gaming, and the broader Industrial Internet of Things (IIoT) markets.

The xSPI IP implementation was carried out using the AMD Vivado development tool. An AMD MicroBlaze soft-core Central Processing Unit (CPU) was used to configure the xSPI IP. Bare-metal testing was performed using the Vitis Integrated Design Environment (IDE). Table 2 shows the achieved read/write speeds for each operating mode on the implemented solution.

 SPIDual SPIQuad SPIOctal SPI
Data RateSingleSingleSingleDoubleSingleDouble
Bandwidth (Mbps)Up to 50Up to 100Up to 200Up to 400Up to 400Up to 800

Table 2. Read/write bandwidth of AMD Spartan-7 with EVERSPIN EMxxLX MRAM

Due to the speed limitations of the AMD Spartan-7, only 800 Mbps bandwidth was achieved using a 50 MHz xSPI clock. Higher bandwidth can be attained by porting the IP to a faster FPGA device.

The xSPI IP has been designed to allow straightforward adaptation to other FPGA development platforms. Future improvements include porting the xSPI IP to other AMD FPGA families as well as to devices from other vendors, such as Lattice and Altera.

Currently, the xSPI IP supports only a single device. Future versions could enable support for multiple xSPI slave devices.

To ensure robustness and correctness, the Universal Verification Methodology (UVM) was used for verification in a simulation environment.

Conclusion

The xSPI IP core delivers a highly flexible direct-replacement solution, supporting multiple operating modes including SPI, DSPI, QSPI, OSPI, and mixed modes to ensure broad compatibility with serial peripheral devices. Designed entirely in-house, the IP is robust, highly configurable, and resource-efficient, eliminating the need for third-party IP licenses. Its implementation on an AMD Spartan-7 FPGA with a MicroBlaze CPU and successful testing with Everspin EMxxLX STT-MRAM confirmed high-speed performance and straightforward integration. Planned future enhancements will further increase adaptability, making the xSPI IP suitable for a wide range of market applications.

With the new runner on the board, we executed the model again to generate the performance dump:

executor_runner --model_path 
/sharefs/mv2.pte --inputs /sharefs/dog_input.bin--etdump_path /sharefs/model.etdump

Once the run finished, we pulled the model.etdump file back to our host PC for analysis:

scp root@BoardsIP:/sharefs/model.etdump.

To analyze the data, we utilized ExecuTorch's Inspector APIs, which provide a clean interface for parsing ETRecord and ETDump files. By using Inspector.to_dataframe, we generated an Excel spreadsheet detailing all recorded events, their execution calls, and their exact runtimes.

However, to map these events back to the original Python source code, specifically capturing exact ATen operator names and stack_traces - we needed to generate an ETRecord file during the initial model export phase. This links back profiling details to the original Python source code (including stack traces and module hierarchy).

To implement this by following the official ETRecord Documentation, we created an updated export script, modifying the original section in export.py from:

prog = export_to_exec_prog(
    model,
    example_inputs,
    dynamic_shapes=dynamic_shapes,
    backend_config=backend_config,
    strict=args.strict,
)

...to the following implementation:

m = model.eval()
m = export(m, example_inputs, strict=True).module()

core_aten_ep = _to_core_aten(
    m,
    example_inputs,
    strict=args.strict,
)

edge_manager = _core_aten_to_edge(
    core_aten_ep,
    edge_compile_config=EdgeCompileConfig(_check_ir_validity=False),
)

edge_manager_copy = copy.deepcopy(edge_manager)
prog = edge_manager.to_executorch(config=backend_config)
generate_etrecord("mv2.etrecord", edge_manager_copy, prog)

We then ran this modified export script to generate both the .pte model and its corresponding mv2.etrecord file:

(.venv) ubuntu@ubuntu:~/executorch$ python3 -m examples.portable.scripts.exportEtRecord --model_name="mv2"
(.venv) ubuntu@ubuntu:~/executorch$ ls -la mv2.etrecord mv2.pte
-rw-r--r-- 1 user nisusers 15509467 May 6 11:52 mv2.etrecord
-rw-r--r-- 1 user nisusers 14233120 May 6 11:52 mv2.pte
(.venv) ubuntu@ubuntu:~/executorch$ scp -v mv2.pte root@BoardsIP:/sharefs

After repeating the inference on the board and pulling the new model.etdump, we loaded both files into the Inspector API:

inspector = Inspector(etdump_path="/path_to/model.etdump", etrecord="/path_to/mv2.etrecord")
df = inspector.to_dataframe()
df.to_csv("data.csv")

The resulting table included full ATen operator names, source stack traces, and module hierarchies. Reviewing this dataframe clearly showed that aten.convolution.default was our slowest operator.

Optimization via RISC-V Vector (RVV) Intrinsics

Our next task was to locate and optimize the underlying source function behind native_call_convolution.out. A thorough search through the ExecuTorch codebase pointed us to the default portable convolution kernel located at ~/executorch/kernels/portable/cpu/op_convolution.cpp.

To accelerate this, we rewrote the intensive parts of the kernel using RISC-V vector intrinsics, creating a new implementation file at /executorch/kernels/portable/cpu/op_convolutionRVV.cpp. We also added a custom .yaml configuration file, according to Kernel Registration Documentation, in /executorch/kernels/portable to register our new kernel:

- op: convolution.out
  kernels:
    - arg_meta: null
      kernel_name: torch::executor::convolutionRVV_out

To guarantee that the build system picked up our optimized kernel instead of the default fallback, we modified ~/praksa/executorch/kernels/portable/CMakeLists.txt to merge our custom configurations:

set(_my_yaml "${CMAKE_CURRENT_SOURCE_DIR}/my_functions.yaml")
set(_yaml "${CMAKE_CURRENT_SOURCE_DIR}/functions.yaml")

merge_yaml(
  FUNCTIONS_YAML ${_my_yaml}
  FALLBACK_YAML ${_yaml}
  OUTPUT_DIR ${CMAKE_CURRENT_BINARY_DIR}
)

gen_selected_ops(
  LIB_NAME "portable_ops_lib"
  OPS_SCHEMA_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

generate_bindings_for_kernels(
  LIB_NAME "portable_ops_lib"
  FUNCTIONS_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

After rebuilding the runtime, we confirmed that the mappings were correctly bound to aten::convolution.out by checking the generated code files: RegisterCodegenUnboxedKernelsEverything.cpp and NativeFunctions.h inside the build directory.

Performance and Benchmark Comparisons

With the optimizations complete, we transferred the newly compiled executable back to the board and ran a direct benchmark.

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003407 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.086602 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.094620 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.101219 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.115040 executorch:executor_runner.cpp:467] Model loaded in 99.634370 ms.
I 00:00:34.384195 executorch:executor_runner.cpp:525] Iteration 1 of 1: 34261.195795 ms
I 00:00:34.391929 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 34261.195795 ms.
I 00:00:34.401600 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:34.408598 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

Optimized: RISC-V Vector (RVV) Kernel

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003408 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.085513 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.093531 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.100130 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.113959 executorch:executor_runner.cpp:467] Model loaded in 98.929963 ms.
I 00:00:02.965595 executorch:executor_runner.cpp:525] Iteration 1 of 1: 2843.675211 ms
I 00:00:02.973243 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 2843.675211 ms.
I 00:00:02.982827 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:02.989899 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

It is important to mention that we used a vector multiplier of LMUL = m4 for our RISC-V vector intrinsic functions. We selected LMUL = m4 because it delivered the best performance on the CanMV-K230 board during testing, where the command:

executor_runner --model_path /sharefs/mv2.pte --inputs
/sharefs/dog_input.bin --etdump_path /sharefs/model.etdump

was executed multiple times using an automated script.

When analyzing the raw 1000-element output tensor, we noticed slight numerical differences between the unoptimized version and the vector version beginning at the 7th or 8th decimal place.

The benchmarks confirm that our RVV-optimized kernel runs about 10 times faster than the default, unoptimized executor_runner.

Figure 1. Measurement results: baseline (unoptimized) kernel vs. optimized RVV kernel

Dataset Accuracy Evaluation (ImageNet Validation)

To ensure that the minor numerical deviations in the 7th and 8th decimal places did not degrade model performance, we decided to run an accuracy evaluation using the full ImageNet validation dataset.

We downloaded the ImageNet validation subset from Kaggle: ImageNet Mini 1000 Dataset on Kaggle

Resizing the Storage Partition

After converting the validation images into raw formats, we attempted to copy the dataset onto the board's /sharefs folder. However, we quickly hit storage limits. We wrote an automation script to loop executor_runner through all raw images, but it regularly crashed due to a lack of disk space.

To resolve this issue, we extended the storage partition hosting /sharefs by following the instructions from the Kendryte K230 FAQ Guide:

[root@canaan /sharefs ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     44.0K     51.7M   0% /run
/dev/mmcblk1p4         255.9M    198.1M     57.8M  77% /sharefs

[root@canaan ~ ]#parted -l /dev/mmcblk1
Warning: Not all of the space available to /dev/mmcblk1 appears to be used...
Fix/Ignore? fix
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End     Size    File system  Name        Flags
1      10.5MB  31.5MB  21.0MB               rtt
2      31.5MB  83.9MB  52.4MB               linux
3      134MB   268MB   134MB   ext4         rootfs
4      268MB   537MB   268MB   fat16        fat32appfs  msftdata

[root@canaan ~ ]#umount /sharefs/
[root@canaan ~ ]#parted -a minimal /dev/mmcblk1 resizepart 4 8.5GB
[root@canaan ~ ]#parted -l /dev/mmcblk1
[root@canaan ~ ]#mkfs.ext2 /dev/mmcblk1p4
[root@canaan ~ ]#parted -l /dev/mmcblk1
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End      Size     File system  Name        Flags
1      10.5MB  31.5MB   21.0MB               rtt
2      31.5MB  83.9MB   52.4MB               linux
3      134MB   268MB    134MB    ext4         rootfs
4      268MB   8500MB   8232MB   ext2         fat32appfs  msftdata

[root@canaan ~ ]#mount /dev/mmcblk1p4 /sharefs/

[root@canaan ~ ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     48.0K     51.6M   0% /run
/dev/mmcblk1p4          7.5G     17.3M      7.1G   0% /sharefs

Final Accuracy Results

Because the dataset's subdirectories were named using standard ImageNet synset IDs (e.g., n01440764), we used a reference file named LOC_synset_mapping.txt to map these IDs to human-readable names. For instance, the entry n01440764 tench, Tinca tinca maps the folder ID to class index 1, representing a "tench" fish.

We calculated the Top-1 and Top-5 accuracy metrics, for both runners, across the entire validation dataset for both runners. The evaluations confirmed that the minor float precision variations from vector calculations caused no change in classification accuracy.

Both setups gave identical evaluation results:

Unoptimized Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Optimized RVV Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Our results align with the official PyTorch MobileNetV2 Model Documentation, which reports:

  • Top-1 Accuracy:878% (~71.9%)
  • Top-5 Accuracy:286% (~90.3%)

Dušan Stojković

You may also like