Accelerating ExecuTorch with RISC-V Vector Extensions on CanMV-K230

Accelerating ExecuTorch with RISC-V Vector Extensions on CanMV-K230


In the last blog post we talked about how to successfully load and execute the example addition model on the RT-Smart core of CanMV-K230-V1.1 board. Following the completion of these preparation steps, now that the ExecuTorch runtime is successfully running on the RT-Smart core, it is time to implement the ATen operators. However, before adding dedicated RVV kernels to ExecuTorch, it is necessary to define the model that will be used for both testing and optimization. While there is a wide variety of models available, because this task focuses on enabling ExecuTorch exclusively on the CPU, our choices for testing models are limited. Among the few models selected for evaluation, we will be focusing on MobileNet-V2.

Exporting MobileNetV2 to ExecuTorch (.pte)

MobileNet V2 improves mobile performance through a highly efficient architecture. It uses inverted residual blocks and linear bottlenecks to start with a smaller representation of the data, expands it for processing, and shrinks it again to reduce the number of computations. To preserve accuracy, despite its simplified design, the model also removes non-linearities. Also, it retains the depthwise separable convolutions introduced in MobileNetV1 to ensure efficiency.

Our first step was to export MobileNetV2 into the .pte format, as this binary file is the standard format consumed by the ExecuTorch runtime.

To achieve this, we followed the ExecuTorch in Portable Mode guide (README.md) found in the official Executorch repository. After completing the initial setup described in the Setting up ExecuTorch section, we used the provided export script to generate our model binary.

The portable/scripts/export.py script allowed us to generate a model binary file by choosing a target model from the models directory. We ran the following commands:

cd executorch # Navigate to the top-level directory
# Display the list of available example models
python3 -m examples.portable.scripts.export -h
# Generate the specific .pte binary for MobileNetV2
python3 -m examples.portable.scripts.export --model_name="mv2"

We initially encountered some dependency issues. However, after successfully installing torchvision-0.26.0+cpu and ensuring the rest of the environment was properly configured, we made a few adjustments and everything ran smoothly. Successful execution of the export command generated our mv2.pte file.

Running Inference on the Board

Next, the .jpg image was preprocessed (resized, cropped, and normalized) and saved as a raw binary file of float32 values and executed it on the target board using executor_runner:

msh /sharefs>executor_runner --model_path /sharefs/mv2.pte --inputs /sharefs/cat_input.bin

To ensure the image was exported to the exact raw format expected by our model, we followed the preprocessing guidelines from the official PyTorch MobileNetV2 Documentation.

Running executor_runner with the model and the .bin image input produced an output tensor containing 1000 elements. We then extracted the top 5 highest values and mapped their indexes against the official ImageNet labels from this PyTorch GitHub file. The classification results matched our expectations.

Profiling and Generating ETRecord Files

To identify performance bottlenecks and see which ATen operators were the slowest, we needed to profile the model following the PyTorch Documentation. This required rebuilding executor_runner with profiling enabled so it could generate an .etdump file.

First, we modified the k230rsmart.cmake configuration file to enable the event tracer and developer tools:


set(EXECUTORCH_ENABLE_EVENT_TRACER       ON CACHE BOOL "" FORCE)
set(EXECUTORCH_BUILD_DEVTOOLS           ON CACHE BOOL "" FORCE)

We then compiled the updated executor_runner and transferred it to the board using the following commands:


cd executorch
rm -rf cmake-outCanMV
cmake -B cmake-outCanMV   -DCMAKE_TOOLCHAIN_FILE= ~/path_to/executorch/k230_rtsmart.cmake   -DCMAKE_BUILD_TYPE=Release   -DEXECUTORCH_BUILD_EXECUTOR_RUNNER=ON -DEXECUTORCH_ENABLE_LOGGING=ON  -DPYTHON_EXECUTABLE=python3
cmake --build cmake-outCanMV --target executor_runner -j8
    
cd cmake-outCanMV/
scp -v executor_runner root@BoardsIP:/sharefs



With the new runner on the board, we executed the model again to generate the performance dump:


executor_runner --model_path /sharefs/mv2.pte --inputs 
/sharefs/dog_input.bin --etdump_path /sharefs/model.etdump

Once the run finished, we pulled the model.etdump file back to our host PC for analysis:


scp root@BoardsIP:/sharefs/model.etdump .

To analyze the data, we utilized ExecuTorch's Inspector APIs, which provide a clean interface for parsing ETRecord and ETDump files. By using Inspector.to_dataframe, we generated an Excel spreadsheet detailing all recorded events, their execution calls, and their exact runtimes.

However, to map these events back to the original Python source code, specifically capturing exact ATen operator names and stack_traces - we needed to generate an ETRecord file during the initial model export phase. This links back profiling details to the original Python source code (including stack traces and module hierarchy).

To implement this by following the official ETRecord Documentation, we created an updated export script, modifying the original section in export.py from:


prog = export_to_exec_prog(
        model,
        example_inputs,
        dynamic_shapes=dynamic_shapes,
        backend_config=backend_config,
        strict=args.strict,
    )


...to the following implementation:


m = model.eval()
m = export(m, example_inputs, strict=True).module()
    
core_aten_ep = _to_core_aten(
    m,
    example_inputs,
    strict=args.strict,
)
    
edge_manager = _core_aten_to_edge(
    core_aten_ep,
    edge_compile_config=EdgeCompileConfig(_check_ir_validity=False),
)
    
edge_manager_copy = copy.deepcopy(edge_manager)
prog = edge_manager.to_executorch(config=backend_config)
generate_etrecord("mv2.etrecord", edge_manager_copy, prog)



We then ran this modified export script to generate both the .pte model and its corresponding mv2.etrecord file:


(.venv) ubuntu@ubuntu:~/executorch$ python3 -m examples.portable.scripts.exportEtRecord --model_name="mv2"
(.venv) ubuntu@ubuntu:~/executorch$ ls -la mv2.etrecord mv2.pte
-rw-r--r-- 1 user nisusers 15509467 May  6 11:52 mv2.etrecord
-rw-r--r-- 1 user nisusers 14233120 May  6 11:52 mv2.pte
(.venv) ubuntu@ubuntu:~/executorch$ scp -v mv2.pte root@BoardsIP:/sharefs



After repeating the inference on the board and pulling the new model.etdump, we loaded both files into the Inspector API:


inspector = Inspector(etdump_path="/path_to/model.etdump", etrecord="/path_to/mv2.etrecord")
df = inspector.to_dataframe()
df.to_csv("data.csv")


The resulting table included full ATen operator names, source stack traces, and module hierarchies. Reviewing this dataframe clearly showed that aten.convolution.default was our slowest operator.

Optimization via RISC-V Vector (RVV) Intrinsics

Our next task was to locate and optimize the underlying source function behind native_call_convolution.out. A thorough search through the ExecuTorch codebase pointed us to the default portable convolution kernel located at ~/executorch/kernels/portable/cpu/op_convolution.cpp.

To accelerate this, we rewrote the intensive parts of the kernel using RISC-V vector intrinsics, creating a new implementation file at /executorch/kernels/portable/cpu/op_convolutionRVV.cpp. We also added a custom .yaml configuration file, according to Kernel Registration Documentation, in /executorch/kernels/portable to register our new kernel:


- op: convolution.out
  kernels:
    - arg_meta: null
      kernel_name: torch::executor::convolutionRVV_out


To guarantee that the build system picked up our optimized kernel instead of the default fallback, we modified ~/praksa/executorch/kernels/portable/CMakeLists.txt to merge our custom configurations:


set(_my_yaml "${CMAKE_CURRENT_SOURCE_DIR}/my_functions.yaml")
set(_yaml "${CMAKE_CURRENT_SOURCE_DIR}/functions.yaml")

merge_yaml(
  FUNCTIONS_YAML ${_my_yaml}
  FALLBACK_YAML  ${_yaml}
  OUTPUT_DIR     ${CMAKE_CURRENT_BINARY_DIR}
)

gen_selected_ops(
  LIB_NAME "portable_ops_lib"
  OPS_SCHEMA_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

generate_bindings_for_kernels(
  LIB_NAME "portable_ops_lib"
  FUNCTIONS_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)



After rebuilding the runtime, we confirmed that the mappings were correctly bound to aten::convolution.out by checking the generated code files: RegisterCodegenUnboxedKernelsEverything.cpp and NativeFunctions.h inside the build directory.

Performance and Benchmark Comparisons

With the optimizations complete, we transferred the newly compiled executable back to the board and ran a direct benchmark.

Baseline: Unoptimized Portable Kernel


msh /sharefs>executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003407 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.086602 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.094620 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.101219 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.115040 executorch:executor_runner.cpp:467] Model loaded in 99.634370 ms.
I 00:00:34.384195 executorch:executor_runner.cpp:525] Iteration 1 of 1: 34261.195795 ms
I 00:00:34.391929 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 34261.195795 ms.
I 00:00:34.401600 executorch:executor_runner.cpp:544] 1 outputs: 
I 00:00:34.408598 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.


Optimized: RISC-V Vector (RVV) Kernel


msh /sharefs>executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003408 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.085513 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.093531 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.100130 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.113959 executorch:executor_runner.cpp:467] Model loaded in 98.929963 ms.
I 00:00:02.965595 executorch:executor_runner.cpp:525] Iteration 1 of 1: 2843.675211 ms
I 00:00:02.973243 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 2843.675211 ms.
I 00:00:02.982827 executorch:executor_runner.cpp:544] 1 outputs: 
I 00:00:02.989899 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.


It is important to mention that we used a vector multiplier of LMUL = m4 for our RISC-V vector intrinsic functions. We selected LMUL = m4 because it delivered the best performance on the CanMV-K230 board during testing, where the command:


executor_runner --model_path /sharefs/mv2.pte --inputs 
/sharefs/dog_input.bin --etdump_path /sharefs/model.etdump

was executed multiple times using an automated script.

When analyzing the raw 1000-element output tensor, we noticed slight numerical differences between the unoptimized version and the vector version beginning at the 7th or 8th decimal place.

The benchmarks confirm that our RVV-optimized kernel runs about 10 times faster than the default, unoptimized executor_runner.

Measurement results: baseline (unoptimized) kernel vs. optimized RVV kernel

Figure 1. Measurement results: baseline (unoptimized) kernel vs. optimized RVV kernel

Dataset Accuracy Evaluation (ImageNet Validation)

To ensure that the minor numerical deviations in the 7th and 8th decimal places did not degrade model performance, we decided to run an accuracy evaluation using the full ImageNet validation dataset.

We downloaded the ImageNet validation subset from Kaggle: ImageNet Mini 1000 Dataset on Kaggle

Resizing the Storage Partition

After converting the validation images into raw formats, we attempted to copy the dataset onto the board's /sharefs folder. However, we quickly hit storage limits. We wrote an automation script to loop executor_runner through all raw images, but it regularly crashed due to a lack of disk space.

To resolve this issue, we extended the storage partition hosting /sharefs by following the instructions from the Kendryte K230 FAQ Guide:


[root@canaan /sharefs ]#df -h
Filesystem                Size      Used Available Use% Mounted on
/dev/root               118.5M     86.4M     28.2M  75% /
devtmpfs                 13.0M         0     13.0M   0% /dev
tmpfs                    51.7M         0     51.7M   0% /dev/shm
tmpfs                    51.7M     52.0K     51.6M   0% /tmp
tmpfs                    51.7M     44.0K     51.7M   0% /run
/dev/mmcblk1p4          255.9M    198.1M     57.8M  77% /sharefs

[root@canaan ~ ]#parted   -l /dev/mmcblk1
Warning: Not all of the space available to /dev/mmcblk1 appears to be used, you
can fix the GPT to use all of the space (an extra 61285343 blocks) or continue
with the current setting? 
Fix/Ignore? fix                                                            
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt
Disk Flags: 

Number  Start   End     Size    File system  Name        Flags
 1      10.5MB  31.5MB  21.0MB               rtt
 2      31.5MB  83.9MB  52.4MB               linux
 3      134MB   268MB   134MB   ext4         rootfs
 4      268MB   537MB   268MB   fat16        fat32appfs  msftdata
 
[root@canaan ~ ]#umount /sharefs/
[root@canaan ~ ]#parted  -a minimal  /dev/mmcblk1  resizepart 4  8.5GB
[root@canaan ~ ]#parted   -l /dev/mmcblk1
[root@canaan ~ ]#mkfs.ext2 /dev/mmcblk1p4
[root@canaan ~ ]#parted   -l /dev/mmcblk1
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt
Disk Flags: 

Number  Start   End     Size    File system  Name        Flags
 1      10.5MB  31.5MB  21.0MB               rtt
 2      31.5MB  83.9MB  52.4MB               linux
 3      134MB   268MB   134MB   ext4         rootfs
 4      268MB   8500MB  8232MB  ext2         fat32appfs  msftdata

[root@canaan ~ ]#mount /dev/mmcblk1p4 /sharefs/

[root@canaan ~ ]#df -h
Filesystem                Size      Used Available Use% Mounted on
/dev/root               118.5M     86.4M     28.2M  75% /
devtmpfs                 13.0M         0     13.0M   0% /dev
tmpfs                    51.7M         0     51.7M   0% /dev/shm
tmpfs                    51.7M     52.0K     51.6M   0% /tmp
tmpfs                    51.7M     48.0K     51.6M   0% /run
/dev/mmcblk1p4            7.5G     17.3M      7.1G   0% /sharefs


Final Accuracy Results

Because the dataset's subdirectories were named using standard ImageNet synset IDs (e.g., n01440764), we used a reference file named LOC_synset_mapping.txt to map these IDs to human-readable names. For instance, the entry n01440764 tench, Tinca tinca maps the folder ID to class index 1, representing a "tench" fish.

We calculated the Top-1 and Top-5 accuracy metrics, for both runners, across the entire validation dataset for both runners. The evaluations confirmed that the minor float precision variations from vector calculations caused no change in classification accuracy.

Both setups gave identical evaluation results:

Unoptimized Kernel Metrics:


==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)


Optimized RVV Kernel Metrics:


==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)



Our results align with the official PyTorch MobileNetV2 Model Documentation, which reports:

  • Top-1 Accuracy:878% (~71.9%)
  • Top-5 Accuracy:286% (~90.3%)

Dušan Stojković

You may also like