Getting real data BBT HDMI 4K grabber, design and development

Getting real data BBT HDMI 4K grabber, design and development


Nowadays, the development of complex professional electronics systems is constantly expanding. The goal of every manufacturer is to put their product on the market as fast as possible and offer it to the customers – on the other hand, a short production time should not affect the quality and reliability of it. Modern devices have many functionalities, from a coupling of different interfaces to combining multiple devices into a single device. Examination and testing of such complex systems has become a serious problem. During the development of a complex professional electronics system, over 40% of the time goes on verification and testing.

As an answer to such challenges in TV and set-top box (STB) industry, RT-RK provided complete testing framework exploiting Black Box Testing (BBT) approach. The missing puzzle was a device capable to grab video outputs from TV chassis or STB video outputs, in real time, each frame or audio/video sequence, to be saved and analyzed. The device was based on Xilinx Zynq SoC and had numerous audio and video interfaces: (HDMI, (High-Definition Multimedia Interface), YPbPr, S-Video (Separate Video), S/PDIF (Sony/Philips Digital Interconnect Format) optical, S/PDIF coaxial, CVBS (Composite Video, Blanking, and Sync)). The device had the ability to connect to the network via 1Gb LAN (Local Area Network) port, and was provided in three different form factors: as a standalone device, PC built-in peripheria in 5.25” case (as a replacement in CD/DVD optic device compartment), or as a desktop PC PCIe card.

BBT HDMI 4K grabber, design and development

Project overview:

  • Internal R&D project, triggered by RT-RK Marketing and Sales department due to increased number of requests for testing of 4K capable equipment within BBT framework

Engagement of the HW department:

  • Complete product requirements and functional specifications
  • Full MECH, HW and FPGA design, certification and release to manufacture
  • Prototype series for 4k capable video grabbing device, not losing any of the functions used in previous grabber versions (RT-AV100, RT-AV110)
  • Final device integration in BBT framework and PC SW
BBT HDMI 4K grabber, design and development-Final device integration in BBT framework and PC SW

Features

  • Real-time capturing and streaming of audio and video signals up to 4K over 1Gbps LAN, 10Gbps LAN or PCIe, 4x 2nd Gen.
  • The device was intended for head end monitoring as well as for STB/DVD/Blu-ray testing
  • Live preview through LAN interface and Automatic Video Signal and Standard Detection
  • Grabber inputs: HDMI, CVBS, RGB, S-Video, analog audio, S/PDIF (optical and coaxial)
BBT HDMI 4K grabber

Conclusion

RTAV4K was a very good example of complete product development cycle, starting from the business idea and market (testing industry) needs, to defining the technology, design, development and market release. Even more chalenging, the product’s small scale production, maintenance, and service over the product lifetime, has accumulated additional important knowledge within the team.

With the new runner on the board, we executed the model again to generate the performance dump:

executor_runner --model_path 
/sharefs/mv2.pte --inputs /sharefs/dog_input.bin--etdump_path /sharefs/model.etdump

Once the run finished, we pulled the model.etdump file back to our host PC for analysis:

scp root@BoardsIP:/sharefs/model.etdump.

To analyze the data, we utilized ExecuTorch's Inspector APIs, which provide a clean interface for parsing ETRecord and ETDump files. By using Inspector.to_dataframe, we generated an Excel spreadsheet detailing all recorded events, their execution calls, and their exact runtimes.

However, to map these events back to the original Python source code, specifically capturing exact ATen operator names and stack_traces - we needed to generate an ETRecord file during the initial model export phase. This links back profiling details to the original Python source code (including stack traces and module hierarchy).

To implement this by following the official ETRecord Documentation, we created an updated export script, modifying the original section in export.py from:

prog = export_to_exec_prog(
    model,
    example_inputs,
    dynamic_shapes=dynamic_shapes,
    backend_config=backend_config,
    strict=args.strict,
)

...to the following implementation:

m = model.eval()
m = export(m, example_inputs, strict=True).module()

core_aten_ep = _to_core_aten(
    m,
    example_inputs,
    strict=args.strict,
)

edge_manager = _core_aten_to_edge(
    core_aten_ep,
    edge_compile_config=EdgeCompileConfig(_check_ir_validity=False),
)

edge_manager_copy = copy.deepcopy(edge_manager)
prog = edge_manager.to_executorch(config=backend_config)
generate_etrecord("mv2.etrecord", edge_manager_copy, prog)

We then ran this modified export script to generate both the .pte model and its corresponding mv2.etrecord file:

(.venv) ubuntu@ubuntu:~/executorch$ python3 -m examples.portable.scripts.exportEtRecord --model_name="mv2"
(.venv) ubuntu@ubuntu:~/executorch$ ls -la mv2.etrecord mv2.pte
-rw-r--r-- 1 user nisusers 15509467 May 6 11:52 mv2.etrecord
-rw-r--r-- 1 user nisusers 14233120 May 6 11:52 mv2.pte
(.venv) ubuntu@ubuntu:~/executorch$ scp -v mv2.pte root@BoardsIP:/sharefs

After repeating the inference on the board and pulling the new model.etdump, we loaded both files into the Inspector API:

inspector = Inspector(etdump_path="/path_to/model.etdump", etrecord="/path_to/mv2.etrecord")
df = inspector.to_dataframe()
df.to_csv("data.csv")

The resulting table included full ATen operator names, source stack traces, and module hierarchies. Reviewing this dataframe clearly showed that aten.convolution.default was our slowest operator.

Optimization via RISC-V Vector (RVV) Intrinsics

Our next task was to locate and optimize the underlying source function behind native_call_convolution.out. A thorough search through the ExecuTorch codebase pointed us to the default portable convolution kernel located at ~/executorch/kernels/portable/cpu/op_convolution.cpp.

To accelerate this, we rewrote the intensive parts of the kernel using RISC-V vector intrinsics, creating a new implementation file at /executorch/kernels/portable/cpu/op_convolutionRVV.cpp. We also added a custom .yaml configuration file, according to Kernel Registration Documentation, in /executorch/kernels/portable to register our new kernel:

- op: convolution.out
  kernels:
    - arg_meta: null
      kernel_name: torch::executor::convolutionRVV_out

To guarantee that the build system picked up our optimized kernel instead of the default fallback, we modified ~/praksa/executorch/kernels/portable/CMakeLists.txt to merge our custom configurations:

set(_my_yaml "${CMAKE_CURRENT_SOURCE_DIR}/my_functions.yaml")
set(_yaml "${CMAKE_CURRENT_SOURCE_DIR}/functions.yaml")

merge_yaml(
  FUNCTIONS_YAML ${_my_yaml}
  FALLBACK_YAML ${_yaml}
  OUTPUT_DIR ${CMAKE_CURRENT_BINARY_DIR}
)

gen_selected_ops(
  LIB_NAME "portable_ops_lib"
  OPS_SCHEMA_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

generate_bindings_for_kernels(
  LIB_NAME "portable_ops_lib"
  FUNCTIONS_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

After rebuilding the runtime, we confirmed that the mappings were correctly bound to aten::convolution.out by checking the generated code files: RegisterCodegenUnboxedKernelsEverything.cpp and NativeFunctions.h inside the build directory.

Performance and Benchmark Comparisons

With the optimizations complete, we transferred the newly compiled executable back to the board and ran a direct benchmark.

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003407 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.086602 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.094620 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.101219 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.115040 executorch:executor_runner.cpp:467] Model loaded in 99.634370 ms.
I 00:00:34.384195 executorch:executor_runner.cpp:525] Iteration 1 of 1: 34261.195795 ms
I 00:00:34.391929 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 34261.195795 ms.
I 00:00:34.401600 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:34.408598 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

Optimized: RISC-V Vector (RVV) Kernel

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003408 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.085513 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.093531 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.100130 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.113959 executorch:executor_runner.cpp:467] Model loaded in 98.929963 ms.
I 00:00:02.965595 executorch:executor_runner.cpp:525] Iteration 1 of 1: 2843.675211 ms
I 00:00:02.973243 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 2843.675211 ms.
I 00:00:02.982827 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:02.989899 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

It is important to mention that we used a vector multiplier of LMUL = m4 for our RISC-V vector intrinsic functions. We selected LMUL = m4 because it delivered the best performance on the CanMV-K230 board during testing, where the command:

executor_runner --model_path /sharefs/mv2.pte --inputs
/sharefs/dog_input.bin --etdump_path /sharefs/model.etdump

was executed multiple times using an automated script.

When analyzing the raw 1000-element output tensor, we noticed slight numerical differences between the unoptimized version and the vector version beginning at the 7th or 8th decimal place.

The benchmarks confirm that our RVV-optimized kernel runs about 10 times faster than the default, unoptimized executor_runner.

Figure 1. Measurement results: baseline (unoptimized) kernel vs. optimized RVV kernel

Dataset Accuracy Evaluation (ImageNet Validation)

To ensure that the minor numerical deviations in the 7th and 8th decimal places did not degrade model performance, we decided to run an accuracy evaluation using the full ImageNet validation dataset.

We downloaded the ImageNet validation subset from Kaggle: ImageNet Mini 1000 Dataset on Kaggle

Resizing the Storage Partition

After converting the validation images into raw formats, we attempted to copy the dataset onto the board's /sharefs folder. However, we quickly hit storage limits. We wrote an automation script to loop executor_runner through all raw images, but it regularly crashed due to a lack of disk space.

To resolve this issue, we extended the storage partition hosting /sharefs by following the instructions from the Kendryte K230 FAQ Guide:

[root@canaan /sharefs ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     44.0K     51.7M   0% /run
/dev/mmcblk1p4         255.9M    198.1M     57.8M  77% /sharefs

[root@canaan ~ ]#parted -l /dev/mmcblk1
Warning: Not all of the space available to /dev/mmcblk1 appears to be used...
Fix/Ignore? fix
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End     Size    File system  Name        Flags
1      10.5MB  31.5MB  21.0MB               rtt
2      31.5MB  83.9MB  52.4MB               linux
3      134MB   268MB   134MB   ext4         rootfs
4      268MB   537MB   268MB   fat16        fat32appfs  msftdata

[root@canaan ~ ]#umount /sharefs/
[root@canaan ~ ]#parted -a minimal /dev/mmcblk1 resizepart 4 8.5GB
[root@canaan ~ ]#parted -l /dev/mmcblk1
[root@canaan ~ ]#mkfs.ext2 /dev/mmcblk1p4
[root@canaan ~ ]#parted -l /dev/mmcblk1
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End      Size     File system  Name        Flags
1      10.5MB  31.5MB   21.0MB               rtt
2      31.5MB  83.9MB   52.4MB               linux
3      134MB   268MB    134MB    ext4         rootfs
4      268MB   8500MB   8232MB   ext2         fat32appfs  msftdata

[root@canaan ~ ]#mount /dev/mmcblk1p4 /sharefs/

[root@canaan ~ ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     48.0K     51.6M   0% /run
/dev/mmcblk1p4          7.5G     17.3M      7.1G   0% /sharefs

Final Accuracy Results

Because the dataset's subdirectories were named using standard ImageNet synset IDs (e.g., n01440764), we used a reference file named LOC_synset_mapping.txt to map these IDs to human-readable names. For instance, the entry n01440764 tench, Tinca tinca maps the folder ID to class index 1, representing a "tench" fish.

We calculated the Top-1 and Top-5 accuracy metrics, for both runners, across the entire validation dataset for both runners. The evaluations confirmed that the minor float precision variations from vector calculations caused no change in classification accuracy.

Both setups gave identical evaluation results:

Unoptimized Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Optimized RVV Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Our results align with the official PyTorch MobileNetV2 Model Documentation, which reports:

  • Top-1 Accuracy:878% (~71.9%)
  • Top-5 Accuracy:286% (~90.3%)

Dušan Stojković

You may also like