Android in IVI Unlocking the full potential of the In-vehicle experience

Android in IVI Unlocking the full potential of the In-vehicle experience


As the end-consumer requires both safety and convenience, there can be no real progress in autonomous driving without gaining consumers´ trust. To gain the trust, automakers are making enormous efforts to ensure safety throughout a vehicle’s lifecycle. But with safety being an irrefutable argument for any modern car, it’s worth asking whether it’s enough.

If safety comes first, what comes next?

The reality is that times have changed, and end users now expect much more from OEMs. When you can rely on the integrity of a car, the focus is on entertainment, which has a direct impact on the competitiveness of OEMs. To remain competitive, OEMs must provide consumers with an environment similar to the one on their phones. This inevitably brings Android into play as the most widely used entertainment platform.

Read on to learn more about the need for in-vehicle infotainment (IVI) in modern cars and why Android, among other solutions, is clearly winning the competition.

The boundary between driver and passenger is becoming blurred

As a vehicle becomes more automated, the driver’s free time also increases, and with it the need for high-quality content. Our driver is gradually becoming a passenger who no longer drives but tries to make the time spent on the journey as useful (and pleasant) as possible.

Future “drivers” will be able to type, watch videos, or even sleep while their car takes them to their desired destination.

The need for high-quality content equals the need for familiar services, which means bringing apps and services from other devices into the car.

Indeed, if it is natural to switch from a laptop to a phone that is familiar in appearance and user experience, then the ability to switch from one of these devices to IVI should make no difference. This only extends the usability of these services to one of the most common means of transportation in modern life – the personal car.

The article from Snapp Automotive confirms this, “Car manufacturers are waking up to the reality that their systems are now being compared to the other devices their customers are using like iPads and smartphones. A 5-year-old custom Linux system without any connection to the rest of the world just won’t do it anymore.”

Video streaming is one of the most widely used entertainment services today. Users want to be able to access it anywhere, which inevitably includes transportation and, in most cases, a personal car.

Gaming at home, gaming in the park on cell phones, gaming with friends, in person and virtually. Gaming has become one of the most popular forms of daily entertainment. At some point, it might even become a criterion when bying a car.

Sharing thoughts through funny pictures, emoticons, gifs and videos is a very engaging way of communication. Chat services have become an almost mandatory tool for quick and easy communication.

Most frequently used entertainment services

“We expect a continuum between our devices. The smartphone revolution has ingrained this expectation to us. You expect your mobile browser to know your passwords for all the sites you use on your laptop as well as to know your browsing history and bookmarks. This precedence set by smartphones should be taken seriously in the automotive industry. There’s no reason to believe that the same won’t happen there.”, Snapp Automotive insightfully concludes.
Clearly, the profound impact of the mobile industry on the automotive industry has set new standards and poses certain challenges for car manufacturers.

Differentiation lies in personalization and fun cars

“Ultimately, end consumers will seek applications that make driving more convenient and a seamless element of their daily routines and lifestyles.” McKinsey & Company

The McKinsey report1 analyzes the key drivers of the global automotive industry and explores the next horizons for differentiating features and services, particularly in the area of consumer-friendly active safety and infotainment solutions. It states: “Delivering services through the car – internet radio, smartphone capabilities, information/entertainment services, driver-assistance apps, tourism information, and the like – is a promising area for future profits and differentiation. So is the creation of new technical features for safe, comfortable, and eventually, autonomous driving. “
There are different opportunities for OEMs to make their cars desirable to the end consumer. They can benefit tremendously from various in-vehicle services. However, OEMs should be able to do that without compromising their brand identity. As HP Jin, global CEO of Telenav, said for the Reuters report2 /ROKU Automotive Global Report, “The auto industry needs new business models for the digital age. The real solution is not CarPlay and Android Auto. The car needs to have native software with deep integration of embedded software and connected services, along with frequent updates of software and the UX. But the cost of that is a hurdle for consumers.”
OEMs need to provide this high level of personalization and rich entertainment experience for the end consumer. Android, the most widely used operating system (OS), which supports numerous entertaining applications and services, is the right way to achieve this.

Why Android?

Due to the rich application ecosystem, the need for IVI consolidation, and the overall high market demand for Android compatibility, Android is being adopted by OEMs at a fast pace.
According to one of ABI Research3 ‘s recent white papers, 36 million new vehicles will be shipped with Android infotainment systems in 2030. This makes it difficult to be neutral about the introduction of this widely used operating system in the automotive industry.

Challenges of Android in IVI

The question is what the most efficient way is to open up to the Android offering. Most OEMs already have an IVI system in place, usually based on Linux. For them, it is convenient to be compliant with Android services. However, an immediate “switch” to Android OS brings some challenges for OEMs.

The evolution in the automotive industry brings many challenges for OEMs

Losing their unique identity due to a different user interface (UI) and the corresponding user experience (UX) of the Android platform is one of the significant drawbacks when looking at the long-term perspective. OEMs have built their brand and recognizable image over years and are reasonably cautious about compromising the signature of their IVI system. As many OEMs are not ready to make this move overnight, they need to find the formula that solves the win-win equation – to stay competitive while remaining authentic in this transition.

Get Android on board with RT-RK

As a solution to this equation, RT-RK provides the Android Automotive Services that enable OEMs to leverage this versatile and widely used platform

To achieve this, we offer both System Integration Service and Consulting Service, based on more than 10 years of experience and proven track record. RT-RK’s scalable System Integration Services enable Android Automotive to be tailored, optimized, matured, and maintained on different hardware platforms. In addition, customers can set up software development processes for Android Automotive by relying on RT-RK’s extensive expertise.

At RT-RK, we integrate your Android system either natively, with hypervisor, through a container or through an external HW device (cartridge concept). We can also scale your internal teams to integrate Android at the system level, taking into account the connections to the hardware and connecting the platform to ADAS.

Less dependency on Tier-1s

With less reliance on Tier1s and shorter time to market, OEMs can easily leverage Android’s rich offering and update their native infotainment system with the latest Android version available.

Beside flexibility, our service offers a wide range of benefits to our customers, including risk mitigation, third-party component integration, lifecycle management, scalable environment setup, and more.

With years of experience with Android systems as well as safety and real-time expertise, RT-RK is the customers can trust. At RT-RK, we not only enable safe and secure system functionality, but also modern Android-based infotainment systems to provide our customers with a complete set of features for their new generation vehicle.

With the new runner on the board, we executed the model again to generate the performance dump:

executor_runner --model_path 
/sharefs/mv2.pte --inputs /sharefs/dog_input.bin--etdump_path /sharefs/model.etdump

Once the run finished, we pulled the model.etdump file back to our host PC for analysis:

scp root@BoardsIP:/sharefs/model.etdump.

To analyze the data, we utilized ExecuTorch's Inspector APIs, which provide a clean interface for parsing ETRecord and ETDump files. By using Inspector.to_dataframe, we generated an Excel spreadsheet detailing all recorded events, their execution calls, and their exact runtimes.

However, to map these events back to the original Python source code, specifically capturing exact ATen operator names and stack_traces - we needed to generate an ETRecord file during the initial model export phase. This links back profiling details to the original Python source code (including stack traces and module hierarchy).

To implement this by following the official ETRecord Documentation, we created an updated export script, modifying the original section in export.py from:

prog = export_to_exec_prog(
    model,
    example_inputs,
    dynamic_shapes=dynamic_shapes,
    backend_config=backend_config,
    strict=args.strict,
)

...to the following implementation:

m = model.eval()
m = export(m, example_inputs, strict=True).module()

core_aten_ep = _to_core_aten(
    m,
    example_inputs,
    strict=args.strict,
)

edge_manager = _core_aten_to_edge(
    core_aten_ep,
    edge_compile_config=EdgeCompileConfig(_check_ir_validity=False),
)

edge_manager_copy = copy.deepcopy(edge_manager)
prog = edge_manager.to_executorch(config=backend_config)
generate_etrecord("mv2.etrecord", edge_manager_copy, prog)

We then ran this modified export script to generate both the .pte model and its corresponding mv2.etrecord file:

(.venv) ubuntu@ubuntu:~/executorch$ python3 -m examples.portable.scripts.exportEtRecord --model_name="mv2"
(.venv) ubuntu@ubuntu:~/executorch$ ls -la mv2.etrecord mv2.pte
-rw-r--r-- 1 user nisusers 15509467 May 6 11:52 mv2.etrecord
-rw-r--r-- 1 user nisusers 14233120 May 6 11:52 mv2.pte
(.venv) ubuntu@ubuntu:~/executorch$ scp -v mv2.pte root@BoardsIP:/sharefs

After repeating the inference on the board and pulling the new model.etdump, we loaded both files into the Inspector API:

inspector = Inspector(etdump_path="/path_to/model.etdump", etrecord="/path_to/mv2.etrecord")
df = inspector.to_dataframe()
df.to_csv("data.csv")

The resulting table included full ATen operator names, source stack traces, and module hierarchies. Reviewing this dataframe clearly showed that aten.convolution.default was our slowest operator.

Optimization via RISC-V Vector (RVV) Intrinsics

Our next task was to locate and optimize the underlying source function behind native_call_convolution.out. A thorough search through the ExecuTorch codebase pointed us to the default portable convolution kernel located at ~/executorch/kernels/portable/cpu/op_convolution.cpp.

To accelerate this, we rewrote the intensive parts of the kernel using RISC-V vector intrinsics, creating a new implementation file at /executorch/kernels/portable/cpu/op_convolutionRVV.cpp. We also added a custom .yaml configuration file, according to Kernel Registration Documentation, in /executorch/kernels/portable to register our new kernel:

- op: convolution.out
  kernels:
    - arg_meta: null
      kernel_name: torch::executor::convolutionRVV_out

To guarantee that the build system picked up our optimized kernel instead of the default fallback, we modified ~/praksa/executorch/kernels/portable/CMakeLists.txt to merge our custom configurations:

set(_my_yaml "${CMAKE_CURRENT_SOURCE_DIR}/my_functions.yaml")
set(_yaml "${CMAKE_CURRENT_SOURCE_DIR}/functions.yaml")

merge_yaml(
  FUNCTIONS_YAML ${_my_yaml}
  FALLBACK_YAML ${_yaml}
  OUTPUT_DIR ${CMAKE_CURRENT_BINARY_DIR}
)

gen_selected_ops(
  LIB_NAME "portable_ops_lib"
  OPS_SCHEMA_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

generate_bindings_for_kernels(
  LIB_NAME "portable_ops_lib"
  FUNCTIONS_YAML "${CMAKE_CURRENT_BINARY_DIR}/merged.yaml"
)

After rebuilding the runtime, we confirmed that the mappings were correctly bound to aten::convolution.out by checking the generated code files: RegisterCodegenUnboxedKernelsEverything.cpp and NativeFunctions.h inside the build directory.

Performance and Benchmark Comparisons

With the optimizations complete, we transferred the newly compiled executable back to the board and ran a direct benchmark.

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003407 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.086602 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.094620 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.101219 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.115040 executorch:executor_runner.cpp:467] Model loaded in 99.634370 ms.
I 00:00:34.384195 executorch:executor_runner.cpp:525] Iteration 1 of 1: 34261.195795 ms
I 00:00:34.391929 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 34261.195795 ms.
I 00:00:34.401600 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:34.408598 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

Optimized: RISC-V Vector (RVV) Kernel

msh >executor_runner -model_path /sharefs/mv2.pte -inputs /sharefs/dog_input.bin -etdump_path /sharefs/model.etdump -print_output "none"
--- ExecuTorch Start ---
I 00:00:00.003408 executorch:executor_runner.cpp:276] Loading inputs from input file(s).
I 00:00:00.085513 executorch:executor_runner.cpp:375] Model file /sharefs/mv2.pte is loaded.
I 00:00:00.093531 executorch:executor_runner.cpp:385] Using method forward
I 00:00:00.100130 executorch:executor_runner.cpp:436] Setting up planned buffer 0, size 9936896.
I 00:00:00.113959 executorch:executor_runner.cpp:467] Model loaded in 98.929963 ms.
I 00:00:02.965595 executorch:executor_runner.cpp:525] Iteration 1 of 1: 2843.675211 ms
I 00:00:02.973243 executorch:executor_runner.cpp:535] Model executed successfully 1 time(s) in 2843.675211 ms.
I 00:00:02.982827 executorch:executor_runner.cpp:544] 1 outputs:
I 00:00:02.989899 executorch:executor_runner.cpp:157] ETDump written to file '/sharefs/model.etdump'.

It is important to mention that we used a vector multiplier of LMUL = m4 for our RISC-V vector intrinsic functions. We selected LMUL = m4 because it delivered the best performance on the CanMV-K230 board during testing, where the command:

executor_runner --model_path /sharefs/mv2.pte --inputs
/sharefs/dog_input.bin --etdump_path /sharefs/model.etdump

was executed multiple times using an automated script.

When analyzing the raw 1000-element output tensor, we noticed slight numerical differences between the unoptimized version and the vector version beginning at the 7th or 8th decimal place.

The benchmarks confirm that our RVV-optimized kernel runs about 10 times faster than the default, unoptimized executor_runner.

Figure 1. Measurement results: baseline (unoptimized) kernel vs. optimized RVV kernel

Dataset Accuracy Evaluation (ImageNet Validation)

To ensure that the minor numerical deviations in the 7th and 8th decimal places did not degrade model performance, we decided to run an accuracy evaluation using the full ImageNet validation dataset.

We downloaded the ImageNet validation subset from Kaggle: ImageNet Mini 1000 Dataset on Kaggle

Resizing the Storage Partition

After converting the validation images into raw formats, we attempted to copy the dataset onto the board's /sharefs folder. However, we quickly hit storage limits. We wrote an automation script to loop executor_runner through all raw images, but it regularly crashed due to a lack of disk space.

To resolve this issue, we extended the storage partition hosting /sharefs by following the instructions from the Kendryte K230 FAQ Guide:

[root@canaan /sharefs ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     44.0K     51.7M   0% /run
/dev/mmcblk1p4         255.9M    198.1M     57.8M  77% /sharefs

[root@canaan ~ ]#parted -l /dev/mmcblk1
Warning: Not all of the space available to /dev/mmcblk1 appears to be used...
Fix/Ignore? fix
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End     Size    File system  Name        Flags
1      10.5MB  31.5MB  21.0MB               rtt
2      31.5MB  83.9MB  52.4MB               linux
3      134MB   268MB   134MB   ext4         rootfs
4      268MB   537MB   268MB   fat16        fat32appfs  msftdata

[root@canaan ~ ]#umount /sharefs/
[root@canaan ~ ]#parted -a minimal /dev/mmcblk1 resizepart 4 8.5GB
[root@canaan ~ ]#parted -l /dev/mmcblk1
[root@canaan ~ ]#mkfs.ext2 /dev/mmcblk1p4
[root@canaan ~ ]#parted -l /dev/mmcblk1
Model: SD SL32G (sd/mmc)
Disk /dev/mmcblk1: 31.9GB
Sector size (logical/physical): 512B/512B
Partition Table: gpt

Number  Start   End      Size     File system  Name        Flags
1      10.5MB  31.5MB   21.0MB               rtt
2      31.5MB  83.9MB   52.4MB               linux
3      134MB   268MB    134MB    ext4         rootfs
4      268MB   8500MB   8232MB   ext2         fat32appfs  msftdata

[root@canaan ~ ]#mount /dev/mmcblk1p4 /sharefs/

[root@canaan ~ ]#df -h
Filesystem                Size      Used      Available  Use% Mounted on
/dev/root              118.5M     86.4M     28.2M  75% /
devtmpfs                13.0M         0     13.0M   0% /dev
tmpfs                   51.7M         0     51.7M   0% /dev/shm
tmpfs                   51.7M     52.0K     51.6M   0% /tmp
tmpfs                   51.7M     48.0K     51.6M   0% /run
/dev/mmcblk1p4          7.5G     17.3M      7.1G   0% /sharefs

Final Accuracy Results

Because the dataset's subdirectories were named using standard ImageNet synset IDs (e.g., n01440764), we used a reference file named LOC_synset_mapping.txt to map these IDs to human-readable names. For instance, the entry n01440764 tench, Tinca tinca maps the folder ID to class index 1, representing a "tench" fish.

We calculated the Top-1 and Top-5 accuracy metrics, for both runners, across the entire validation dataset for both runners. The evaluations confirmed that the minor float precision variations from vector calculations caused no change in classification accuracy.

Both setups gave identical evaluation results:

Unoptimized Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Optimized RVV Kernel Metrics:

==============================
OVERALL TOP-1 ACCURACY
==============================
2790/3923 (71.12%)

==============================
OVERALL TOP-5 ACCURACY
==============================
3535/3923 (90.11%)

Our results align with the official PyTorch MobileNetV2 Model Documentation, which reports:

  • Top-1 Accuracy:878% (~71.9%)
  • Top-5 Accuracy:286% (~90.3%)

Dušan Stojković

You may also like