TFLite(LiteRT) - Generative AI on Genio 720

This page lists the generative AI models and representative performance data for Genio 720 platforms. For background information about generative workloads and usage notes, refer to TFLite(LiteRT) - Generative AI.

Model Support and Performance

Note

Performance numbers apply to both Android and Yocto (this platform supports both OS). For OS-level GAI support across all platforms, see AI Supporting Scope.

The following symbols are used in the tables below:

  • -- : To be released.

  • Q3/E : Support planned for Q3 end, 2026.

  • Q4/E : Support planned for Q4 end, 2026.

  • V : Supported and verified on this platform.

  • X : Platform does not support this model.

Large Language Models (LLMs)

Performance Comparison (Prompt/Generative) (Unit: tok/s)

Model

Genio 720

Qwen3-0.6B

535.94 / 22.93

Qwen3-1.7B

262.73 / 13.84

Qwen3-4B

125.58 / 7.25

Qwen3-8B

80.17 / 4.76

Qwen2.5-1.5B-Instruct

337.06 / 19.25

Qwen2.5-3B-Instruct

163.23 / 10.64

Qwen2.5-7B-Instruct

69.47 / 4.73

gemma3-1B (Text Only)

583.01 / 26.44

gemma3-4B (Text-Only)

176.79 / 5.93

llama3.2-1B-Instruct

400.57 / 24.92

llama3.2-3B-Instruct

153.56 / 10.36

Phi-3-mini-4k-instruct

127.56 / 7.28

Phi-3.5-mini-instruct

136.63 / 6.29

Vision-Language Models (VLMs)

Performance Comparison (ViT Inference Time / Prompt / Generative)

Model

ViT Inference Time (s)

Prompt Mode (tok/s)

Generative Mode (tok/s)

Qwen3VL-2B

0.43

199.75

14.4286

Speech Recognition Models (ASR)

Performance Data

Model

Genio 720

Whisper-base (8w16a)

160.6 ms / 174.8 tok/s

Qwen3-ASR-0.6B

127.6 / 7.1 tok/s

Moonshine-Tiny

V

Measured latency for these models is reported in asr_cmdline_tool, together with the transcripts each package’s sample clip is expected to produce.

Android-only Models

The following model categories are supported on Android only and are not available on Yocto.

Note

Yocto support for these model categories is planned for 2026/E (estimated).

Stable Diffusion and Image Generation

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

Genio 720

Stable Diffusion v2.1 base model with controlnet

35461 / 30365

Stable Diffusion v.1.5 controlnet

33870 / 32763

CLIP and Embedding Models

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

Genio 720

img_encoder_proj_clip_vit_large_dynamic

473.82 / 309.29

img_encoder_proj_openclip_vit_h_dynamic

1331.19 / 1026.39

text_encoder_clip_vit_large

297.77 / 42.24

text_encoder_openclip_vit_h

607.12 / 125.44