TFLite(LiteRT) - Generative AI on MT8883

This page lists the generative AI models and representative performance data for MT8883 platforms. For background information about generative workloads and usage notes, refer to TFLite(LiteRT) - Generative AI.

Model Support and Performance

Note

For OS-level GAI support across all platforms, see AI Supporting Scope.

The following symbols are used in the tables below:

  • -- : To be released.

  • Q3/E : Support planned for Q3 (estimated).

  • X : Platform does not support this model.

Large Language Models (LLMs)

Performance Comparison (Prompt/Generative) (Unit: tok/s)

Model

MT8883

Qwen3-0.6B

Qwen3-1.7B

834.02 / 25.18

Qwen3-4B

Qwen3-8B

Qwen2.5-1.5B-Instruct

763.66 / 20.16

Qwen2.5-3B-Instruct

502.05 / 19.32

Qwen2.5-7B-Instruct

184.73 / 6.89

gemma3-1B (Text Only)

1125.16 / 38.87

gemma3-4B (Text-Only)

gemma2-2b-it

llama3.2-1B-Instruct

1372.57 / 43.47

llama3.2-3B-Instruct

611.76 / 19.89

llama3-8b

128.36 / 6.53

MiniCPM-2B-sft-bf16-llama-format

Phi-3-mini-4k-instruct

Phi-3.5-mini-instruct

DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek-R1-Distill-Qwen-7B

DeepSeek-R1-Distill-Llama-8B

Android-only Runtime Models

The following models use Android-only speculative decoding runtime.

Performance Comparison (Prompt/Generative) (Unit: tok/s)

Model

MT8883

llava1.5-7b-speculative-decoding

138.76 / 4.46

medusa_v1_0_vicuna_7b_v1.5

vicuna1.5-7b-tree-speculative-decoding-plus

baichuan-7b-int8-cache

Vision-Language Models (VLMs)

Performance Comparison (ViT Inference Time / Prompt / Generative)

Model

ViT Inference Time (s)

Prompt Mode (tok/s)

Generative Mode (tok/s)

Qwen3VL-2B

InternVL3-1B

Stable Diffusion and Image Generation

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

MT8883

Stable Diffusion v2.1 base model with controlnet

Stable Diffusion v.1.5 controlnet

CLIP and Embedding Models

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

MT8883

img_encoder_proj_clip_vit_large_dynamic

img_encoder_proj_openclip_vit_big_g_dynamic

img_encoder_proj_openclip_vit_h_dynamic

text_encoder_clip_vit_large

text_encoder_openclip_vit_h

Speech Recognition Models

Support Status

Model

MT8883

Whisper

Q3/E