TFLite(LiteRT) - Generative AI on MT8875

This page lists the generative AI models and representative performance data for MT8875 platforms. For background information about generative workloads and usage notes, refer to TFLite(LiteRT) - Generative AI.

Model Support and Performance

Note

For OS-level GAI support across all platforms, see AI Supporting Scope.

The following symbols are used in the tables below:

  • -- : To be released.

  • Q3/E : Support planned for Q3 (estimated).

  • X : Platform does not support this model.

Large Language Models (LLMs)

Performance Comparison (Prompt/Generative) (Unit: tok/s)

Model

MT8875

Qwen3-0.6B

Qwen3-1.7B

235.67 / 6.20

Qwen3-4B

Qwen3-8B

Qwen2.5-1.5B-Instruct

406.94 / 17.13

Qwen2.5-3B-Instruct

220.32 / 9.60

Qwen2.5-7B-Instruct

83.07 / 4.13

gemma3-1B (Text Only)

680.41 / 21.02

gemma3-4B (Text-Only)

gemma2-2b-it

llama3.2-1B-Instruct

llama3.2-3B-Instruct

llama3-8b

MiniCPM-2B-sft-bf16-llama-format

Phi-3-mini-4k-instruct

Phi-3.5-mini-instruct

DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek-R1-Distill-Qwen-7B

DeepSeek-R1-Distill-Llama-8B

Android-only Runtime Models

The following models use Android-only speculative decoding runtime.

Performance Comparison (Prompt/Generative) (Unit: tok/s)

Model

MT8875

llava1.5-7b-speculative-decoding

medusa_v1_0_vicuna_7b_v1.5

vicuna1.5-7b-tree-speculative-decoding-plus

baichuan-7b-int8-cache

Vision-Language Models (VLMs)

Performance Comparison (ViT Inference Time / Prompt / Generative)

Model

ViT Inference Time (s)

Prompt Mode (tok/s)

Generative Mode (tok/s)

Qwen3VL-2B

InternVL3-1B

Stable Diffusion and Image Generation

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

MT8875

Stable Diffusion v2.1 base model with controlnet

Stable Diffusion v.1.5 controlnet

CLIP and Embedding Models

Performance Comparison (Main Time/Inference Time) (Unit: ms)

Model

MT8875

img_encoder_proj_clip_vit_large_dynamic

img_encoder_proj_openclip_vit_big_g_dynamic

img_encoder_proj_openclip_vit_h_dynamic

text_encoder_clip_vit_large

text_encoder_openclip_vit_h

Speech Recognition Models

Support Status

Model

MT8875

Whisper

Q3/E