TFLite(LiteRT) - Generative AI on MT8893
This page lists the generative AI models and representative performance data for MT8893 platforms. For background information about generative workloads and usage notes, refer to TFLite(LiteRT) - Generative AI.
Model Support and Performance
Note
For OS-level GAI support across all platforms, see AI Supporting Scope.
The following symbols are used in the tables below:
--: To be released.Q3/E: Support planned for Q3 (estimated).X: Platform does not support this model.
Large Language Models (LLMs)
Model |
MT8893 |
|---|---|
Qwen3-0.6B |
– |
Qwen3-1.7B |
1069.16 / 23.42 |
Qwen3-4B |
– |
Qwen3-8B |
– |
Qwen2.5-1.5B-Instruct |
1621.85 / 38.57 |
Qwen2.5-3B-Instruct |
751.06 / 20.87 |
Qwen2.5-7B-Instruct |
471.95 / 11.74 |
gemma3-1B (Text Only) |
– |
gemma3-4B (Text-Only) |
– |
llama3.2-1B-Instruct |
2093.61 / 61.14 |
llama3.2-3B-Instruct |
1022.95 / 25.05 |
llama3-8b |
426.13 / 11.51 |
MiniCPM-2B-sft-bf16-llama-format |
886.72 / 22.28 |
Phi-3-mini-4k-instruct |
– |
Phi-3.5-mini-instruct |
– |
DeepSeek-R1-Distill-Qwen-1.5B |
1057.25 / 25.68 |
DeepSeek-R1-Distill-Qwen-7B |
448.17 / 11.69 |
DeepSeek-R1-Distill-Llama-8B |
425.79 / 11.36 |
Android-only Models
The following LLM models are not supported on Yocto.
Model |
MT8893 |
|---|---|
llava1.5-7b-speculative-decoding |
267.98 / 6.78 |
medusa_v1_0_vicuna_7b_v1.5 |
501.05 / 22.79 |
vicuna1.5-7b-tree-speculative-decoding-plus |
454.58 / 22.72 |
Vision-Language Models (VLMs)
Model |
ViT Inference Time (s) |
Prompt Mode (tok/s) |
Generative Mode (tok/s) |
|---|---|---|---|
Qwen3VL-2B |
– |
– |
– |
InternVL3-1B |
0.51 |
183.64 |
14.09 |
Stable Diffusion and Image Generation
Model |
MT8893 |
|---|---|
Stable Diffusion v2.1 base model with controlnet |
6969 / 5451 |
Stable Diffusion v.1.5 controlnet |
9395 / 8035 |
CLIP and Embedding Models
Model |
MT8893 |
|---|---|
img_encoder_proj_clip_vit_large_dynamic |
358.61 / 51.14 |
img_encoder_proj_openclip_vit_big_g_dynamic |
1390.56 / 517.13 |
img_encoder_proj_openclip_vit_h_dynamic |
591.93 / 147.47 |
text_encoder_clip_vit_large |
308.72 / 18.94 |
text_encoder_openclip_vit_h |
510.92 / 48.49 |
Speech Recognition Models
Model |
MT8893 |
|---|---|
Whisper |
Q3/E |