Performance

Performance Notes and Limitations

For Generative AI workloads, measured performance on Genio 520 may be slightly lower than on Genio 720.

This gap is primarily due to DRAM bandwidth differences between the two platforms and might affect:

  • Token generation speed for LLMs.

  • End-to-end latency for diffusion-based image generation.

  • Multimodal pipelines that exchange large intermediate tensors between subsystems.

Important

The tables in this section provide representative numbers only. To obtain the most accurate performance for a specific use case, developers must deploy and run the workload directly on the target platform under the intended system configuration.

Note

The following symbols are used in the performance tables below:

  • -- : To be released.

  • Q3/E : Support planned for Q3 (estimated).

  • X : Platform does not support this model.

Unless otherwise noted, all models listed below are supported on both Android and Yocto OS. Models or sections explicitly marked Android-only are not available on Yocto.

LLM Performance Comparison
Performance Comparison (Prompt Mode / Generative Mode) (Unit: tok/s)

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

Qwen3-0.6B

318.66 / 14.22

420.38 / 21.44

526.43 / 23.04

535.94 / 22.93

Qwen3-1.7B

152.67 / 8.69

222.83 / 13.03

258.83 / 13.72

262.73 / 13.84

235.67 / 6.20

834.02 / 25.18

1069.16 / 23.42

Qwen3-4B

73.21 / 3.53

105.18 / 6.87

126.12 / 7.30

125.58 / 7.25

Qwen3-8B

X

65.04 / 4.58

79.95 / 4.77

80.17 / 4.76

Qwen2.5-1.5B-Instruct

199.70 / 12.21

294.22 / 17.89

340.29 / 18.99

337.06 / 19.25

406.94 / 17.13

763.66 / 20.16

1621.85 / 38.57

Qwen2.5-3B-Instruct

94.62 / 6.71

132.30 / 9.98

161.19 / 10.38

163.23 / 10.64

220.32 / 9.60

502.05 / 19.32

751.06 / 20.87

Qwen2.5-7B-Instruct

X

57.79 / 4.61

69.71 / 4.86

69.47 / 4.73

83.07 / 4.13

184.73 / 6.89

471.95 / 11.74

gemma3-1B (Text Only)

359.43 / 16.58

513.69 / 24.58

598.27 / 27.23

583.01 / 26.44

680.41 / 21.02

1125.16 / 38.87

gemma3-4B (Text-Only)

92.76 / 3.44

149.00 / 5.69

176.90 / 5.90

176.79 / 5.93

llama3.2-1B-Instruct

233.97 / 15.94

329.32 / 21.58

385.77 / 24.76

400.57 / 24.92

1372.57 / 43.47

2093.61 / 61.14

llama3.2-3B-Instruct

X

120.17 / 9.92

153.84 / 10.26

153.56 / 10.36

611.76 / 19.89

1022.95 / 25.05

llama3-8b

X

49.23 / 4.39

56.47 / 4.66

55.87 / 4.64

128.36 / 6.53

426.13 / 11.51

MiniCPM-2B-sft-bf16-llama-format

X

168.67 / 5.82

153.14 / 6.48

194.79 / 7.69

886.72 / 22.28

Phi-3-mini-4k-instruct

74.79 / 4.45

101.94 / 6.30

126.82 / 7.26

127.56 / 7.28

Phi-3.5-mini-instruct

77.82 / 3.27

111.01 / 5.30

136.84 / 6.09

136.63 / 6.29

DeepSeek-R1-Distill-Qwen-1.5B

183.94 / 7.60

298.68 / 11.19

341.88 / 11.70

331.12 / 11.62

1057.25 / 25.68

DeepSeek-R1-Distill-Qwen-7B

X

56.47 / 4.60

67.80 / 4.84

67.59 / 4.83

448.17 / 11.69

DeepSeek-R1-Distill-Llama-8B

X

X

X

36.65 / 4.58

425.79 / 11.36

VLM Performance Comparison
Performance Comparison (ViT Inference Time (s) / Prompt (tok/s) / Generative (tok/s))

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

Qwen3VL-2B

0.66 / 120.62 / 13.2906

0.49 / 173.16 / 13.4803

0.42 / 124.25 / 13.7964

0.43 / 199.75 / 14.4286

InternVL3-1B

2.99 / 49.25 / 3.10

2.42 / 69.37 / 6.00

1.77 / 80.49 / 6.20

1.79 / 79.84 / 6.35

0.51 / 183.64 / 14.09

Speech Recognition Performance Comparison
Support Status

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

Whisper

Q3/E

Q3/E

Q3/E

Q3/E

Q3/E

Q3/E

Q3/E

Android-only Models

The following model categories are supported on Android only and are not available on Yocto.

LLM — Android-only Models
Performance Comparison (Prompt Mode / Generative Mode) (Unit: tok/s)

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

llava1.5-7b-speculative-decoding

X

58.62 / 2.99

73.12 / 3.40

73.11 / 3.40

138.76 / 4.46

267.98 / 6.78

medusa_v1_0_vicuna_7b_v1.5

X

X

X

91.82 / 10.56

501.05 / 22.79

vicuna1.5-7b-tree-speculative-decoding-plus

X

X

X

84.90 / 12.65

454.58 / 22.72

Stable Diffusion Performance Comparison
Performance Comparison (Main Time / Inference Time) (Unit: ms)

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

Stable Diffusion v2.1 base model with controlnet

216219 / 143784

39995 / 38425

35480 / 30421

35461 / 30365

Stable Diffusion v.1.5 controlnet

65449 / 53267

43507 / 42111

34483 / 33057

33870 / 32763

CLIP Performance Comparison
Performance Comparison (Main Time / Inference Time) (Unit: ms)

Model

Genio 360

Genio 420

Genio 520

Genio 720

MT8875

MT8883

MT8893

img_encoder_proj_clip_vit_large_dynamic

1362.52 / 435.94

619.67 / 421.55

488.69 / 314.31

473.82 / 309.29

358.61 / 51.14

img_encoder_proj_openclip_vit_big_g_dynamic

14997.09 / 5177.00

6371.25 / 5221.64

12595.05 / 4092.66

4870.33 / 3974.65

1390.56 / 517.13

img_encoder_proj_openclip_vit_h_dynamic

2440.64 / 1457.64

1760.07 / 1392.31

1376.49 / 1023.27

1331.19 / 1026.39

591.93 / 147.47

text_encoder_clip_vit_large

755.48 / 65.62

408.53 / 58.81

366.14 / 44.80

297.77 / 42.24

308.72 / 18.94

text_encoder_openclip_vit_h

1857.49 / 203.44

794.72 / 149.17

679.66 / 128.66

607.12 / 125.44

510.92 / 48.49

Per-platform Performance

Detailed per-model performance data for each platform: