Pil-Seong Jeong
| 2026, 31(7)
| pp.1~11
| number of Cited : 0
This study compares modality-specific specialized models distributed across a heterogeneous edge cluster with single-device omnimodal deployment under identical tasks. A cluster of Jetson AGX Orin, Raspberry Pi 5+Hailo-8, and Orange Pi 5 Plus runs MobileViT-S, Whisper-tiny, and DistilBERT via gRPC, benchmarked against LLaVA-1.5-7B FP16, its INT4 quantization, and Qwen2-VL-2B FP16. The distributed configuration is 80–288× faster on Vision, 3.7–25× faster on Text, and up to 806× more energy-efficient, with Text F1 within 2.8 percentage points across all four configurations. Effect decomposition attributes 3.93× to quantization and 1.75× to model size reduction, with an additional 63.5× from distribution itself, confirming independence from model size and quantization. The 17–195× cost-normalized advantage is retained across five wired-to-wireless network scenarios; distributed clusters thus suit concurrent multi-modal serving, while omnimodal deployment suits cross-modal reasoning.