본문 바로가기
  • Home

Multimodal sLLM Serving Optimization and Hardware-Aware Precision Tuning for Autonomous Cloud Fault Mitigation

  • Journal of Software Forensics
  • Abbr : JSF
  • 2026, 22(3), pp.187~197
  • Publisher : Korea Software Assessment and Valuation Society
  • Research Area : Engineering > Computer Science
  • Received : September 8, 2026
  • Accepted : September 20, 2026
  • Published : September 30, 2026

Hojung Lim 1

1한국전자기술연구원

Accredited

ABSTRACT

Cascading failures caused by non-linear dependencies in microservice clouds expose critical limitations in reducing MTTR through single-metric monitoring [3, 9]. To achieve autonomous cloud fault mitigation, this paper proposes: (1) an Adaptive Gating Cross-Attention alignment scheme integrating multimodal telemetry (metrics, logs, and traces); and (2) a hardware ISA-aware mixed-precision (W8A8/W4A8) lightweight serving framework for sub-3B sLLMs. Evaluated on 120 incident cases, our multimodal integration (M4) on Qwen2.5-Coder-3B achieves a 79.2% component localization rate and a 77.5% root-cause accuracy. Furthermore, on a commodity CPU, our approach resolves the dequantization latency inversion caused by uniform low-bit quantization (INT8: 264.40ms), stabilizing inference latency to 171.86ms (a 35.0% reduction). This provides a practical foundation to economically implement sub-second autonomous fault diagnosis and mitigation SLAs on-premises without security leakages.

Journal Copyright Policy

No CCL information provided

Citation status

* References for papers published after 2025 are currently being built.