The rapid growth of online news and social media has increased the importance of multimodal sentiment analysis that jointly leverages textual and visual information. However, existing studies have mainly focused on social media, online reviews, and video data, and have not sufficiently addressed modality-level sentiment inconsistency, image noise, and uncertainty in weakly supervised labels in news articles. To address these issues, this study constructs CNN-MMSA, a weakly supervised multimodal news sentiment dataset consisting of 8,273 CNN news samples, and proposes a PMIG-Net-based sentiment analysis framework. PMIG-Net utilizes BERT-based text representations and CLIP-based image representations, selectively adjusts the contribution of visual information through prior-guided attention fusion and image gating, and jointly learns modality-specific and final multimodal representations through hierarchical multi-task learning. On CNN-MMSA, PMIG-Net achieved an Accuracy of 0.7164, a Macro-F1 of 0.6732, and a Macro-Recall of 0.6828. Compared with adapted baselines that reimplement the core fusion concepts of TFN, MulT, MAG-BERT, and Self-MM, the proposed model showed relatively stronger performance, particularly in Macro-F1 and Macro-Recall. Ablation experiments showed clear decreases in Macro-F1 and Macro-Recall when the image gate was removed, while the other components exhibited different trade-offs across evaluation metrics. In an additional benchmark evaluation on MVSA-Single, PMIG-Single-Pass also outperformed CLIP-only, BERT-only, and BERT+CLIP Concat under a controlled experimental setting, indicating its applicability to image–text sentiment analysis in social media data. By integrating weak-label construction and selective multimodal fusion within a unified framework, this study provides a practical multimodal sentiment analysis approach for news opinion analysis and intelligent media content analysis.