Uncertainty-Aware Lightweight Breast Ultrasound Modeling on BreastMNIST+: Parameter-Efficient Label-Prompt Alignment, Controlled-Shift Robustness, and Verifiable Evidence Cards

Authors

Keywords:

Breast Ultrasound, Medmnist+, Medical Artificial Intelligence, Label-Prompt Alignment, Reduced-Rank Modeling, Probability Calibration, Conformal Prediction, Distribution Shift, Evidence Cards

Abstract

On BreastMNIST+ 224, we tested visual and class-prompt models with calibration, conformal sets, routing, perturbations, and evidence cards. The 546/78/156 train/validation/test artifact contained one conflicting duplicate. The ensemble achieved AUROC/AUPRC 0.885/0.738, balanced accuracy 0.842, sensitivity 0.833, and specificity 0.851. Temperature scaling reduced Brier/ECE 0.120/0.093-to-0.116/0.079. Label-conditional 90% conformal prediction achieved overall/malignant coverage 0.897/0.952 and size 1.359. Rank-2 prompts reached AUROC 0.790; blending raised sensitivity to 0.905 but lowered specificity to 0.763. Rotation and contrast degraded most. 156 cards passed checks. Results support routing, not clinical validity.

References

1] J. Mongan, L. Moy, and C. E. Kahn, Jr., "Checklist for artificial intelligence in medical imaging (CLAIM): A guide for authors and reviewers," Radiology: Artificial Intelligence, vol. 2, no. 2, art. e200029, 2020, doi: 10.1148/ryai.2020200029.

[2] G. Varoquaux and V. Cheplygina, "Machine learning for medical imaging: Methodological failures and recommendations for the future," npj Digital Medicine, vol. 5, art. 48, 2022, doi: 10.1038/s41746-022-00592-y.

[3] J. Yang, R. Shi, D. Wei, Z. Liu, L. Zhao, B. Ke, H. Pfister, and B. Ni, "MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification," Scientific Data, vol. 10, art. 41, 2023, doi: 10.1038/s41597-022-01721-8.

[4] J. Yang, R. Shi, D. Wei, Z. Liu, L. Zhao, B. Ke, H. Pfister, and B. Ni, "[MedMNIST+] 18x standardized datasets for 2D and 3D biomedical image classification with multiple size options: 28 (MNIST-Like), 64, 128, and 224," Zenodo, ver. 3.0, Jan. 2024, doi: 10.5281/zenodo.10519652.

[5] W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy, "Dataset of breast ultrasound images," Data in Brief, vol. 28, art. 104863, 2020, doi: 10.1016/j.dib.2019.104863.

[6] S. Doerrich, F. Di Salvo, J. Brockmann, and C. Ledig, "Rethinking model prototyping through the MedMNIST+ dataset collection," Scientific Reports, vol. 15, art. 7669, 2025, doi: 10.1038/s41598-025-92156-9.

[7] F. Di Salvo, S. Doerrich, and C. Ledig, "MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions," arXiv:2406.17536, 2024, doi: 10.48550/arXiv.2406.17536.

[8] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "MobileNetV2: Inverted residuals and linear bottlenecks," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2018, pp. 4510-4520, doi: 10.1109/CVPR.2018.00474.

[9] A. Dosovitskiy et al., "An image is worth 16 x 16 words: Transformers for image recognition at scale," in Proc. ICLR, 2021.

[10] G. Mi, T. Ye, and D. Wood, "A lightweight medical foundation model for cross-modal multi-task pretraining and parameter-efficient few-shot transfer on MedMNIST," J. Technol. Informatics Eng., vol. 4, no. 3, pp. 572-589, 2025, doi: 10.51903/jtie.v4i3.492.

[11] S. Lu and T. Zou, "Uncertainty-aware medical vision-language classification on a lightweight MedMNIST-compatible biomedical patch benchmark," J. Technol. Informatics Eng., vol. 5, no. 2, pp. 1-19, 2026, doi: 10.51903/jtie.v5i2.530.

[12] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, "A simple framework for contrastive learning of visual representations," in Proc. 37th Int. Conf. Machine Learning, vol. 119, 2020, pp. 1597-1607.

[13] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, "Emerging properties in self-supervised vision transformers," in Proc. IEEE/CVF Int. Conf. Computer Vision, 2021, pp. 9650-9660, doi: 10.1109/ICCV48922.2021.00951.

[14] Q. Wu, G. Mi, and D. Wood, "Calibration-light subject-independent motor imagery BCI via self-supervised pretraining and Conformer," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 239-262, 2025, doi: 10.51903/jtie.v5i1.493.

[15] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, "LoRA: Low-rank adaptation of large language models," in Proc. ICLR, 2022.

[16] N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, "Parameter-efficient transfer learning for NLP," in Proc. 36th Int. Conf. Machine Learning, vol. 97, 2019, pp. 2790-2799.

[17] A. Radford et al., "Learning transferable visual models from natural language supervision," in Proc. 38th Int. Conf. Machine Learning, vol. 139, 2021, pp. 8748-8763.

[18] Z. Wang, Z. Wu, D. Agarwal, and J. Sun, "MedCLIP: Contrastive learning from unpaired medical images and text," in Proc. EMNLP, 2022, pp. 3876-3887, doi: 10.18653/v1/2022.emnlp-main.256.

[19] S.-C. Huang, L. Shen, M. P. Lungren, and S. Yeung, "GLoRIA: A multimodal global-local representation learning framework for label-efficient medical image recognition," in Proc. IEEE/CVF Int. Conf. Computer Vision, 2021, pp. 3942-3951, doi: 10.1109/ICCV48922.2021.00391.

[20] K. Zhang et al., "A generalist vision-language foundation model for diverse biomedical tasks," Nature Medicine, vol. 30, no. 11, pp. 3129-3141, 2024, doi: 10.1038/s41591-024-03185-2.

[21] C. Kim, S. U. Gadgil, A. J. DeGrave, J. A. Omiye, Z. R. Cai, R. Daneshjou, and S.-I. Lee, "Transparent medical image AI via an image-text foundation model grounded in medical literature," Nature Medicine, vol. 30, pp. 1154-1165, 2024, doi: 10.1038/s41591-024-02887-x.

[22] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, "Cancer image classification based on DenseNet model," J. Phys.: Conf. Ser., vol. 1651, no. 1, art. 012143, 2020, doi: 10.1088/1742-6596/1651/1/012143.

[23] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, "Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma," Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, 2024, doi: 10.54097/zJ4MnbWW.

[24] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On calibration of modern neural networks," in Proc. 34th Int. Conf. Machine Learning, vol. 70, 2017, pp. 1321-1330.

[25] G. W. Brier, "Verification of forecasts expressed in terms of probability," Monthly Weather Review, vol. 78, no. 1, pp. 1-3, 1950, doi: 10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2.

[26] Q. Xin, "Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning)," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215-238, 2025, doi: 10.51903/jtie.v4i1.485.

[27] B. Lakshminarayanan, A. Pritzel, and C. Blundell, "Simple and scalable predictive uncertainty estimation using deep ensembles," in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 6402-6413.

[28] Y. Ovadia et al., "Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift," in Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 13991-14002.

[29] Z. S. Zhong, X. Pan, and Q. Lei, "Bridging domains with approximately shared features," in Proc. 28th Int. Conf. Artificial Intelligence and Statistics, PMLR, vol. 258, pp. 559-567, 2025.

[30] Y. Romano, M. Sesia, and E. J. Candès, "Classification with valid and adaptive coverage," in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 3581-3591.

[31] A. N. Angelopoulos, S. Bates, J. Malik, and M. I. Jordan, "Uncertainty sets for image classifiers using conformal prediction," in Proc. ICLR, 2021.

[32] Y. Geifman and R. El-Yaniv, "Selective classification for deep neural networks," in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4878-4887.

[33] J. Jin, T. Huang, and S. Lu, "A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI credit card default dataset," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 74-85, 2024, doi: 10.69987/JACS.2024.40606.

[34] Y. Chen, Y. Zhang, D. Chau, and M. Sherman, "Credit card default risk tiering with probability calibration and uncertainty-driven rejection: A reproducible study on the UCI credit card clients dataset," J. Adv. Comput. Syst., vol. 3, no. 4, pp. 31-47, 2023, doi: 10.69987/JACS.2023.30403.

[35] J. Jin, T. Huang, and S. Lu, "Cost-sensitive learning, simulated PU learning, and one-class autoencoding for extreme-imbalance credit card fraud detection," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 64-73, 2024, doi: 10.69987/JACS.2024.40605.

[36] D. Hendrycks and T. Dietterich, "Benchmarking neural network robustness to common corruptions and perturbations," in Proc. ICLR, 2019.

[37] N. Arun et al., "Assessing the trustworthiness of saliency maps for localizing abnormalities in medical imaging," Radiology: Artificial Intelligence, vol. 3, no. 6, art. e200267, 2021, doi: 10.1148/ryai.2021200267.

[38] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, "Model cards for model reporting," in Proc. Conf. Fairness, Accountability, and Transparency, 2019, pp. 220-229, doi: 10.1145/3287560.3287596.

[39] Z. S. Zhong, Q. Wu, and G. Mi, "Uncertainty-aware medical image explanation cards: LLM-generated visual explanations for AI-assisted radiology interfaces," Int. J. Graph. Des., vol. 3, no. 2, pp. 415-436, 2025, doi: 10.51903/ijgd.v3i2.3616.

[40] T. Ye, X. Chang, and E. Zhong, "Uncertainty-aware breast ultrasound explanation cards: A visual communication framework for image-based AI diagnostic support using BreastMNIST_224," Int. J. Graph. Des., vol. 3, no. 2, pp. 365-380, 2025, doi: 10.51903/ijgd.v3i2.3701.

[41] B. Zhou, C. Li, and L. Liu, "Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication," Int. J. Graph. Des., vol. 3, no. 2, pp. 365-380, 2025, doi: 10.51903/ijgd.v3i2.3696.

[42] C. Li, B. Zhou, and K. Gao, "Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces," Int. J. Graph. Des., vol. 3, no. 2, pp. 381-394, 2025, doi: 10.51903/ijgd.v3i2.3709.

[43] J. Jin, "LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy," Int. J. Graph. Des., vol. 3, no. 2, pp. 397-414, 2025, doi: 10.51903/ijgd.v3i2.3698.

[44] C. Li, J. Bai, and S. Wang, "Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications," J. Adv. Comput. Syst., vol. 4, no. 2, pp. 76-92, 2024, doi: 10.69987/JACS.2024.40207.

[45] Z. S. Zhong, J. Chen, E. Zhong, and X. Sun, "Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 502-520, 2025, doi: 10.51903/jtie.v4i2.536.

[46] T. Saito and M. Rehmsmeier, "The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets," PLoS ONE, vol. 10, no. 3, art. e0118432, 2015, doi: 10.1371/journal.pone.0118432.

[47] S. Lu and D. Zhou, "TinyLLM-assisted intrusion detection for real-time IoT networks," J. Adv. Comput. Syst., vol. 4, no. 8, pp. 72-87, 2024, doi: 10.69987/JACS.2024.40809.

[48] Q. Xin, "Hybrid cloud architecture for efficient and cost-effective large language model deployment," J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182-2195, 2025, doi: 10.51519/journalisi.v7i3.1170.

[49] S. He, C. Li, and H. Rao, "Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 306-324, 2025, doi: 10.51903/jtie.v4i1.546.

[50] J. Zhang, "From general human activity recognition to volleyball-oriented wearable transfer learning: Cross-dataset evidence from UCI HAR and WISDM for domain adaptation and edge deployment," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 263-283, 2025, doi: 10.51903/jtie.v4i1.524.

[51] S. He, X. Chang, and E. Sun, "Cross-cloud transfer learning for AI training capacity forecasting under workload and topology distribution shift," J. Adv. Comput. Syst., vol. 4, no. 1, pp. 100-120, 2024, doi: 10.69987/JACS.2024.40108.

[52] Q. Xin, "Probabilistic bike-sharing demand forecasting under changing weather and seasonal regimes with transformer-based models," Findings, Mar. 2026, doi: 10.32866/001c.157499.

[53] S. Chen, S. He, and E. Sun, "Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters," J. Adv. Comput. Syst., vol. 4, no. 5, pp. 119-134, 2024, doi: 10.69987/JACS.2024.40509.

[54] Q. Xin, "Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation," AVITEC, vol. 8, no. 2, pp. 247-264, 2026, doi: 10.28989/avitec.v8i2.3974.

[55] J. Jin, "Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 520-533, 2025, doi: 10.51903/jtie.v4i2.535.

[56] Q. Xin, "Explainable and fair credit risk scoring with counterfactual explanations: A reproducible evaluation on the German Credit dataset (HELOC-motivated)," J. Inf. Technol., vol. 14, no. 2, pp. 215-231, 2026, doi: 10.32664/j-intech.v14i02.2228.

[57] X. Sun, Z. S. Zhong, and Q. Wu, "Retrieval-grounded HDFS log anomaly detection and deterministic failure narrative generation," J. Comput. Syst. Appl., vol. 3, no. 1, pp. 15-30, 2026, doi: 10.64229/j6d7fr94.

[58] Q. Xin, "Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI Online Retail," J. Inf. Technol., vol. 14, no. 1, pp. 20-37, 2026, doi: 10.32664/j-intech.v14i01.2229.

[59] W. Su, H. Rao, and E. Ma, "Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks," Int. J. Graph. Des., vol. 4, no. 1, pp. 186-191, 2026, doi: 10.51903/ijgd.v4i1.3699.

[60] B. Zhang, H. Rao, and D. Zhao, "Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents," J. Adv. Comput. Syst., vol. 4, no. 3, pp. 109-125, 2024, doi: 10.69987/JACS.2024.40308.

[61] G. Liu, C. Li, and E. Zhang, "OpsLLM for cloud incident triage: Bilingual RAG-based root cause analysis and alert summarization for AI infrastructure operations," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 97-111, 2024, doi: 10.69987/JACS.2024.40408.

[62] W. Su, S. Chen, and C. Zhao, "Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark," J. Technol. Informatics Eng., vol. 4, no. 3, pp. 649-662, 2025, doi: 10.51903/jtie.v4i3.543.

[63] D. Zheng, B. Zhang, and J. Geibel, "VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification," J. Adv. Comput. Syst., vol. 4, no. 1, pp. 67-82, 2024, doi: 10.69987/JACS.2024.40106.

[64] W. Su, S. Chen, and E. Qian, "Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck," J. Technol. Informatics Eng., vol. 5, no. 1, pp. 327-340, 2026, doi: 10.51903/jtie.v5i1.549.

[65] B. Zhou, J. Jin, and D. Zhao, "Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets," J. Technol. Informatics Eng., vol. 4, no. 3, pp. 625-648, 2025, doi: 10.51903/jtie.v4i3.529.

[66] Y. Chen, Y. Zhang, and M. Sherman, "Going concern and bankruptcy prediction under extreme class imbalance: Cost-sensitive learning, resampling, and focal loss with explainable financial-ratio portraits," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 80-96, 2024, doi: 10.69987/JACS.2024.40407.

[67] Z. S. Zhong, C. Li, and H. Rao, "Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench," J. Technol. Informatics Eng., vol. 5, no. 1, pp. 341-360, 2026, doi: 10.51903/jtie.v5i1.539.

[68] Y. Chen and H. Xu, "Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access," Int. J. Graph. Des., vol. 4, no. 1, pp. 141-164, 2026, doi: 10.51903/ijgd.v4i1.3552.

[69] B. Zhou, H. Wang, and X. Chang, "Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447-463, 2025, doi: 10.51903/jtie.v4i2.522.

[70] C. Li, G. Liu, and Z. Zhao, "Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files," J. Technol. Informatics Eng., vol. 5, no. 2, pp. 91-103, 2026, doi: 10.51903/jtie.v5i2.538.

[71] X. Chang, Y. Lu, and Z. S. Zhong, "Review-grounded explainable recommendation with faithfulness evaluation on Amazon reviews," J. Electr. Eng. Comput. Sci., vol. 11, no. 1, pp. 9-22, 2026, doi: 10.54732/jeecs.v11i1.2.

[72] D. Zheng and C. Li, "Behavior-level jailbreak resistance via multi-stage refusal + utility preservation," J. Adv. Comput. Syst., vol. 4, no. 1, pp. 83-99, 2024, doi: 10.69987/JACS.2024.40107.

[73] Z. Li, K. Zhang, and A. Wong, "Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data," J. Technol. Informatics Eng., vol. 5, no. 2, pp. 75-90, 2026, doi: 10.51903/jtie.v5i2.541.

[74] B. Zhang, X. Sun, G. Liu, and B. Zhou, "LLM-style DevOps copilot for cloud-native troubleshooting: Retrieval-augmented runbook generation and command-safety evaluation," J. Technol. Informatics Eng., vol. 5, no. 2, pp. 104-118, 2026, doi: 10.51903/jtie.v5i2.534.

[75] M.-J. Kuo, D. Zheng, and J. Hires, "Federated topic-preference learning for knowledge-grounded chat with differential privacy," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 385-401, 2025, doi: 10.51903/jtie.v4i2.502.

[76] Y. Zhang and H. Zhang, "A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 464-486, 2025, doi: 10.51903/jtie.v4i2.547.

[77] J. Nie, G. Liu, C. Li, and T. Zou, "Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces," Int. J. Graph. Des., vol. 4, no. 1, pp. 179-185, 2026, doi: 10.51903/ijgd.v4i1.3703.

[78] Y. Zhang and H. Zhang, "Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces," Int. J. Graph. Des., vol. 3, no. 1, pp. 214-229, 2025, doi: 10.51903/ijgd.v3i1.3722.

[79] Z. Li, S. Zhou, and Z. Zhou, "Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench," Int. J. Graph. Des., vol. 3, no. 1, pp. 196-210, 2025, doi: 10.51903/ijgd.v3i1.3715.

[80] Y. Li, S. Lu, and L. Zhao, "LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment," Int. J. Graph. Des., vol. 3, no. 1, pp. 196-215, 2025, doi: 10.51903/ijgd.v3i1.3661.

Downloads

Published

2026-08-07

How to Cite

Sun, J., Henderson, T., Zhao, K., & Cooper, A. (2026). Uncertainty-Aware Lightweight Breast Ultrasound Modeling on BreastMNIST+: Parameter-Efficient Label-Prompt Alignment, Controlled-Shift Robustness, and Verifiable Evidence Cards. Journal of Information Systems and Business Technology, 2(4), 48-57. https://journal.jci.co.id/jisbt/article/view/584