From Enterprise UI Screenshots to Trustworthy Front-End Prototypes: Structural–Visual Consistency, Accessibility, Progressive Rendering, and an Evidence Layer for LLM-Assisted Design Critique
Keywords:
Design2Code, Screenshot-To-Code, Multimodal AI, Front-End Generation, Accessibility, Progressive Rendering, Prototype Retrieval, LLM Design Critique, Structural–Visual ConsistencyAbstract
Screenshot-to-code systems are often judged by visual resemblance, yet deployable prototypes also require coherent DOM structure, accessible semantics, and predictable rendering. We evaluated these dimensions on 484 Design2Code screenshot–HTML pairs using deterministic screenshot features, static markup audits, and CSS-hash-grouped five-fold cross-validation. Random Forest and Extra Trees predicted DOM size with R² values of 0.215 and 0.244, respectively, but accessibility-burden R² remained below zero, confirming that visual fidelity cannot substitute for markup inspection. Train-fold-only visual retrieval scored 78.26 on a composite consistency measure; structural reranking scored 79.59, and a deterministic evidence reranker combining predicted structure, accessibility, and static critical-render-path profiles scored 79.63. Its 1.37-point gain over visual retrieval was significant (95% bootstrap confidence interval 0.92–1.86; one-sided Wilcoxon p < 0.001) and cost 0.36 visual-similarity points. Conservative remediation removed 218 of 1,981 detected issues, raised zero-issue pages from 59 to 79, and preserved normalized body text and embedded CSS on all pages. On a fixed 25-page diagnostic subset, mean SSIM relative to each page's complete-CSS render increased from 0.693 for semantic HTML to 0.784 with typography; the complete renders independently reached median SSIM 0.996 against the stored screenshots. These findings establish an auditable evidence layer for subsequent LLM-assisted critique rather than an evaluated language-model critic.
References
[1] C. Si, Y. Zhang, R. Li, Z. Yang, R. Liu, and D. Yang, “Design2Code: Benchmarking multimodal code generation for automated front-end engineering,” in Proc. NAACL, 2025, pp. 3956–3974, doi: 10.18653/v1/2025.naacl-long.199.
[2] SALT-NLP, “Design2Code-hf,” Hugging Face Datasets, 2024. [Online]. Available: https://huggingface.co/datasets/SALT-NLP/Design2Code-hf
[3] C. Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, 2020.
[4] T. Beltramelli, “pix2code: Generating code from a graphical user interface screenshot,” in Proc. EICS, 2018, Art. no. 3, doi: 10.1145/3220134.3220135.
[5] A. Robinson, “Sketch2Code: Generating a website from a paper mockup,” arXiv:1905.13750, 2019.
[6] Y. Chen and M. Li, “From Hand-Drawn Sketches to Interactive Web Prototypes: A Reproducible Vision-Language Approach with Structural and Visual Consistency Evaluation,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 364–384, Aug. 2025, doi: 10.51903/jtie.v4i2.490.
[7] H. Laurençon, L. Tronchon, and V. Sanh, “Unlocking the conversion of web screenshots into HTML code with the WebSight dataset,” arXiv:2403.09029, 2024.
[8] S. Yun et al., “Web2Code: A large-scale webpage-to-code dataset and evaluation framework for multimodal LLMs,” in Advances in Neural Information Processing Systems, 2024.
[9] Y. Wan et al., “Divide-and-conquer: Generating UI code from screenshots,” Proc. ACM Softw. Eng., vol. 2, FSE, Art. FSE094, pp. 2099–2122, 2025, doi: 10.1145/3729364.
[10] D. Soselia, K. Saifullah, and T. Zhou, “Learning UI-to-code reverse generator using visual critic without rendering,” arXiv:2305.14637, 2023.
[11] J. Zhang, “Adaptive User Interface Design for Volleyball Learning Apps: Empirical Evidence from Google Play Reviews and Mobile Screen Analysis,” Int. J. Graph. Des., vol. 3, no. 1, pp. 175–195, May 2025, doi: 10.51903/ijgd.v3i1.3618.
[12] K. Lee et al., “Pix2Struct: Screenshot parsing as pretraining for visual language understanding,” in Proc. ICML, 2023, pp. 18893–18912.
[13] G. Baechler et al., “ScreenAI: A vision-language model for UI and infographics understanding,” in Proc. IJCAI, 2024, doi: 10.24963/ijcai.2024/339.
[14] X. Chen et al., “WebSRC: A dataset for web-based structural reading comprehension,” in Proc. EMNLP, 2021, pp. 4173–4185, doi: 10.18653/v1/2021.emnlp-main.343.
[15] J. Li, Y. Xu, L. Cui, and F. Wei, “MarkupLM: Pre-training of text and markup language for visually rich document understanding,” in Proc. ACL, 2022, doi: 10.18653/v1/2022.acl-long.420.
[16] Y. Jiang, C. Zhou, V. Garg, and A. Oulasvirta, “Graph4GUI: Graph neural networks for representing graphical user interfaces,” in Proc. CHI, 2024, doi: 10.1145/3613904.3642822.
[17] P. Duan, C.-Y. Chen, G. Li, B. Hartmann, and Y. Li, “UICrit: Enhancing automated design evaluation with a UI critique dataset,” in Proc. UIST, 2024, doi: 10.1145/3654777.3676381.
[18] P. Duan, J. Warner, Y. Li, and B. Hartmann, “Generating automatic feedback on UI mockups with large language models,” in Proc. CHI, 2024, doi: 10.1145/3613904.3642782.
[19] Y. Li, S. Lu, and L. Zhao, “LLM-as-Design-Critic: Aligning AI-Generated UI Feedback with Human Graphic Design Judgment,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–215, May 2025, doi: 10.51903/ijgd.v3i1.3661.
[20] Q. Wu, S. Meng, and J. Zhao, “Text-Grounded LLM-Assisted Design Rationale Interfaces: Turning Advertising Layout Metadata into Explainable UI/UX Decision Cards,” Int. J. Graph. Des., vol. 3, no. 1, pp. 216–240, May 2025, doi: 10.51903/ijgd.v3i1.3713.
[21] H. Tu, S. Zhao, and A. Zhou, “Visual Brief Cards for Advertising Design: A Structured UI/UX Framework for Turning Creative Intentions into Graphic Design Decisions,” Int. J. Graph. Des., vol. 3, no. 1, pp. 210–226, May 2025, doi: 10.51903/ijgd.v3i1.3714.
[22] J. Mu, Y. Lu, and E. Hwang, “Structured Visual Brief Interfaces for Advertising Design: A UI/UX Framework for Turning Creative Intentions into Designer-Editable Graphic Design Cards,” Int. J. Graph. Des., vol. 4, no. 1, pp. 192–208, Apr. 2026, doi: 10.51903/ijgd.v4i1.3702.
[23] A. Madaan et al., “Self-Refine: Iterative refinement with self-feedback,” in Advances in Neural Information Processing Systems, 2023.
[24] Z. Gou et al., “CRITIC: Large language models can self-correct with tool-interactive critiquing,” in Proc. ICLR, 2024.
[25] L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” in Advances in Neural Information Processing Systems, 2023, pp. 46595–46623.
[26] J. Huang et al., “Large language models cannot self-correct reasoning yet,” in Proc. ICLR, 2024.
[27] Q. Xin, “Auditable Automated Essay Scoring and Formative Feedback: A Rubric-Grounded Pipeline for Secondary and Higher Education,” J. Appl. Artif. Intell. Educ., vol. 2, no. 1, pp. 1–19, Jul. 2026, doi: 10.66053/jaaie.v2i1.348.
[28] A. Radford et al., “Learning transferable visual models from natural language supervision,” in Proc. ICML, 2021, pp. 8748–8763.
[29] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004, doi: 10.1109/TIP.2003.819861.
[30] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proc. CVPR, 2018, pp. 586–595, doi: 10.1109/CVPR.2018.00068.
[31] M. R. Luo, G. Cui, and B. Rigg, “The development of the CIE 2000 colour-difference formula: CIEDE2000,” Color Res. Appl., vol. 26, no. 5, pp. 340–350, 2001, doi: 10.1002/col.1049.
[32] B. Zhou, H. Wang, and X. Chang, “Distilling VMAF into an Edge-Deployable Quality Predictor: A Pilot Shot-Level Proxy with LLM-Ready Quality Tokens,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447–463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.
[33] D. F. Crouse, “On implementing 2D rectangular assignment algorithms,” IEEE Trans. Aerosp. Electron. Syst., vol. 52, no. 4, pp. 1679–1696, 2016, doi: 10.1109/TAES.2016.140952.
[34] M. Pawlik and N. Augsten, “Tree edit distance: Robust and memory-efficient,” Inf. Syst., vol. 56, pp. 157–173, 2016, doi: 10.1016/j.is.2015.08.004.
[35] W3C, “Web Content Accessibility Guidelines (WCAG) 2.2,” Oct. 2023. [Online]. Available: https://www.w3.org/TR/WCAG22/
[36] W3C, “Accessible Rich Internet Applications (WAI-ARIA) 1.2,” Jun. 2023. [Online]. Available: https://www.w3.org/TR/wai-aria-1.2/
[37] Deque Systems, “axe-core: Accessibility engine for automated Web UI testing,” 2024. [Online]. Available: https://github.com/dequelabs/axe-core
[38] H. Xu, Y. Chen, and A. Med, “Automatic Detection and Explanation of Dark Patterns from Interface Microcopy: Empirical Comparison of BERT-Style Encoders, RoBERTa-Style Encoders, and LLM-Style Decoders on the ec-darkpattern Dataset,” J. Technol. Informatics Eng., vol. 4, no. 3, pp. 590–612, Dec. 2025, doi: 10.51903/jtie.v4i3.491.
[39] P. Panchekha, A. T. Geller, M. D. Ernst, Z. Tatlock, and S. Kamil, “Verifying that web pages have accessible layout,” in Proc. PLDI, 2018, pp. 1–14, doi: 10.1145/3192366.3192407.
[40] Y. Weiss and N. Peña Moreno, “Largest Contentful Paint,” W3C Working Draft, Sep. 2023. [Online]. Available: https://www.w3.org/TR/largest-contentful-paint/
[41] W. Liu, X. Yang, H. Lin, Z. Li, and F. Qian, “Fusing Speed Index during web page loading,” Proc. ACM Meas. Anal. Comput. Syst., vol. 6, no. 1, Art. no. 23, 2022, doi: 10.1145/3511214.
[42] H. Wang, Y. Ren, and X. Chang, “Layout-Aware Progressive PDF Rendering: AI Prioritization of PDF Slices to Reduce Time-to-Functional-First-Frame on FUNSD,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425–446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.
[43] J. Jin, “LLM-Style Evidence Cards for Scientific Search Interfaces: A UI/UX Design Framework for Retrieval Transparency, Ranking Trust, and Visual Evidence Hierarchy,” Int. J. Graph. Des., vol. 3, no. 2, pp. 397–414, Oct. 2025, doi: 10.51903/ijgd.v3i2.3698.
[44] Z. S. Zhong, Q. Wu, and G. Mi, “Uncertainty-Aware Medical Image Explanation Cards: LLM-Generated Visual Explanations for AI-Assisted Radiology Interfaces,” Int. J. Graph. Des., vol. 3, no. 2, pp. 415–436, Oct. 2025, doi: 10.51903/ijgd.v3i2.3616.
[45] B. Zhou, C. Li, and L. Liu, “Risk-Calibrated Patient-Facing AI Safety Cards: A UI/UX Design Framework for Rubric-Based Medical Risk Communication,” Int. J. Graph. Des., vol. 3, no. 2, pp. 365–380, Oct. 2025, doi: 10.51903/ijgd.v3i2.3696.
[46] C. Li, B. Zhou, and K. Gao, “Risk-Calibrated Patient-Facing AI Safety Cards: A UI/UX Benchmark for Explainable Medical AI Response Interfaces,” Int. J. Graph. Des., vol. 3, no. 2, pp. 381–394, Oct. 2025, doi: 10.51903/ijgd.v3i2.3709.
[47] B. Zhang, Y. Ren, and J. Zou, “LLM-Style Explainable E-Commerce Recommendation Cards: A UI/UX Design Framework for Trust-Calibrated Product Recommendation,” Int. J. Graph. Des., vol. 3, no. 2, pp. 381–396, Oct. 2025, doi: 10.51903/ijgd.v3i2.3697.
[48] X. Chang, Y. Lu, and Z. S. Zhong, “Review-Grounded Explainable Recommendation with Faithfulness Evaluation on Amazon Reviews,” J. Electr. Eng. Comput. Sci., vol. 11, no. 1, pp. 9–22, Jun. 2026, doi: 10.54732/jeecs.v11i1.2.
[49] Z. Li, S. Zhou, and Z. Zhou, “Financial Risk Dashboard Design for Institutional RWA Investors: Visual Hierarchy, Chart Comprehension, and Explainability in FinChart-Bench,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–210, May 2025, doi: 10.51903/ijgd.v3i1.3715.
[50] G. Liu, S. He, and H. Wong, “LLM-Compatible Visual Brief Cards for AI Infrastructure Capacity Dashboards: A UI/UX Framework for Turning Forecast Risk into Graphic Design Decisions,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–213, May 2025, doi: 10.51903/ijgd.v3i1.3723.
[51] Y. Zhang and H. Zhang, “Visualizing the Right Counseling Support: Evidence-Linked Recommendation Cards for Explainable Mental Health Intake Interfaces,” Int. J. Graph. Des., vol. 3, no. 1, pp. 214–229, May 2025, doi: 10.51903/ijgd.v3i1.3722.
[52] Y. Zhang and H. Zhang, “A Therapist-Facing Session Copilot for Live Counseling Support: Reasoning-Guided Retrieval and Ranking from Multi-Turn Counseling Dialogues,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 464–486, Aug. 2025, doi: 10.51903/jtie.v4i2.547.
[53] K. Zhang, Y. Chen, and A. Qian, “Evidence-Grounded Accounting Disclosure Review Cards: A Visual Communication Framework for LLM-Style Explanations over SEC Financial Statements and Notes,” Int. J. Graph. Des., vol. 3, no. 2, p. 395, Oct. 2025, doi: 10.51903/ijgd.v3i2.3710.
[54] Y. Chen and H. Xu, “Trust-Calibrated Multilingual RAG for Humanitarian Information Platforms: Empirical Evaluation on OMoS-QA for Migration Information Access,” Int. J. Graph. Des., vol. 4, no. 1, pp. 141–164, Apr. 2026, doi: 10.51903/ijgd.v4i1.3552.
[55] Y. Li, “Findable then Explainable: Retrieval–Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[56] B. Zhang, H. Rao, and D. Zhao, “Evidence-Grounded RAG for Cloud-Native DevOps: Hallucination-Resistant AIOps Question Answering over Private Operations Documents,” J. Adv. Comput. Syst., vol. 4, no. 3, pp. 109–125, Mar. 2024, doi: 10.69987/JACS.2024.40308.
[57] C. Li, J. Bai, and S. Wang, “Evidence-Chain Reliable RAG: Word-Level Hallucination Detection, Source Attribution, and Provenance Explanation for LLM Applications,” J. Adv. Comput. Syst., vol. 4, no. 2, pp. 76–92, Feb. 2024, doi: 10.69987/JACS.2024.40207.
[58] J. Jin, “Evidence-Chain Reliable RAG: Hallucination Detection, Source Attribution, and Deterministic Provenance Explanations,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 520–533, Aug. 2025, doi: 10.51903/jtie.v4i2.535.
[59] Z. S. Zhong, J. Chen, E. Zhong, and X. Sun, “Evidence-Calibrated RAG for Unanswerable Question Answering: Retrieval Coverage, Abstention Calibration, and Hallucination-Proxy Analysis on SQuAD 2.0,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 502–520, Aug. 2025, doi: 10.51903/jtie.v4i2.536.
[60] R. Zhang, Z. Wen, C. Wang, C. Tang, P. Xu, and Y. Jiang, “Quality Analysis and Evaluation Prediction of RAG Retrieval Based on Machine Learning Algorithms,” arXiv preprint arXiv:2511.19481, 2025, doi: 10.48550/arXiv.2511.19481.
[61] W. Su, H. Rao, and E. Ma, “Privacy and Data-Integrity Risk Cards for LLM Agents: A UI/UX Design Framework for Secure Human Oversight under Prompt-Injection Attacks,” Int. J. Graph. Des., vol. 4, no. 1, pp. 186–191, Apr. 2026, doi: 10.51903/ijgd.v4i1.3699.
[62] D. Zheng, B. Zhang, and J. Geibel, “VerifySafe: Toxicity-Safe Agent Responses under Adversarial Prompts with Evidence-Based Self-Verification,” J. Adv. Comput. Syst., vol. 4, no. 1, pp. 67–82, Jan. 2024, doi: 10.69987/JACS.2024.40106.
[63] J. Nie, G. Liu, C. Li, and T. Zou, “Evidence-Constrained Incident Visualization Cards for Distributed Cloud Logs: A UI/UX Framework for Turning Hadoop, OpenStack, and ZooKeeper Logs into Actionable SRE Design Interfaces,” Int. J. Graph. Des., vol. 4, no. 1, pp. 179–185, Apr. 2026, doi: 10.51903/ijgd.v4i1.3703.
[64] W. Su, S. Chen, and E. Qian, “Narrative-Aware Scientific Claim Verification Agent with Evidence Ranking for ClimateCheck,” J. Technol. Informatics Eng., vol. 5, no. 1, pp. 327–340, Apr. 2026, doi: 10.51903/jtie.v5i1.549.
[65] D. Zheng, C. Li, and H. Davidson, “Continual Red-Teaming for In-the-Wild Jailbreaks via Online Guardrail Updates and Guardrail Distillation,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 35–49, Feb. 2023, doi: 10.69987/JACS.2023.30203.
[66] D. Zheng and C. Li, “Behavior-Level Jailbreak Resistance via Multi-Stage Refusal + Utility Preservation,” J. Adv. Comput. Syst., vol. 4, no. 1, pp. 83–99, Jan. 2024, doi: 10.69987/JACS.2024.40107.
[67] B. Zhang, X. Sun, G. Liu, and B. Zhou, “LLM-Style DevOps Copilot for Cloud-Native Troubleshooting: Retrieval-Augmented Runbook Generation and Command-Safety Evaluation,” J. Technol. Informatics Eng., vol. 5, no. 2, pp. 104–118, Aug. 2026, doi: 10.51903/jtie.v5i2.534.
[68] Z. S. Zhong, C. Li, and H. Rao, “Trajectory Reliability Prediction for Generalist AI Agents: Tool-Use Failure Analysis and Success Forecasting on ZClawBench,” J. Technol. Informatics Eng., vol. 5, no. 1, pp. 341–360, Apr. 2026, doi: 10.51903/jtie.v5i1.539.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Wei Dong, Sierra Campbell, Lei Wu, Gabriel Ross (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a