LANTERN, a specification-guided dynamic conformance testing framework that mutates CTS tests using constraints extracted from the WebGPU specification, is introduced, demonstrating that syntactically seeded, semantics-aware mutation of conformance tests provides a way to uncover browser bugs during WebGPU testing.
Abstract
WebGPU is a low-level graphics and compute API that exposes modern GPU functionality to web applications. While the official WebGPU Conformance Test Suite (CTS) focuses on well-formed usage under the WebGPU specification, it is not designed to stress implementations with semantic edge cases or adversarial inputs. General-purpose fuzzers, in contrast, struggle with WebGPU because of its complex graphics stack and multi-process architecture. We introduce LANTERN, a specification-guided dynamic conformance testing framework that mutates CTS tests using constraints extracted from the WebGPU specification. LANTERN extracts explicit syntactic API rules from WebIDL definitions and recovers semantic constraints, such as command ordering and object lifetimes, from natural-language specification text. Selected rules guide AST-located textual transformations that generate both valid and intentionally invalid CTS variants. We execute the resulting tests at scale on AddressSanitizer-instrumented Chromium to discover bugs. Our evaluation discovers three reproducible bugs, including a heap corruption, in Chromium versions current at the time of study. These results demonstrate that syntactically seeded, semantics-aware mutation of conformance tests provides a way to uncover browser bugs during WebGPU testing.
We present WGSLsmith, a tool for randomised testing of compilers for the WebGPU Shading Language (WGSL). Under the WebGPU API---now supported by all three major browsers---GPU programs are written in the WebGPU Shading Language (WGSL), and every implementation ships a WGSL compiler that validates shaders and translates...
Michał Andryskowski, Amber Gorzynski, Hasan Mohsin et al.· Companion Proceedings of the...· 0 citations
Coverage feedback is an important source of guidance for fuzzing. However, obtaining such feedback normally requires application-level instrumentation that is specific to the language and runtime of the application. Given that modern web applications span multiple languages and runtimes, this application-level instrume...
I. P. A. Dharmaadi, Elias Athanasopoulos, Fatih Turkmen· 0 citations
The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
The security of the modern web depends on the correctness of JavaScript (JS) engines, yet these complex systems remain vulnerable to high-impact bugs. A critical limitation of state-of-the-art fuzzers is the coverage plateau: once a fuzzer saturates the control-flow graph, edge coverage loses its ability to guide disco...
Wai-Kin Wong, Dong-Wei Xiao, Anthony Cheuk Tung Lai et al.· Proceedings of the ACM SIGOP...· 0 citations
This paper investigates the capability of large language models (LLMs) to perform automated Wasm deobfuscation and introduces a three-tier evaluation hierarchy for assessing deobfuscation quality, consisting of syntax correctness, execution validity, and semantic similarity.
A security assessment on 75 FastAPI backends generated by three contemporary LLMs revealed a disconnect between functional correctness and secure logic, which is interpreted as a review-risk pattern, which is called the human-in-the-loop paradox.
Abdul Ali Khan, S. Rauti, T. Mäkilä· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.