Bloom benchmarks LLMs by executing side-by-side generative p5.js code
Tactic · dev.to · stat: 2 LLMs Bloom, a generative-art model arena developed by Harish Kotra, evaluates LLM capabilities by executing model-generated p5.js code side-by-side. The system routes API…
Tactic · dev.to · stat: 2 LLMs
Bloom, a generative-art model arena developed by Harish Kotra, evaluates LLM capabilities by executing model-generated p5.js code side-by-side. The system routes API calls through a backend to protect keys, then passes the code to sandboxed iframes via postMessage. To prevent malicious execution, the tool scans the code using an acorn AST parser.
Visual execution beats synthetic benchmarks for testing real-world LLM code generation Founders building code-generation tools can use AST parsing and sandboxed iframes to safely execute untrusted AI output in the browser.
Every claim ties to a primary source. See our methodology.