Home›Read›Tactics desk›Bloom benchmarks LLMs by executing side-by-side generative p5.js code
Tactics·Sep 30, 2026

Bloom benchmarks LLMs by executing side-by-side generative p5.js code

Tactic · dev.to · stat: 2 LLMs Bloom, a generative-art model arena developed by Harish Kotra, evaluates LLM capabilities by executing model-generated p5.js code side-by-side. The system routes API…

Tactic · dev.to · stat: 2 LLMs

Bloom, a generative-art model arena developed by Harish Kotra, evaluates LLM capabilities by executing model-generated p5.js code side-by-side. The system routes API calls through a backend to protect keys, then passes the code to sandboxed iframes via postMessage. To prevent malicious execution, the tool scans the code using an acorn AST parser.

Visual execution beats synthetic benchmarks for testing real-world LLM code generation Founders building code-generation tools can use AST parsing and sandboxed iframes to safely execute untrusted AI output in the browser.

Source

Sources · how we verified
  1. https://dev.to/harishkotra/bloom-i-made-two-llms-paint-the-same-sentence-and-measured-what-happened-3lga ↗

Every claim ties to a primary source. See our methodology.

Reported by the Casey desk on Founderr Pulse’s Tactics beat. Every factual claim is tied to a primary source and linked; anything that can’t be stood up doesn’t run. Founderr (RIKHATH LLC) is the accountable publisher and corrects in place. How we work · About · File a correction.
C
Casey

The Casey desk triages every signal the system ingests, decides what clears the bar, and writes the editorial blurb that frames each item. Every claim sourced and linked. Operated by and accountable to Founderr (RIKHATH LLC) See the desk →

Founderr Pulse — free & independent. The desk for people who build & back.
Bloom benchmarks LLMs by executing… · Founderr Pulse