Build, our new mobile-first creation tool, allows a Roblox user to create a game using text prompts and publish it to players around the world. Generating a holistic game requires an end-to-end agentic system for game generation: one that coordinates code, visual assets, gameplay systems, UI, and controls into a single compound artifact. This system must robustly handle the various scenarios that emerge with player input and multiplayer interactions. Since the alpha launch in New Zealand, over 100,000 players have used Build to try creating a game, and approximately 9,000 Build games have been published. Of the thousands who’ve used Build to create a game, 71% had never used Roblox Studio before, which shows they’re trying something new with Build.
But creators will only use this new tool if it’s easy to use and the quality of the games it generates is high enough to build upon. To understand the quality of the games created with Build, we evaluate the end-to-end creation journey and use those evaluations as well as creator feedback to improve the product. Our evaluation (eval) framework has helped us prioritize and improve intent understanding, 3D playability, and polish. Since the alpha launch of Build in New Zealand, creator satisfaction has improved 10%, unintended outputs are down from 18% to 3%, and approximately 14% of prompted sessions and 19% of playtests are published. These are descriptive indicators of progress, not causal estimates of evaluation impact alone.
Why Evaluating Game Creation Is Different
Games share some of the same evaluation challenges found in other systems that generate compound artifacts, like apps, videos, and world models. For example, quality depends on how multiple components work together, and no single component score fully describes the outcome.
What sets games apart is that the artifact is an executable, real-time system driven by players. Properties such as responsiveness, agency, progression, and game feel emerge through interaction and cannot be fully assessed by inspecting code or assets alone. A game therefore has to clear a uniquely high bar of low-level correctness (defect-free and playable), as well as high-level player experience (a game that feels finished, not a prototype).
The question we seek to answer through evaluation is not whether any one piece is correct, but whether the holistic artifact matches creator intent, and whether players can play through the whole experience. Answering that takes more than one metric.
Four Signals, One Framework
Our Build creation eval framework combines four signals. Each answers a different question, catches a different class of failure, and yields a concrete quantitative score:
1. Game specification quality (0 - 100 score): Build is the end-to-end creation agent: It translates the creator’s prompt into a specification, then uses that to generate the game. The evaluator assesses whether the specification makes important decisions explicitly, preserves the creator’s intent, considers feasibility within Roblox’s runtime, and avoids inventing requirements that were never expressed.
2. Static game quality (0 - 100 score): The static place evaluator inspects the generated game in Roblox Studio edit mode. This includes game structure, scripts, assets, properties, connections, and client/server boundaries. We seek to answer two questions at this stage. The first is whether the game is built correctly: implementation of required systems, engine-API and networking choices, presence of expected controls and UI components, robustness and coherence of the game architecture.
The second is whether it is designed to be fun: core-loop depth, content variety and escalation, progression, replayability, and innovation beyond a genre default. A game can be architecturally clean but have no content depth, or be full of ideas but poorly wired.
Separating the two lenses and measuring against best practices gives us a more granular idea of where the game is failing.
3. Playtest (0 - 100% pass rate): The playtest evaluator runs the game and executes a test plan with real player-like input, including taps, clicks, keystrokes, and navigation, rather than firing events from code. A per-case test plan checks that the game starts, responds, progresses, recovers from failure, and stays free of runtime errors. Each item is a binary pass/fail, reported as a tier-weighted pass rate alongside a plain playable/not-playable verdict. We repeat runs to measure consistency.
4. Visual and experiential quality (0 - 100 score): Some of what matters most to a game’s look and feel cannot be read from code, so the visual evaluator scores presentation on its own using captured gameplay on viewports representing different devices. This rubric looks only at how a game looks and sounds based on asset craft, UI finish, motion, lighting, and audio. The core question is: Does the game read like a shippable product, or like a prototype? Keeping visuals separate from mechanics means that “the code works” and “the game looks finished” don’t average into one misleading score.
Alongside the scored signals, every run captures informative-only metrics, like token usage, tool calls, skills triggered, and latency. These metrics do not move the eval score, but they may explain how a result was produced and catch regressions that the rubrics miss, like a build that got twice as expensive for the same game.
The signals map onto the failures they catch:
Signal
Typical failure
Game spec
Vague loop; missing fail/retry or mobile control assumptions
Static place:
Engineering and design
Engineering: Systems present but miswired, unsafe client authority, UI covering the playfield, not mobile-ready
Design: Architecturally clean but shallow; no progression, variety, or hook
Playtest
Game does not start, respond, score, recover, or stay free of runtime errors
Visual
Runs correctly but looks unfinished, cluttered, or unreadable on a phone
Walking Through a Prompted 3D Obby Course
The signals are easiest to see on a real prompt. For example, “A true 3D obby game, with a start platform, moving hazards, checkpoints, and finish line.” Obbys are a popular platforming game type on Roblox. The creator intent is clear, even if it does not prescribe every implementation detail. Build needs to fill in the details with reasonable assumptions.
Plan and Spec
Build produced a spec doc for a new game, called Skyforge Ascent, detailing a five-section course with checkpoints and three hazard types. It defined the core loop, difficulty curve, art direction, screens, overlays, and controls. Spec eval scored it 91/100, earning full marks on completeness, prompt adherence, feasibility, and internal consistency. The deduction came from two specific dimensions: implementation detail and acceptance-criteria coverage.
The spec never pinned its numeric constants, such as spin speeds, slide distances, tile timings, platform dimensions, or reset-plane height, so an engineer still has to invent them. Additionally, a few acceptance criteria (like “about 24 platforms” or “celestial-forge palette”) aren’t strictly falsifiable. A high spec score only proves that the document is build-ready. Whether those commitments survive contact with actual game development is what the next three stages measure.
Static Checks: Code Quality and Content Quality
The virtual judge inspected the placefile of Skyforge Ascent without playing it, and the two evaluation lenses reached different conclusions.
The engineering lens, reading the same placefile, scored it 98/100. It found an exemplary build with five sections and three moving hazards, fully satisfying the creator’s intent using server-authoritative sequential checkpoints, fall-reset systems, and a UI built on the StyleSheet framework. Everything outlined in the spec was delivered. The only issue was two unused helper function exports.
The design lens scored it 31/100. Everything worked, but the game scored zero on progression, replayability, player agency, and original hook. The core loop was simple and linear—reach each checkpoint, cross the finish—with little variety across runs and no reason to come back for a second one. It is an obby genre staple, not a game with a hook of its own.
That gap is the finding. Code quality and content quality are fundamentally different dimensions, and a build can be near-perfect on one while failing the other. A single blended quality score would have averaged these into something meaningless.
Playtest
The playtest agent played the running game using real player-like input and passed all nine items on the test plan: spawning a character on the start platform, navigating 3D geometry across gaps, taking hazard impacts, triggering checkpoints and automatic fall recovery, completing the run, and restarting, all with zero runtime errors. In this specific run, Skyforge Ascent scored 100/100 on the playtest.
A perfect playtest score simply proves that the mechanics work and the game is technically playable from start to finish. It says nothing about whether the game is actually engaging or fun to play—that distinction is precisely why the design lens scored it a 31.
Visual Quality
Finally, the eval system scored the recorded session on a phone-sized viewport for visual presentation. Skyforge Ascent scored 86/100: clean, cohesive, and comfortably in line with the visuals of a typical obby; readable platforming geometry; legible hazards; and a modern, unobtrusive HUD. Minor visual limits keep it from perfection, but there is nothing broken, and nothing that would make a player stop scrolling either.
Trace Metrics
The same run recorded process metrics too. The game took roughly eight minutes to build, there were zero tool errors, and Build used the 3D, UI, script editing, asset insertion, and game audio skills. These metrics do not affect the eval score directly, but they tell us how the game was made and would flag a future build that got slower or more expensive for no gain.
Across five stages, Skyforge Ascent scored differently on game spec, code quality, content quality, playtest, and visual quality. Four of those dimensions say Build can turn a single prompt into a complete, correct, playable, presentable Roblox experience. One dimension says it isn’t a game that anyone would choose to play twice. Measuring that gap, instead of averaging it away, is the whole point of Build eval.
How Early Evaluation Has Improved Build
Evaluation results do more than generate scores, they also provide actionable direction for product improvements. By combining individual component debugging with holistic system-level testing, offline batch evaluations across standardized prompt suites have directly shaped Build’s development.
- Rapidly resolving alpha creator pain points: In our New Zealand alpha launch, users gave us feedback that Build sometimes created a 2D game when they wanted a 3D game. The eval framework helped our team quickly improve Build’s 3D capabilities. Within two weeks, we improved Build’s ability to understand a user’s intent to create a 3D game—even if they didn’t explicitly state it. We also refined 3D playability and polish by fixing common problems with object positioning, lighting, etc. Since then, “not what I want” complaints dropped from 18% to 3% among all complaints; overall dissatisfied reaction after playtest dropped 10% points. We also accumulated learnings when iterating on the framework itself.
- Comparative benchmarking across models and architectures: We use our eval framework to rigorously benchmark Build system configurations against standardized prompt suites. Before shipping changes to production, offline A/B eval reports compare candidate releases against our production baselines across all four dimensions:
For example, these offline diffs allow us to isolate how various model upgrades and agent routing decisions impact individual quality signals:
- Base model performance: Upgrading our base model delivered consistent gains across every signal—which resulted in fewer broken builds on the first try, superior visual polish, and deeper gameplay content.
- Agent routing architecture: Conversely, delegating specific sub-tasks to a lightweight model in a dual-agent setup preserved content depth and diversity, but caused sharp regressions in playability and visual finish.
Prod vs. Candidate
Playtest
Visual
Content
Model 1 vs Model 2
+7.6
+9.0
+5.2
Single agent vs dual agent
-11.2
-13.2
0.1
- Tracking quality baselines across creation paradigms: We also use the framework to understand the current quality of build at different times with the most relevant evaluation set. These sets provide snapshots of increasingly broad, complex creation tasks.The earliest set consists of entirely 2D games with relatively simple mechanisms. The game logic was easy to build, therefore the playtest score tended to be higher; on the other hand, the content score remained low for the exact same reason.As we move the focus to 3D games, the score distribution shifts naturally: bug-free play becomes more challenging, reducing playtest scores. Visual scores remain flat due to a lack of asset polish, giving games a prototype feel. Content scores are higher as 3D games have the potential to be deeper—but overall scores tend to remain low since single-turn prompts offer limited room for content expansion.
Set
Playtest
Visual
Content
Early 2D game set
75.9
67.1
19.2
Current 3D game set
51.5
66.3
33.9
- Uncovering genre-specific bottlenecks: Evaluation data also highlights significant performance variations across genres. Genres that are heavy on play mechanism (e.g., RPGs, shooters, racing games) tend to encounter more playability issues. Genres with relatively simple logic (e.g., survival and tycoon games) often perform better in playtest, making visual polish and UI design the next milestone to solve. Knowing the difference between genres incentivizes us to broaden our evaluation set with more balanced sampling across Build games.
- Lessons learned in evaluation harness design: Iterating on the framework itself taught us critical lessons about harness reliability:
- System metrics vs. component metrics: Clean code structure or API usage doesn't guarantee a playable game. We keep correctness, design, and presentation distinct to avoid hiding actionable bugs.
- Harness weaknesses and human calibration: Static checks can miss game-crashing runtime blockers, while automated playtests struggle with complex physics. Virtual judges are continuously calibrated against professional game developers and experienced community reviewers.
- Expanding to multiturn and real traffic: Because real creators iterate continuously (adding systems, adjusting difficulty, fixing bugs), we are expanding our eval harness to evaluate multiturn edits and real-world prompt distributions.
What Comes Next
We’re currently focused on extending our Build creation eval system in five directions:
- Broader coverage across genres, 2D/2.5D/3D experiences, multiplayer scenarios, and complexity tiers
- Stronger support for multi-turn edits
- More reliable playtest, screenshot, video, and audio evidence
- Calibration to human ratings at scale and creator-led benchmark cases
- Better integration between offline evaluation, online product signals, and regression testing
Our long-term goal is a transparent, continuously improving eval system that measures Build at the level creators care about: the quality of the experiences they can make, test, share, and continue developing. Today, these signals are used to evaluate Build-generated games. Over time, as we further refine and calibrate them against human judgment, they may generalize to other creation workflows and help provide a shared language for understanding game quality across the broader Roblox ecosystem.
For a related look at evaluating agentic systems for Roblox Assistant, see Using OpenGameEval to Benchmark Agentic AI Assistants for Roblox Studio.
Based on data through Sept. 1, 2026.
As of Aug. 29, 2026.












