Robots have been learning in invented worlds for far longer than the current enthusiasm for simulation might suggest. In 1989, Dean Pomerleau's ALVINN at Carnegie Mellon was trained on simulated road images before being successfully tested on the university's autonomous navigation vehicle. The graphics were crude, but the lesson has proved durable: a synthetic world does not have to be a perfect copy of reality, as long as it contains the structure, variation and feedback the task demands.
That, at heart, is the premise of an asset factory. A conventional 3D model tells us what an object looks like. Computer-aided design (CAD) goes further by specifying engineering dimensions and how an object is assembled. A simulation-ready asset must also tell a robot what the object weighs, how it moves, where it collides, how its surfaces respond and what its sensors will perceive.
If that asset is intended for reinforcement learning, it must go further still by defining what the robot can observe and do, what counts as progress and which states represent success or failure.
The challenge is creating these assets efficiently and at scale. Our Asset Factory Blueprint brings together the principles behind an auditable, repeatable agentic pipeline for creating simulation- and learning-ready assets.
From descriptions to machine-operable assets
Inputs to an asset factory can take many forms. They might be structured sources such as CAD or building information modeling (BIM) files, bills of materials, inventory catalogs or robot descriptions. They may also be photographs, partial scans, maintenance manuals, operator notes or even a one-line description such as “a wheeled stainless-steel cart with two locking doors.”
Whatever the input, the factory preserves the evidence and determines what is known, what is missing and what remains uncertain. Provenance and record-keeping are fundamental to the process because small errors can have significant consequences. An incorrect scale can change grasp geometry, while a flipped axis can introduce subtle simulation errors that undermine reinforcement learning.
The Blueprint maintains a running manifest of sources and tracks units, axes, source rights and checksums throughout the process.
The source then moves through the stages it requires, from reconstruction or conditioning through mesh verification, semantic segmentation, material and physical-property inference, texturing, physics and articulation to SimReady verification.
What emerges is not simply a representation of the object but a package of verified geometry, named parts, materials, collision bodies, mass properties, joints, affordances, permitted variants, provenance and uncertainty. The task determines which of these elements are required and what simulation-readiness means in practice.
Consider a warehouse tote with a hinged lid. What does a robot need to know about it?
That depends on the task. A robot opening the tote needs the correct scale, hinge axis and limits, lid inertia, collision geometry, friction and handle location. A robot sorting it may care more about appearance and sensor response. A robot stacking it may prioritize center of mass and deformation.
Fidelity is task-dependent. Appearance and behavior also need to agree: a surface that looks like rubber but behaves like polished steel teaches the wrong contact expectations.
This is why the Blueprint can draw on bill-of-materials entries, CAD material names, specifications, images and operator notes to propose physical properties with units, ranges, evidence and confidence attached.
NVIDIA's VoMP, for example, predicts fine-grained mechanical properties including density, Young's modulus and Poisson's ratio across 3D objects.
Rather than relying on a single inference, evidence can be accumulated and used to refine values until an acceptable level of certainty is reached. The result is a rebuildable, machine-operable asset.
From object to objective
A complete physical asset, however faithful, is still not a reinforcement-learning environment. It must also be connected to a robot, sensors, a scene and a task contract.
The Blueprint's reinforcement-learning extension provides that contract by defining what can be observed, what actions are available, how reward is calculated, when episodes end, how the scene resets, which variations are permitted and which tests verify that the environment starts correctly.
This creates a goal-driven training landscape.
In a pick-and-place task, physical states are not equal in value. Closing the gripper, achieving a stable lift and reaching the destination should move the policy toward higher reward. Collision, slip, a crushed package or an exploited simulator loophole should move it away from that goal or end the attempt.
In this operational sense, the environment “understands” the task by mapping spatial and physical conditions to training signals.
The scene becomes part of the learning system itself. Geometry determines what is reachable and physics determines the consequences of action. Sensors determine what can be observed and the reward signal determines which behaviors are reinforced.
A poorly authored world can teach the wrong lesson with tremendous computational efficiency. Conversely, GPU-accelerated simulation environments can in some cases reduce reinforcement-learning training time by two to three orders of magnitude compared with conventional reinforcement-learning training using a CPU-based simulator.
As the factory controls variation, it can also support adaptive curricula. Training might begin with stable lighting and uncluttered layouts before introducing new viewpoints, loads, friction, wear, positions, sensor noise and distractors. This allows complexity to increase in line with the learning objective.
Neither a digital twin nor a simulator
It is important to distinguish asset factories from two related concepts: digital twins and simulators.
A digital twin typically represents a particular asset, facility or process. An asset factory may consume digital twin inputs or produce assets that eventually live inside one, but its primary purpose is different: it creates the assets and variations a robot needs to learn.
That distinction changes the target. One exact replica of a single blue tote may be less useful for training than 10,000 plausible totes spanning different dimensions, wear, surface properties, loads and placements. The goal is to create a distribution of experiences that helps prepare a robot for real-world variation rather than always reproducing one exact instance.
An asset factory is also not a simulator. The simulator executes physics, renders sensors and advances time. The factory builds and validates the objects, scenes and task contracts that the simulator executes.
In software terms, the simulator is the runtime, while the factory provides the compiler, build system and test suite.
More realism is not automatically better either. Grasping may depend on contact geometry, friction and mass far more than on a perfectly reproduced logo, while visual inspection may reverse those priorities. Fidelity should be determined by the task and the training signal rather than by how realistic a render appears.
The factory is the point
Historically, simulation content has often required handoffs among CAD engineers, 3D artists, material specialists, physics engineers, roboticists and reinforcement-learning teams. That approach can work for individual demonstrations but becomes difficult to scale across thousands of training environments.
“A 3D model shows the thing. CAD specifies it. An asset factory gives it consequences and purpose.”
The purpose of the “factory” model is to introduce repeatability, routing, throughput and quality control.
An agentic orchestration layer can convert a request into a dependency-complete plan of stages, providers, typed inputs and outputs, validation gates, retries and stop conditions. Different source types can take different routes through the pipeline, while missing units, rights or task-critical evidence create a blocked state instead of being replaced with unsupported assumptions.
None of this removes human judgment. It relocates it to the points where it creates the most value.
Deterministic checks can validate schemas, checksums, units, package completeness, OpenUSD loading and physics configuration. Vision-language review can identify perceptual defects that formal validators may miss. Bounded repairs can address known fault classes, while unresolved issues, physics validation and release decisions can be escalated to a person.
This is also how public and proprietary capabilities can work together. OpenUSD and public reconstruction, geometry, simulation and validation technologies provide an interoperable foundation. We combine these with proprietary material rendering, physical-property inference, agentic coordination and environment design.
A modular architecture also allows new technologies to be integrated into the pipeline as they mature.
Simulation as a first-class source of experience
Simulation has often been treated as a substitute for reality. In Physical AI, it can also serve as a first-class source of training experience.
The real world remains the final authority, but it is a difficult sole training environment. It can be slow, costly, difficult to reset and, in some cases, dangerous. Rare failures remain rare, internal state can remain hidden and labels can be expensive to generate.
Simulation changes those constraints. It can reproduce edge cases on demand, expose hidden state, generate exact labels and run variations in parallel. This makes it possible to create datasets that would be impractical, unsafe or impossible to collect physically.
The approach can extend from highly detailed engineering environments to much larger operational settings. CAD and BIM data, for example, can help create training environments for robots operating in mines, plants and other infrastructure environments.
Healthcare offers another example. The 1000 Lungs dataset contains 1,000 synthetic virtual bronchoscopy episodes designed to support navigation research in synthetic human airways. Each case includes generated anatomy, airway geometry, a procedure-style camera path, rendered condition streams and trajectory data.
Where real-world data is costly or difficult to obtain at the required scale, simulation can provide additional training experience.
None of this eliminates the sim-to-real challenge. Incorrect physics can simply produce incorrect experience at scale.
That is why uncertainty, provenance, validation results and task fitness need to travel with each asset, while real-world observations continue to calibrate the factory and close the loop.
The Blueprint provides a reference framework for this process, while implementations, including our proprietary tools, can add the physical accuracy required for specific use cases.
Industrializing world-building
An asset factory is not simply a content generator with a physics layer added. It is a system for converting descriptions of physical objects into machine experience spanning geometry, materials, behavior, sensing, tasks and goals.
As robotic models improve, the bottleneck increasingly shifts toward the environments in which they learn.
Training policies need experience that is diverse yet controlled, synthetic yet evidence-based and designed around the task rather than visual spectacle. Asset factories make that process more repeatable and scalable.
Robots do not need a world that merely looks convincing. They need one that responds correctly, varies intelligently and knows what success means.






