Extrinsic World Modeling with Opus, Astra & Grok
One photo → a furnished 3D room. Five rooms, shared meshes, three build-and-furnish agents.
One workflow, three models.
Each model builds the room and places the same object meshes from a single reference photo. The renders below show the resulting scenes from nine camera angles.
Scores and execution averages cover all five rooms. Full comparison conditions are included below.
| Per scene | Grok 4.7 | Astra | Opus 5.5 |
|---|---|---|---|
| Input fidelity VLM / 5 ↑ | 3.225 | 3.575 | 3.400 |
| Orbit quality VLM / 5 ↑ | 3.036 | 3.267 | 3.053 |
| Model cost estimated USD | $36.39 | $7.81 | $2.83 |
| Job time build + furnish | 75.3 min | 15.8 min | 18.5 min |
| Average turns build + furnish | 267.4 | 46.0 | 31.2 |
Grok includes its continued initial attempts; scene 1 reused its completed 100-turn stages. Costs exclude compute, shared meshes and evaluation.
Click to enlarge. Rows match the requested orbit angle; model cameras and pivots differ from ground truth. Scores shown compare model renders against the input photo, not these ground-truth views.
Grok vs AstraOrbit A/B: Astra wins 56–19, with 5 ties.
Grok vs OpusOrbit A/B: Grok wins 49–24, with 7 ties.
ALL3D maintains the source scenes, execution traces and evaluation pipeline. These are available for a closer technical comparison.
Full scores, costs and comparison conditions
| Agent | Input CLIP ↑ | Orbit CLIP ↑ | SSIM ↑ | LPIPS ↓ | EdgeMBA ↑ |
|---|---|---|---|---|---|
| Grok 4.7 | 0.916 | 0.843 | 0.540 | 0.654 | 0.342 |
| Astra | 0.940 | 0.863 | 0.561 | 0.519 | 0.436 |
| Opus 5.5 | 0.927 | 0.847 | 0.561 | 0.551 | 0.385 |
| Opponent | Input | Orbit |
|---|---|---|
| Astra | 0 / 0 / 10 | 19 / 5 / 56 |
| Opus 5.5 | 4 / 0 / 6 | 49 / 7 / 24 |
How to read the comparison
Grok 4.7, GPT-6 Astra and Claude Opus 5.5 use shared pre-generated Meshy assets. Five selected scenes, one reconstruction per model per scene, nine native views: 0°, ±15°, ±45°, ±90° and ±135°. This changes the build/furnish agents, not every model in ALL3D.
Gemini 3.8 Flash rates each image twice against the input; pairs are anonymous and judged in both orders. A/B counts are judge calls, not independent scenes. CLIP uses a pinned OpenAI ViT-L/14 checkpoint. Pixel metrics cover the input view only. Unequal budgets, provider adaptations and camera estimates limit causal claims. This is not a general model leaderboard.
Grok and Opus disabled scene-library access. Astra scene 1 allowed it; reviewed reuse included a French-door reference and oak-floor materials. All tables include all five scenes.
Cost and time accounting
Four of five Grok builders needed continuation. Raising the cap from 100 to 500 turns let them finish in 125–225 total build turns. Astra and Opus ran with 100-turn caps.
Logged SDK model-cost estimates, not invoices. Build + furnish job time includes tool waits; it is not GPU-billed time or end-to-end pipeline latency. Grok includes its initial attempts and continuations; Astra and Opus use the selected scored runs. Earlier discarded experiments are excluded.
Grok total: $181.96 and 6.28 summed job hours. Parallel batch span: 103.1 minutes. Zero new Meshy credits. Compute, rendering, evaluation and preflight charges are excluded.
WebP previews are for display; scoring used the original PNGs. ALL3D conducted this comparison without provider endorsement.
Download scores, image hashes and accounting (JSON)


































