Extrinsic World Modeling with Opus, Astra & Grok

One photo → a furnished 3D room. Five rooms, shared meshes, three build-and-furnish agents.

One workflow, three models.

Each model builds the room and places the same object meshes from a single reference photo. The renders below show the resulting scenes from nine camera angles.

Scores and execution averages cover all five rooms. Full comparison conditions are included below.

Five-scene comparison · raw 3D outputs
Per sceneGrok 4.7AstraOpus 5.5
Input fidelity VLM / 5 ↑3.2253.5753.400
Orbit quality VLM / 5 ↑3.0363.2673.053
Model cost estimated USD$36.39$7.81$2.83
Job time build + furnish75.3 min15.8 min18.5 min
Average turns build + furnish267.446.031.2

Grok includes its continued initial attempts; scene 1 reused its completed 100-turn stages. Costs exclude compute, shared meshes and evaluation.

All nine angles · ground truth and model renders
AngleGround truth Authored sceneGrok 4.7AstraOpus 5.5
Input view · 0°Scene 1: ground truth at Input view · 0°Grok 4.7, scene 1, Input view · 0°3.25 / 5Astra, scene 1, Input view · 0°3.50 / 5Opus 5.5, scene 1, Input view · 0°3.50 / 5
−15°Scene 1: ground truth at −15°Grok 4.7, scene 1, −15°3.13 / 5Astra, scene 1, −15°3.50 / 5Opus 5.5, scene 1, −15°3.31 / 5
+15°Scene 1: ground truth at +15°Grok 4.7, scene 1, +15°3.25 / 5Astra, scene 1, +15°3.63 / 5Opus 5.5, scene 1, +15°3.38 / 5
−45°Scene 1: ground truth at −45°Grok 4.7, scene 1, −45°3.13 / 5Astra, scene 1, −45°3.50 / 5Opus 5.5, scene 1, −45°3.69 / 5
+45°Scene 1: ground truth at +45°Grok 4.7, scene 1, +45°2.50 / 5Astra, scene 1, +45°3.50 / 5Opus 5.5, scene 1, +45°3.25 / 5
−90°Scene 1: ground truth at −90°Grok 4.7, scene 1, −90°3.13 / 5Astra, scene 1, −90°3.65 / 5Opus 5.5, scene 1, −90°3.19 / 5
+90°Scene 1: ground truth at +90°Grok 4.7, scene 1, +90°2.65 / 5Astra, scene 1, +90°2.75 / 5Opus 5.5, scene 1, +90°2.85 / 5
−135°Scene 1: ground truth at −135°Grok 4.7, scene 1, −135°3.17 / 5Astra, scene 1, −135°3.38 / 5Opus 5.5, scene 1, −135°3.50 / 5
+135°Scene 1: ground truth at +135°Grok 4.7, scene 1, +135°3.25 / 5Astra, scene 1, +135°2.50 / 5Opus 5.5, scene 1, +135°2.63 / 5

Click to enlarge. Rows match the requested orbit angle; model cameras and pivots differ from ground truth. Scores shown compare model renders against the input photo, not these ground-truth views.

Grok vs AstraOrbit A/B: Astra wins 56–19, with 5 ties.

Grok vs OpusOrbit A/B: Grok wins 49–24, with 7 ties.

ALL3D maintains the source scenes, execution traces and evaluation pipeline. These are available for a closer technical comparison.

Full scores, costs and comparison conditions
All five scenes · equal weight · additional image metrics
AgentInput CLIP ↑Orbit CLIP ↑SSIM ↑LPIPS ↓EdgeMBA ↑
Grok 4.70.9160.8430.5400.6540.342
Astra0.9400.8630.5610.5190.436
Opus 5.50.9270.8470.5610.5510.385
Grok’s direct A/B results · all five scenes · wins / ties / losses
OpponentInputOrbit
Astra0 / 0 / 1019 / 5 / 56
Opus 5.54 / 0 / 649 / 7 / 24

How to read the comparison

Grok 4.7, GPT-6 Astra and Claude Opus 5.5 use shared pre-generated Meshy assets. Five selected scenes, one reconstruction per model per scene, nine native views: 0°, ±15°, ±45°, ±90° and ±135°. This changes the build/furnish agents, not every model in ALL3D.

Gemini 3.8 Flash rates each image twice against the input; pairs are anonymous and judged in both orders. A/B counts are judge calls, not independent scenes. CLIP uses a pinned OpenAI ViT-L/14 checkpoint. Pixel metrics cover the input view only. Unequal budgets, provider adaptations and camera estimates limit causal claims. This is not a general model leaderboard.

Grok and Opus disabled scene-library access. Astra scene 1 allowed it; reviewed reuse included a French-door reference and oak-floor materials. All tables include all five scenes.

Cost and time accounting

Four of five Grok builders needed continuation. Raising the cap from 100 to 500 turns let them finish in 125–225 total build turns. Astra and Opus ran with 100-turn caps.

Logged SDK model-cost estimates, not invoices. Build + furnish job time includes tool waits; it is not GPU-billed time or end-to-end pipeline latency. Grok includes its initial attempts and continuations; Astra and Opus use the selected scored runs. Earlier discarded experiments are excluded.

Grok total: $181.96 and 6.28 summed job hours. Parallel batch span: 103.1 minutes. Zero new Meshy credits. Compute, rendering, evaluation and preflight charges are excluded.

WebP previews are for display; scoring used the original PNGs. ALL3D conducted this comparison without provider endorsement.

Download scores, image hashes and accounting (JSON)