We gave two AI setups the same request: build a little 3D living room you can decorate and rearrange. Here is what happened—and what you can try yourself.
Start with a little exploring
You do not need to understand programming. The useful question is: does the room do what you want?
1
Look around. Turn the room and zoom closer.
2
Make it yours. Change the sofa colour and move the chair.
3
Change the mood. Try evening lighting, then reset the room.
Try the rooms before reading the scores if you want your own first impression.
Astra’s room (private, not linked)
This room is private: Devon can show it using his signed-in account. Its layout adjusts to a phone-sized screen; physical-phone testing is still pending. On a narrow screen, the controls are below the room. Spark’s room was sent to Devon as a separate HTML attachment for a computer browser.
The result of this test
GPT-5.3 Codex Spark · Light
A quick, rough first version
38 / 100
Worked: looking around, zooming, changing the sofa colour and switching lighting.
Fell short: dragging furniture failed, walls were misplaced and some furniture floated. Reset left a colour indicator wrong.
About 1 minute 17 seconds for the resumed run. Extra help was then needed just to open the demo.
GPT-6 Astra · Ultra · Fast
A much more complete room
89 / 100
Worked: selecting, colouring, moving and turning furniture; lighting, camera controls and reset.
Fell short: more scrolling on small screens; the mouse wheel can zoom when you mean to scroll; headings can overlap a close-up view.
About 18 minutes 57 seconds for the resumed run. The delivered website opened without reviewer repairs.
The trade-off we observed: Spark finished much sooner. Astra took longer and delivered a room that was substantially more usable.
Why could the results be so different?
There are three separate choices behind those names.
A
The model: which AI does the work
Spark and Astra are different models. OpenAI describes Spark as a fast, less-capable model designed for rapid coding changes. A fast answer can still be useful—but a whole interactive room asks for many parts to work together. Source
B
The effort: how much reasoning is requested
“Light” requests lighter reasoning. “Ultra” requests much deeper reasoning. Extra effort gives a model more opportunity to work through a complicated task; it does not guarantee a correct result.
C
Fast mode: how quickly processing happens
“Fast” is separate from reasoning effort. It speeds up a supported model’s processing; it does not turn Astra into Spark or mean the whole project will finish instantly. Building and checking a larger result still takes time. Source
Our interpretation: the different model, effort setting and amount of work likely contributed to the gap. We changed several things at once, so this experiment cannot tell us exactly how much each caused.
What this tells us about AI
Free AI is still AI
A free chatbot counts as using AI. This experiment compares two coding setups; it does not compare all free services with all paid ones.
The tools matter too
Here, the AI could write files and work with a browser to build something you can use. Asking questions in a chat is another way of using AI.
Both wrote code for 3D rooms. The difference was how well the finished pieces fitted together—and how much outside help was needed.
One more experiment: your request
“Make a cosy reading corner by the window. Keep the sofa colour I chose.”
Ask Devon to give both AI tasks your exact words. Then check: did each understand you, preserve what you liked, and leave the controls working?
The room itself does not have a chat box. Changes like this are requested in the AI task that built it. Your temporary colour choice may need to be described again.
How the grades were decided
What was judged
Spark
Astra
Working controls / 40
18
38
Meeting the request / 20
12
20
Easy for a beginner / 20
5
15
Appearance / 15
3
12
Ready to open and checked / 5
0
4
Total / 100
38
89
These are an AI reviewer’s judgments from browser checks and source inspection, not an official benchmark. The reviewer knew the model identities. Appearance and ease-of-use scores involve judgment.
What keeps this comparison honest?
Both received the same original request, but they had different settings and did not use equal time budgets.
A Windows setup failure affected both before the resumed runs. That outage was excluded.
Spark needed a running preview server and a missing library mapping supplied by the reviewer. Its grade does not pretend those were Spark’s successes.
Astra’s eight objects passed tests at all four room corners while rotated 45°. Their reported bounds stayed inside the room and on the floor.
Furniture can overlap in Astra. Avoiding furniture-to-furniture overlap was not requested, so that was not penalized.
One room-building task cannot establish which model is best at every kind of work. More attempts and equal budgets would make a stronger comparison.