AI Spots IKEA Assembly Mistakes: What the New Furniture Test Shows

Handwerkzeuge und Arbeitshandschuhe auf einer Holzfläche
Photo by Todd Quackenbush on Unsplash

A reversed side panel often becomes obvious only when the furniture is almost finished. AI could eventually catch such mistakes earlier: On September 23, Epoch AI published a test in which models compare photos of IKEA furniture with the corresponding assembly manuals. GPT-6 Astra answers 80 percent of the tasks correctly. The finding shows tangible progress in spatial reasoning, though it does not yet establish a reliable digital furniture assembler.

Key takeaways

  • The Furniture Assembly Benchmark contains 60 photos from three furniture builds, with and without deliberately introduced mistakes.
  • GPT-6 Astra scores 80 percent, followed by Claude Fable 5.1 at 70 percent and Claude Opus 5 at 61 percent.
  • Alongside the photo, models receive a PDF assembly manual and tools for enlarging and analyzing images.
  • The fastest AI tested takes a median of three minutes per photo. The test does not demonstrate a finished real-time assembly assistant.

Understanding instructions takes more than recognizing parts

Finding an assembly mistake requires connecting several representations. The photo shows a three-dimensional object, while the manual contains schematic drawings. A board can look quite different from the camera’s perspective than it does on paper. What matters is not just whether the model recognizes a screw or drawer. It must understand which part belongs where and how it should be oriented.

Epoch AI uses photos of a STÄLL shoe cabinet, a TONSTAD bed frame, and a GULLABERG dresser. The team deliberately introduced mistakes and sometimes continued assembling before taking a photo. That prevents the task from being reduced to the most recently installed component. A visible inconsistency may originate in an earlier step. The test is independent of IKEA; it is neither a joint product announcement nor an assistant offered by the furniture manufacturer.

Models must identify the steps containing mistakes and describe the problems reasonably. If the build is correct, they should direct the user to the next step. They are also told which step was just completed. That is a useful constraint: In a living room, a system might first need to determine the current stage. The experimental conditions and a casual question accompanied by a phone photo are therefore not interchangeable.

The official manual remains the foundation. IKEA provides assembly documents on the TONSTAD bed’s product page. In a practical application, matching the exact furniture variant with the right document would be essential. A plausible explanation based on the wrong manual is no more useful during assembly than an accurately read drawing that belongs to a different component.

What the 80 percent score actually means

In November 2025, the best score among the models Epoch tested was 28 percent, achieved by Claude Opus 4.5. The current report therefore shows a substantial jump within ten months. It reconstructs progress using selected models rather than retrospectively measuring every system available at the time. That qualification matters more to the comparison than an impressive formula for the increase.

Today’s leading score is also a task accuracy figure on a small dataset. It does not mean that AI could safely assemble 80 percent of all furniture, or that every individual mistake would be found with that probability. One photo can contain several mistakes. The assessment asks whether the required answer for that image is correct. It offers no guarantee for other furniture, other rooms, or hidden assembly errors.

Grading has several stages. First, the answer is checked for the correct mistaken steps. GPT-5.6 Sol then evaluates the description, with instructions to be lenient if the mistake is roughly described correctly. An AI grader makes larger comparisons manageable, but it is not an independent human expert. Epoch has also not yet measured human performance on this test. Claims such as “better than experienced installers” would therefore be unsupported.

The most interesting signal is less the overall ranking than the interaction between document and photo. A model must check a technical specification against a real object. For future assistance products, that would offer a different benefit from merely describing an image: They could point out a concrete inconsistency and identify the relevant part of the manual. Whether that benefit can be delivered reliably must be demonstrated by an application outside the benchmark.

How to explore the idea sensibly today

A cautious personal experiment could focus on a single, manageable question during a furniture build. It requires a service that can process both photos and the relevant document. This is an adaptation of the test idea, not an Epoch-verified tutorial for a particular chat product. The published model scores likewise apply only to the experimental setup described.

It would help to provide the exact furniture name, the most recently completed step, and the relevant manual page together with a well-lit photo. The question should not assume there is already a mistake: “Does the visible assembly match this step? Identify the part of the manual supporting your assessment.” If the model claims that a part faces the wrong way, the reasoning can then be checked directly against the document.

For this experiment, the quality of the visible view matters. An overview shows how parts are arranged; a close-up may clarify a connection. Screws that cannot be seen, concealed fittings, and the strength of a connection remain uncertain from a photo. An answer should therefore be treated as a checkable suggestion. It replaces neither the complete manufacturer’s instructions nor an actual inspection of the assembled furniture.

The next advance must work during real assembly

A median of three minutes may be acceptable for an occasional question, but it is a noticeable pause for continuous guidance. There is also the unresolved question of how well the finding transfers to additional furniture and everyday photos. A useful assembly assistant would need to find mistakes, confirm correct steps, recognize missing visual information, and request specific follow-up photos when needed.

The new test makes that development measurable. It shows that connecting a manual to a physical object is no longer just an appealing idea for the future. The next convincing step would be broader testing with varied furniture, different camera angles, and people actually assembling it. Only then would a strong laboratory score become a dependable everyday benefit: less disassembly because a reversed part was spotted in time.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top