Claude Opus 5.5 vs GPT-6 Astra: Which AI Model Should Manufacturers Use to Read Drawings?

Every custom manufacturing job starts with a drawing set. Before a cabinet or countertop shop can quote, someone has to go through the architect’s sheets and pull out what needs to be built: how many cabinets, how much countertop, which elevations and callouts matter. At a mid-size casework shop that takeoff can eat two to six hours per project, and it is exactly the kind of work leaders now want AI to take on.

The question is which model. Two recent releases have made that choice harder. OpenAI’s GPT-6 Astra set a new bar for reading drawings, and Anthropic’s Claude Opus 5.5 came close at a fraction of the price. We ran both, alongside 13 other models, through the same test on real architectural drawings. Here is how they compare, and how to decide between them.

How we tested

Our AI drawing benchmark uses 120 sheets from nine real casework projects. An expert marked every object of five types by hand: cabinets, countertops, floor plans, interior elevations and callouts. That came to 1,430 objects, and it is the answer key.

Every model got the same pages and the same instructions: find these objects and draw a box around each one. A box only counts if it lands close enough to the expert’s box. The score is F1, which rises only when a model finds more real objects and invents fewer fake ones. A score of 100 would mean every object found and nothing made up.

The head-to-head

GPT-6 AstraClaude Opus 5.5
Overall (F1)9279
Cabinets8260
Countertops8451
Floor plans9991
Interior elevations9490
Callouts9684
Cost per page (median)$0.186$0.040
Time per page (median)28 s9 s

GPT-6 Astra is the best model we have tested. Claude Opus 5.5 is second of 15. The gap between them is real, but so is the gap in price and speed.

Where GPT-6 Astra wins

Astra is the first model we have seen that reads the hard parts of a drawing about as well as the easy parts. Floor plans are big, clearly bounded shapes, and most models handle them. Countertops are thin strips that run behind other lines, and callouts are small bubbles scattered across a busy sheet. Astra scores 84 on countertops and 96 on callouts. Of the other 14 models, only Opus 5.5 clears 35 on countertops.

For a manufacturer, the most important row is cabinets. They are what the shop builds and what the quote is priced on. Astra scores 82 there. Opus 5.5 scores 60, which means it misses or misplaces a meaningful share of the cabinets on a sheet. If the output feeds straight into a quote, that difference is money.

Where Claude Opus 5.5 wins

Opus 5.5 costs about $0.04 a page, roughly a fifth of what Astra costs. It is also the fastest model in the test at a median of 9 seconds a page, about three times quicker than Astra. On a 200-sheet project, that is roughly $8 and half an hour of processing versus $37 and an hour and a half.

It is also a big step forward from its own predecessor. Claude Opus 5 scored 40 overall and just 15 on callouts. Opus 5.5 scores 79 overall and 84 on callouts. Apart from Astra, it is the only model in the test to score above 50 on countertops and above 70 on callouts.

And on the large, structural parts of a drawing set it is close to Astra: 91 on floor plans and 90 on elevations. That makes it a strong fit for the jobs that come before the detailed count, such as sorting a 300-page set, finding the sheets that matter and flagging which rooms carry casework scope.

So which should you use?

It depends on the job, not on which model is newest.

Use GPT-6 Astra when the output goes into a quote. If a missed cabinet or countertop run turns into a pricing error, pay the extra 15 cents a page. On a typical project that is still a few dollars, and it is small next to a single estimating mistake.

Use Claude Opus 5.5 for volume and triage. When you are processing large drawing sets, screening bids you may not pursue, or finding the right sheets before a detailed takeoff, its speed and cost win. Many teams will get the best result by using both: Opus 5.5 to sort and scope, Astra to count.

Re-test every few months. Rankings in this category move fast: we have re-run the benchmark three times since August, and the leader has changed. A newer version is not automatically better, either: in the same run, GPT-6 Sol scored 49, below the GPT-5.6 Sol it replaced (65), and Grok 4.7 scored 20, below Grok 4.6 (30).

Test on your own drawings. A public benchmark tells you which models are worth trying. It cannot tell you how a model handles your architects, your drawing standards or your product mix. Before you commit, run a few of your real projects through the models on your shortlist and check the output against a takeoff your estimators have already done.

The bigger picture

A general-purpose model is a starting point, not a finished estimating tool. The manufacturers seeing the largest gains build the model into a workflow that matches how their estimators already work. One U.S. casework manufacturer cut cabinet takeoff from about six hours to roughly ten minutes this way.

Model choice still matters, though. And right now manufacturers have a clear choice: the strongest drawing reader on the market, or a very capable one at a fifth of the price.