Comparison finds wide cost and quality gaps across AI models

Netlify’s comparison of 11 AI models shows that the cost of generating a simple website can vary dramatically, while the visible quality does not always rise in step with spending. Elad Rosenheim writes for Netlify that the company tested models with the same prompt to help users weigh output quality against the credits consumed by AI coding agents.

The test asked each model to build a one page website for a neighbourhood coffee shop. The site needed opening hours, an address, a short menu and a photo. The prompt also made clear that the page should remain static unless edited manually, so no database or content management system was needed.

Netlify ran every model three times with its default settings. It used its open source AXIS evaluation tool to assess whether generated projects work correctly and use platform services only when needed. The coffee shop comparison focused more directly on design, content and credit consumption than on the automated functional checks.

Costs range from 2.4 to 519 credits

Claude Opus 5 had the highest average cost at 519 credits per run. That figure was heavily affected by one unusually expensive run that consumed 1,055 credits. Its other two runs used 253 and 249 credits. Netlify says the expensive output included more detailed visual elements, a custom map and working dark mode, but it cautions that higher expenditure does not reliably produce a proportionately better result.

At the other end of the scale, DeepSeek V4 Flash used an average of 2.4 credits. Its outputs varied, but Netlify found that one of the low cost results resembled work from a more expensive mid tier closed model. Kimi K2.7 Code averaged 19 credits, GLM 5.2 averaged 27, DeepSeek V4 Pro averaged 37 and GPT 5.6 Terra averaged 39.

The middle of the range also produced different trade offs. GPT 5.6 Sol, set to low effort by default, averaged 141 credits and, in Netlify’s assessment, delivered richer content and stronger basic design choices than Claude Sonnet 5, which averaged 143 credits. Gemini 3.6 Flash averaged 103 credits and produced more polished pages than Gemini 3.1 Pro, which averaged 53 credits but generated notably plain results.

Netlify observed that GPT 5.6 Terra produced a distinct visual style despite its lower cost, although some runs showed issues such as a missing image or hard to read text. DeepSeek V4 Pro also produced a broken image in one test, where the generated page referenced a file that was not present in the project.

Design is only one part of the decision

The company stresses that a static marketing page is a limited test. For more advanced applications, it says the central question becomes whether a model can choose and correctly implement the necessary services. That includes databases, user accounts, file uploads, AI features, security and testing.

Some models also differ in the information they can work with. Netlify notes that GLM 5.2 is text only and cannot use uploaded screenshots as visual inspiration, unlike the Kimi models tested. That limitation may matter for teams that want AI to follow an existing brand style or layout.

The practical conclusion is not that one model wins every task. Netlify suggests that users with limited credits may prefer a less expensive model and use follow up prompts to refine the result. Those seeking a more complete first draft may accept the higher and sometimes unpredictable cost of a frontier model. The company plans further comparisons involving shared data, image uploads and AI powered applications.

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×